Friday, 24 November 2023

Vice-chancellor narcissism and university performance

Over the last two decades (or more), universities have increasingly come to be managed like businesses. In that case, the role of the university vice-chancellor (or president) has come to resemble that of a business CEO. As a consequence, the types of skills a successful vice-chancellor must possess have changed. And, the types of academics attracted to becoming a vice-chancellor have also changed. If I claimed that current vice-chancellors are, on average, more narcissistic than vice-chancellors from ten years ago, I think many academics would agree.

In fact, that is one of the findings of this new article by Shee-Yee Khoo (Bangor University), Pietro Perotti (University of Bath), Thanos Verousis (University of Essex), and Richard Watermeyer (University of Bristol), published in the journal Research Policy (sorry, I don't see an ungated version online). Khoo et al. aren't interested in the level of narcissism of vice-chancellors per se, but whether vice-chancellor narcissism is related to university performance. They use data from British universities from 2009/10 to 2019/20 to investigate this question. Their sample includes 133 universities, and 261 vice-chancellors.

Interestingly, they:

...measure narcissism based on the size of the signature of the VC...

First, we obtain the signature of each VC from the university annual report, or the university strategic plan or letter when the signature is unavailable in the annual report. Second, we draw a rectangle around each VC's signature, where the signature touches its furthermost endpoint, ignoring any dot at the end of the signature or/and underline below the signature... Third, we measure the area covered by the signature by multiplying the length and width (in centimetres) of the rectangle. Fourth, we divide the area by the number of letters in the VC's name to control for the length of the VC's name.

Apparently, this is a widely used measure of narcissism <quickly checking the size of my signature>, and Khoo et al. note that it has advantages over survey-based measures because it is "unobtrusive". It is also less subject to manipulation than survey-based measures, although I was worried that the size of the signature on an electronic document (like an annual report) would create a lot of measurement error. However, in the section on robustness checks, Khoo et al. report that:

...we compare the handwritten and electronic signature sizes of the same VC, for the sample where we have both types. The size of the signature remains the same irrespective of the signature type.

Ok then. For their measures of university performance:

Firstly, the Research Excellence Framework (REF) and its predecessor until 2014, the Research Assessment Exercise (RAE), is a system for assessing the research quality of UK universities and other HE institutions... We use the overall quality of research based on the REF (formerly known as the RAE) as our research quality indicator...

Secondly, we employ the National Student Survey (NSS) which assesses teaching quality in UK universities... In particular, we employ the Student Satisfaction Score, which is the average score from across the organisation and management, learning resources, learning community and student voice sections of the NSS...

Thirdly, we use the overall university ranking, based on the Guardian newspaper.

Based on the way that they describe their results though, it seems like they use the ranking of each university, rather than the scores themselves for research quality and teaching quality. For the analysis, they perform the analysis separately for 'old universities' (those that were created before 1992) and 'new universities' (those created after 1992). The main difference between those groups of universities is that the old universities tend to be more research-focused, while the new universities tend to be more teaching-focused. The analysis for each type of university compares the university ranking two years before and two years after a change in vice-chancellor:

In particular, we rely on the universities that face a VC transition, i.e., a change from a low narcissist VC to a high narcissist VC, within the sample period. We define a high narcissist VC as one in the top quartile of the distribution. In our analysis, our treatment group consists of universities that appointed a high narcissist VC during the sample period (i.e., low-to-high transition universities). The baseline group consists of universities that appointed a low narcissist VC during the sample period (i.e., low-to-low transition universities).

Collapsing vice-chancellor narcissism to a dichotomous variable (equal to one only when a university transitions from a low-narcissism vice-chancellor to a high-narcissism vice-chancellor) deals with some of the measurement error issues I noted above. However, it does open up their analysis to some criticism, since there are many alternative arbitrary cut-offs that they could have used (and they don't report the robustness of their results to this particular choice). With that limitation in mind, Khoo et al. find that:

...a change from a low narcissist VC to a high narcissist VC is associated with a deterioration in research performance for both New and Old universities. VC Change×Narcissism Change is negative and significant, confirming that VC narcissism has a negative effect on research performance. Controlling for university as well as VC characteristics, and year and university fixed effects, a change from a low narcissist VC to a high narcissist VC is associated with a drop of approximately 16 places (VC Change×Narcissism Change = –16.07) in research performance for the sample of New universities and nine places for the sample of Old universities (VC Change×Narcissism Change = –9.57). This finding also demonstrates that New universities are more susceptible to VC transitions.

Then for teaching quality:

The coefficient of VC Change×Narcissism Change is negative and significant, indicating a drop of approximately 12 places for the group of New universities and 19 places for the group of Old universities (VC Change×Narcissism Change = –12.06 for New universities and –19.52 for Old universities).

And for overall ranking:

The transition from a low to a high narcissist VC is associated with a drop of approximately 27 places in the Guardian ranking for the group of New universities but has no effect on the Guardian ranking of Old universities.

So, narcissistic vice-chancellors lower the research and teaching performance of universities, but have more negative impact on new universities (except for teaching quality, where the impact is greater for old universities). Why do vice-chancellors negatively impact performance? It turns out that the answer may be different for old universities and new universities. Khoo et al. go on to show several additional results, including that:

...for Old universities, financial risk substantially increases with VC narcissism. Specifically, the appointment of a highly narcissistic VC deteriorates the financial sustainability of Old universities by approximately five to six [Financial Security Index] points...

Hence, the appointment of a highly narcissistic VC is associated with higher financial risk (i.e., lower financial sustainability) and lower effectiveness of the use of the resources. These results are consistent with excessive risk-taking behaviour. For Old universities, the findings suggest that highly narcissistic VCs take unnecessary risk, which might lead to a decrease in university performance...

The results for Capital Expenditures are insignificant. However, when using Expenses to Revenue as the dependent variable, the coefficient on VC Change×Narcissism Change is positive and significant at the 5 % level for the group of New universities. This evidence, although based on only one of the two measures, is consistent with highly narcissistic VCs engaging in empire-building strategies in New universities, which might be detrimental to the performance of the organisation.

I don't find the ratio of expenses to revenue very convincing as a variable, but Khoo et al. use it to suggest that narcissistic vice-chancellors in new universities are engaging in empire building. In contrast, Khoo et al argue that narcissistic vice-chancellors in old universities are engaging in excessive risk-taking behaviour. Both of these behaviours have been noted of narcissistic business CEOs as well.

So, what should universities do in order to mitigate the negative impacts of vice-chancellor narcissism? Khoo et al. recommend that:

...university councils and relevant committees should take into account and, if possible, measure the narcissism of the candidates for the role of VC... Given, however, that narcissists tend to appeal to recruiters, we also recommend that VC selection committees should undertake rigorous training that will allow them to control for this implicit bias in favour of narcissistic applicants.

I don't find those recommendations to be very convincing (and they aren't really supported by Khoo et al.'s results). What is supported is higher-quality governance, since in their final set of results:

The core finding that VC narcissism has a negative effect on research performance still holds after controlling for the effect of university governance; however, we note a decrease in the magnitude of the effect...

The results here are not consistent across all measures of university performance, but in all cases the measure of university governance does appear to moderate the effect of a shift from a low-narcissism vice-chancellor to a high-narcissism vice-chancellor.

So, the takeaway message from this research is the narcissistic vice-chancellors harm university performance, but having strong university governance can limit the damage. The only question remaining is, how did the vice-chancellors of these researchers' universities respond when they learned of this research?

Thursday, 23 November 2023

New results on the bat-and-ball problem

A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?

If you guessed ten cents, you would be in the majority. You would also be quite wrong. The correct answer is five cents. This 'bat-and-ball' problem is quite famous (and you may have seen it, or a question like it, before - a variant was in a pub quiz that I competed in a few weeks ago, for example). The problem is one of three questions included in the Cognitive Reflection Test, which purports to measure whether people engage in cognitive reflection, or are more prone give into 'intuitive thinking'. It also relates to what Daniel Kahneman referred to in his book Thinking Fast and Slow as System 1 and System 2 thinking. System 1 is intuitive and automatic (and gives a ready answer of ten cents to the bat-and-ball problem), while System 2 is slower and reflective (and is more likely to lead to the correct answer of five cents).

However, a new article by Andrew Meyer (Chinese University of Hong Kong) and Shane Frederick (Yale University), published in the journal Cognition (open access), may give us reason to question the theory of System 1 and System 2 thinking (or reason to question the validity of the bat-and-ball question). Frederick is the author who introduced the Cognitive Reflection Test, so the results reported in this paper should be considered especially notable.

Meyer and Frederick conducted a number of studies of the bat-and-ball problem, showing a number of increasingly disquieting results. First:

...verifying the intuitive response requires nothing more than adding $1.00 and $0.10 to ensure that they sum to $1.10 (they do) and subtracting $0.10 from $1.00 to ensure that they differ by $1.00 (they don't). Since essentially everyone can perform these verification tests, the high error rate means that they aren't being performed or that respondents are drawing the wrong conclusion despite performing them.

If respondents aren't attempting to verify their answer, encouraging them to do so may help. We tested this in five studies involving a total of 3219 participants who were randomly assigned to either a control condition or to one of four warning conditions shown below. Two studies were administered to students who used paper and pencil. The rest were web-based surveys of a broader population...

The warnings improved performance, but not by much... This suggests that they failed to engage a checking process, or that the checking process was insufficient to remedy the error...

Specifically, only 13 percent of research participants in the pure control group got the bat-and-ball problem correct. In the treatment group that received the simplest warning (which simply warned: "Be careful! Many people miss this problem"), this increased to 23 percent. There were modest increases in performance across other studies that Meyer and Frederick report (with various different wordings of the warning), ranging from -9 percentage points to +17 percentage points. They don't report a measure of statistical significance, but the magnitude of the change is not large, and warnings to check the answer don't eliminate the intuitive response. Evidently, research participants aren't great at checking their answer. Or maybe, they simply don't perform any check at all. What about being more directive that research participants should check their answer if their original answer was ten cents:

Since these warnings were ineffective, we next tried an even stronger manipulation by telling respondents that 10 cents is not the answer. We conducted eight such experiments, with a total of 7766 participants. In five studies (three online and two paper and pencil), participants were randomly assigned to either the control condition or to a Hint condition in which the words “HINT: 10 cents is not the answer” appeared next to the response blank...

In three other studies (two online and one in-lab), we used a within-participant design in which the Hint was provided after the participant's initial response. In those studies, respondents could revise their initial (unhinted) response, and we recorded both their initial and final responses...

The hint that the answer wasn't 10 cents helped substantially, but, more notably, many – and sometimes most – still failed to solve the problem...

Receiving the hint increased performance in the bat-and-ball problem by between +17 percentage points and +23 percentage points in a between-subjects comparison (comparing research participants who received the hint with those that didn't receive the hint), and between +16 and +22 percentage points in a within-subjects comparison (where research participants could change their answer after they received the hint). The latter results lead Meyer and Frederick to note that:

Though the bat and ball problem is often used to categorize people as reflective (those who say 5) or intuitive (those who say 10), these results suggest that the “intuitive” group can – and should – be further divided into the “careless” (who answer 10, but revise to 5 when told they are wrong) and the “hopeless” (who are unable or unwilling to compute the correct response, even when told that 10 is not the answer).

Why would so many research participants still maintain that the answer is ten cents, even when they are explicitly told that ten cents is not the correct answer? Meyer and Frederick suggest that:

This result has hallmarks of simultaneous contradictory belief (Sloman, 1996), because respondents who report that $1.00 and $0.10 differ by $1.00 obviously do not actually believe this. It is also akin to research on Wason's four card task showing that participants will rationalize their faulty selections, rather than change them (Beattie & Baron, 1988; Wason & Evans, 1974). It could also be considered as an Einstellung effect (Luchins, 1942), in which prior operations blind respondents to an important feature of the current task or as an illustration of confirmation bias, in which initial erroneous interpretations interfere with the processes needed to arrive at a correct interpretation (Bruner & Potter, 1964; Nickerson, 1998).

I would put a lot of this down to motivated reasoning. However, it gets even worse:

...we ran two studies on GCS in which we asked respondents to either consider the correct answer (N = 2002) or to simply enter it (N = 1001)...

Asking respondents to consider the correct answer more than doubled solution rates, but only to 31%. Asking them to simply enter the correct answer worked better, as 77% did so, though, notably, the intuitive response emerged even here.

So, when research participants are asked to consider if the answer could be five cents, more than half still get it wrong. And even when research participants were told that the answer is five cents, and directed to write down five cents as the answer, nearly a quarter of research participants still get the answer wrong. That leads Meyer and Frederick to conclude that:

...the very existence of such manipulations (and their lack of complete efficacy) undermines a conclusion many draw from dual process theories of reasoning: that judgmental errors can be avoided merely by getting respondents to slow down and think harder...

Meyer and Frederick use all of these results (and others) to suggest that people engage in an 'approximate checker' process, wherein if the intuitive result provided by System 1 is approximately correct, then the more deliberative System 2 doesn't go through a complete process of checking. They demonstrate this with some further results that show that:

As the price difference between the bat and ball decreases, participants slow down... and solution rates rise markedly – from 14% to 57%...

So, perhaps these results are not fatal for the idea of System 1 and System 2 thinking, but psychologists and behavioural scientists need to re-think the conditions under which System 2 operates, and whether it always operates optimally. The results also suggest that the bat-and-ball problem may not actually show quite what it purports to - at least, it doesn't necessarily show cognitive reflection, as even when such reflection is explicitly invoked (through asking research participants to check their answer, or telling them to consider if the answer might be five cents), many do not exhibit such reflection (or else, they reflect and still get the answer wrong. Meyer and Frederick finish by noting that:

...the remarkable durability of that error paints a more pessimistic picture of human reasoning than we were initially inclined to accept; those whose thoughts most require additional deliberation benefit little from whatever additional deliberation can be induced.

[HT: Marginal Revolution. back in September]

Wednesday, 22 November 2023

AI and the research production function

I've been a bit quiet on the blog lately, due to travelling, and attending the North American Regional Science Conference in San Diego. Regional science is a multidisciplinary field that overlaps economics, geography, planning, and many other disciplines. There were many interesting sessions at the conference, although apparently not the session that I was presenting in, which had an audience of two (one of which was another speaker in the session, with the other three speakers in my session no-showing).

Anyway, one of the most interesting sessions was the day before the conference proper started, on "The Potential of AI for Regional Science", hosted by The Regional Science Academy. Speakers included many of the contemporary great minds in regional science such as Rick Church (University of California, Santa Barbara), Tomaz Dentinho (University of Azores), Peter Nijkamp (Open University), and John Östh (Oslo Metropolitan University), among others. They noted (with examples) the great opportunities that artificial intelligence offers, especially as an aid for research. For example, Östh demonstrated a really interesting application of AI to generating data on neighbourhoods, while Patricio Aroca (Universidad Andres Bello) showed how AI could help researchers whose first language is not English to improve their chances of publication in top-ranked English language journals.

I came away from this session in equal parts excited and concerned (which might be an appropriate mix of reactions to any sufficiently disruptive technology). There is great potential in using AI as a research tool, and I've already seen many examples (and blogged about some ideas here and here). However, my concern comes from how AI will change the research production function.

Before we get that far, let's go back to an earlier technological revolution that greatly changed how research in economics was conducted. This story is not mine, but I forget where I got it from. Consider a simple research production function in economics, with two inputs: (1) econometric modelling; and (2) economic explanations. As computers increased in computational power, the relative price of econometric modelling decreased. So, researchers reallocated their scarce research resources from economic explanations (which had become relatively more expensive as an input) towards econometric modelling (which had become relatively less expensive as an input). The result was a rapid increase in econometric modelling, and a corresponding decline in the quality of economic explanations to go along with the econometric modelling. It is not clear that this was a positive change for the overall quality of economics research (but there certainly was an increase in the quantity of research).

Now consider a similar research production function in economics, where the inputs are: (1) human intuition and explanations; and (2) artificial intelligence. New AI models have made AI radically less expensive to use. We can probably expect research in economics (and in other fields) to adjust to much more use of AI, and much less use of human intuition and explanations. It remains to be seen whether that leads to an increase or a decrease in the quality of research overall. My money is on an overall increase in the quantity of research, but a reduction in the average quality and an increase in the variance (with some very high quality research resulting from AI, as well as a lot of very low quality research).

Unlike the science fiction and fantasy magazine Clarkesworld, which shut off submissions earlier this year due to a flood of AI-generated stories were submitted, I'm not aware of any academic publishers that are feeling the strain, yet. It will be coming though, and may intersect with the increasing volume of predatory publishers (see here and here). As the Managing Editor of a journal myself (the Australasian Journal of Regional Studies), I'm really not looking forward to policing the submission of AI-generated articles of low quality.

Taken altogether, it is clear that there is a trade-off inherent in the impact of AI on research (in regional science, in economics, and in other fields). We will need to accept (and act to mitigate) the negative impacts, while endeavouring to maximise the good.

Wednesday, 15 November 2023

Jibbitz trading bans offer a missed opportunity to introduce some basic economics to children

I read this article from the New Zealand Herald from earlier this week with some interest:

Jibbitz - accessories that clip on to Crocs - are being banned in schools in Northland due to escalating arguments between youngsters over the sought-after items.

Kamo Primary School principal Sally Wilson was forced to take action after students became upset over Jibbitz trades and some resorted to stealing...

Wilson said attempts to create a safe environment for trades were a learning curve for tamariki and sometimes “ended in tears”.

Often, tamariki trade an item in the hopes of getting it back, and when they realise that isn’t going to happen, they “emotionally can’t cope”, she said.

Eventually, Wilson banned Jibbitz from the school because they had become “disruptive”.

“They were getting stashes and holding on to them, and there was an uneven trade for a certain one that they were after.”

While some kids have their “eye on the prize” and trade cheap Jibbitz for more expensive ones, Wilson said there have also been cases of stealing.

“It’s a learning curve about possessions.”

My children are well beyond the age of adding accessories to Crocs (or going to school, for that matter), but I can remember past crazes for trading Pokemon or Yu-Gi-Oh cards, where trading for sought-after cards got a bit out of hand. So, I can understand the attractiveness of a ban, to protect vulnerable, younger traders from being taken advantage of by older, more savvy traders, who understand that trades are 'for keeps' and better recognise the real value of what is being traded.

However, I can't help but feel that there is a missed opportunity here. The gains from trade is a cornerstone of economic principles, and can be taught very easily. And Jibbitz offer the opportunity to teach the gains from trade in a way that young children can readily understand. Jibbitz trading also offers an opportunity for young children to understand that there are two sides to every trade, and that trading is a voluntary activity. So, thinking about Jibbitz trading, whenever there is a trade of Jibbitz from one child to another (there are two traders), the trade will only happen if both children agree to the trade (either trader can say "no" to a trade), and each child will only agree to the trade if they think that what they are receiving is better than what they are giving away (there are gains from trade for both traders).

Instead of a ban, putting some simple rules for trading Jibbitz in place could really help. Here's a few. First. all trades are voluntary. Both children have to agree. Forced trades can be cancelled by a teacher. Second, all trades are 'for keeps'. There are no take-backs. As an extra rule perhaps a 'current price list' could be maintained, where the number of common Jibbitz expected to be traded for particularly rare or valuable Jibbitz are recorded. This may help avoid problems of a thin market for rare Jibbitz.

There are also opportunities for young children to better understand demand (some Jibbitz are more sought after than others), relative prices (the most sought-after Jibbitz may be traded for several less-sought-after Jibbitz), and scarcity (the rarest Jibbitz are likely to be the most valuable). All of these economic principles can be taught simply, and without recourse to economic jargon, and would help children to better understand some simple economics.

Of course, then there is this objection to Jibbitz trading:

Dargaville mother Taiāwhio Wati-Kaipo was first annoyed when Jibbitz were recently banned at Dargaville Primary School, worrying for her children’s ability to express their “individuality”.

But Wati-Kaipo soon considered the issue and realised the ban had a “deeper meaning”.

She believed the ownership of Jibbitz is a “social indication” of where someone is “sitting on the financial bracket”.

Wati-Kaipo said the price hike in Crocs themselves has created a “has and has not” situation among students.

“Without the Jibbitz, the Crocs were already speaking volumes about someone’s identity,” she said.

Wilson said the craze created a social comparison, as it was about who had the coolest ones and who had the most.

I know that at least some people really believe that inequality can be addressed by banning markets, but it's not correct. In this case, banning Jibbitz will simply shift the outward expression and indicators of social status to some other margin. There are many ways that social status is conveyed. Why stop at banning Jibbitz? Why not ban Crocs altogether? Or premium school bags? Or premium stationery? Or phone games or apps where players can buy special skins or other in-game virtual merchandise? Or phones entirely? Anyway, I'm getting off topic. The Jibbitz ban is a missed opportunity to help young children to better understand some key economic concepts that will be helpful in their development as economic citizens. A ban isn't necessary, and there are better responses to protect children from exploiting each other in the Jibbitz market.