Showing posts with label Research quality. Show all posts
Showing posts with label Research quality. Show all posts

Monday, 30 March 2026

The tone and expression of academics on X (or Twitter)

In my previous post, I highlighted the apparent contribution of X (formerly Twitter) to toxicity on the Economics Job Market Rumors (EJMR) website. A natural follow-up question is whether and to what extent academics on X contribute to the toxicity on that platform and, by extension, to other forums such as EJMR. This recent article by Prashant Garg (Imperial College London) and Thiemo Fetzer (University of Warwick), published in the journal Nature Human Behaviour (open access), goes some way towards providing an answer.

Garg and Fetzer constructed a dataset of nearly 100,000 academics, including all of their Twitter [*] activity from 2016 to 2022. They then use large language models (ChatGPT-3.5 and GPT-4) to characterise each tweet in relation to content and tone. They assess each academic's stance on climate change, economic policy, and cultural issues. In terms of tone, they measure egocentrism (how often the academic refers to themselves in the first person), toxicity (based on the probability a tweet is classified as toxic by Google's Perspective API), and the balance between reason and emotion (measured as a ratio of 'affective terms' to 'cognitive terms' based on the Linguistic Inquiry and Word Count tool). The analysis is then largely descriptive, but nonetheless interesting.

Garg and Fetzer first find that:

...leading academics are not typically social media influencers... We found weak correlations between citation counts and Twitter metrics: citations and likes... citations and followers... and citations and content creation...

Garg and Fetzer observe that:

The weak correlation underscores that many prominent public intellectuals online gain visibility through public engagement rather than scholarly achievements, often holding lower academic credentials while commanding significant public attention, thus widening the gap between social media influencers and established academic experts.

I think that Garg and Fetzer overstate the case here. The weak correlations suggest that Twitter includes a cross-section of academics (in terms of academic quality), rather than that the top academics eschew Twitter (which would instead lead to negative correlations between measures of academic quality and Twitter engagement).

I'll put aside their results on political expression, which I round rather uninteresting. In contrast, the results in terms of tone demonstrate some interesting correlations. First, in terms of egocentrism (using self-referential terms such as 'I', 'me', 'my', and 'myself'):

Female academics... exhibit higher egocentrism than male academics...

Egocentrism increases with university ranking: academics at top-100 institutions... exhibit higher egocentrism than those from institutions ranked 101-500... US-based academics... show higher egocentrism than non-US academics

Then, in terms of toxicity:

Academics with high reach but low academic credibility... exhibit lower toxicity than those with the contrasting profile, that is, ones with low reach but high credibility...

Academics at top-100 universities... exhibit higher toxicity than those at institutions ranked 101-500... Moreover, US-based academics... exhibit higher toxicity than non-US academics...

And in terms of emotionality (or reason):

Emotionality is significantly higher among female academics... than male academics... In terms of reach and credibility, high-reach/low-credibility scholars... show significantly higher emotionality than low-reach/high-credibility scholars...

Finally, US-based academics... exhibit higher emotionality than non-US scholars...

Many of those differences will surprise no one, such as US-based academics being more egocentric and toxic in their expression on Twitter. Other differences seem to confirm familiar stereotypes, such as female academics using more emotional language than male academics. No doubt, some of the differences relate to differences in norms across different disciplines in terms of communication styles (both on Twitter and in general academic discourse). Garg and Fetzer don't control those other factors that might affect tone and expression. And before we get carried away about how toxic academics are on Twitter, Garg and Fetzer provide an important comparison with the general population. From Figure 6 in the paper:

Notice that academics (the blue line) exhibit far less toxicity (in the graph in the top middle) than the general population of Twitter users (the red line). Moreover, the trend in toxicity is downwards (for academics over the whole period from 2016 to 2023, and for the general population from 2021 to 2023). So, academics are not the main problem in terms of toxicity in the discourse on Twitter.

Nevertheless, there are important differences across academics, and one difference in particular stands out. Academics with high reach (those that are very active on Twitter) but low academic credibility (they are not highly credible academics, as measured by citations) exhibit less toxic expression on Twitter than other academics, particularly those who have low reach but high academic credibility. In their conclusion, Garg and Fetzer focus on this as a problem because:

...those with the greatest public reach may not represent top scholars, potentially distorting public perceptions

However, I see the opposite problem. In terms of tone, the top scholars with the lowest reach have the most toxic expression. Are those the sorts of academics that we want to promote even further on social media? I would suggest not.

What is a better option? First, more highly credible academics should be encouraged to engage in the social media discourse. However, it is important to recognise that credibility alone is not enough. What is needed are credible academics who also model constructive discourse without the toxicity, raising the standard of debate. However, as noted in yesterday's post, many high-quality (especially female) scholars are targets of hostility on social media. These are not separate issues.

Alternatively, we could raise the standard of academic discourse on Twitter more generally, without changing who is represented on the platform. That would reduce the toxic nature of the interactions. Stop laughing! It could happen. The tone and expression of academics on X (or Twitter) matters. Academics can set the standards for everyone else. We don't need to descend into the toxic culture wars that play out each day on social media. We are better than that, and if we show ourselves to be such, maybe more people will listen.

[HT: Marginal Revolution, last year]

*****

[*] I refer to the platform mostly as Twitter, because it didn't change names to X until July 2023, after Garg and Fetzer's dataset ends.

Tuesday, 24 March 2026

Evidence that artificial intelligence is increasing the impact, but narrowing the scope, of research

There is growing evidence of positive impacts of generative artificial intelligence on productivity. This includes productivity in research (see this post, for example), including my own. However, some have questioned whether increasing research productivity comes at a cost of narrowing the scope of research.

So, I was interested to read this article by Qianyue Hao (Tsinghua University) and co-authors, published in the prestigious journal Nature (ungated earlier version here) late last year. They look at the impact of AI tools (not limited to generative AI) on the productivity of researchers and the quality of research. Specifically, they look at authors publishing in six representative fields: biology, medicine, chemistry, physics, materials science, and geology, across three 'eras': (1) the 'machine learning era ' (from 1980 to 2014), the 'deep learning era' (from 2015 to 2022), and the 'generative AI era' (from 2023 onwards). Hao et al. compare authors who publish 'AI augmented papers' with those who do not. An 'AI augmented paper' is one that uses methods such as:

...support vector machines and principal component analysis from the machine learning era, and convolutional neural networks and generative adversarial networks from the deep learning era. Large language models, which have emerged in recent years, also rank among the most frequently used methods...

Using a dataset that includes over 27 million papers with complete records that were published between 1980 and 2025, of which about 310,000 were 'AI augmented', Hao et al. find that:

...annual citations to AI papers are 98.70% higher than those to non-AI papers on average...

So, AI augmented research gathers more citations, which suggests that authors using AI in their research achieve greater impact. This is reinforced by evidence that AI augmented papers are published in higher quality journals (with Q1 journals being the highest ranked). Hao et al. report that:

...the proportion of AI papers in Q1 journals is 18.60% higher than that of non-AI papers in all journals; in Q2 journals, the AI proportion is 1.59% higher; whereas Q3 and Q4 journals hold a relatively lower proportion of papers with AI... These results indicate a heterogeneous distribution of AI-augmented papers across journals, with a higher prevalence in high-impact journals.

And AI appears to make authors more productive, as:

On average, researchers adopting AI annually publish 3.02 times more papers... and garner 4.84 times more citations... than those not adopting AI, with consistency.

All of these results seem to hold across all of the disciplines that Hao et al. consider. However, it is not all good news. Hao et al. use machine learning to create a measure of the 'breadth of scholarly attention'. Using that measure, they find that:

Compared with conventional research, AI research is associated with a 4.63% contracted median collective knowledge extent across science, which is consistent across all six disciplines... Moreover, when dividing these disciplines into more than two hundred sub-fields, the contraction of knowledge extent can be observed in more than 70% of them...

Of course, some of the differences here may be due to selection, as the types of researchers, and the types of research, involving AI use may be meaningfully different from those that don't. However, putting the selection issues aside, Hao et al. note that there is a tension between the individual researcher's incentive to produce a greater quantity of research that has higher impact, which would suggest greater use of AI, and the social incentive to produce a greater breadth of research.

So, the takeaway from this paper is that we need to consider researcher incentives, not just productivity. Specifically, this research suggests that the use of AI in research is leading to a 'prisoners' dilemma' outcome: each individual researcher acting in their own best interests (and using AI in their research) leads to an outcome that is worse for society overall (less breadth of research and more incremental gains).

Hao et al. conclude that:

The substantial academic benefits of AI use may be a driving force behind its accelerated rate of adoption; however, we also find unintended consequences from the increased prevalence of AI-augmented research. In all fields, AI-augmented research focuses on a narrower scope of scientific topics and reduces the scientific engagement of follow-on research, leading to more overlapping research work that slows the expansion of knowledge. Further, with a greater concentration of collective attention to the same AI papers, the adoption of AI seems to induce authors to converge on the same solutions to known problems rather than create new ones.

So, what is the solution here? Society probably wants research to be higher quality and have a broad scope. But individual researchers' incentives to use AI in their research appears inconsistent with that outcome. The traditional prisoners' dilemma is a repeated game (see here or here, for example), and the players of that game can avoid the worst outcome by cooperating. In this case, the researchers could cooperate by agreeing not to use AI in their research. The problem is that every researcher has an incentive to cheat on that agreement, since if they use AI, then that will be good for their career. This prisoners' dilemma is more difficult to ensure cooperation in than the traditional game, because there are not just two players who need to cooperate, but thousands (or millions). Ensuring cooperation in a prisoners' dilemma game with many players, each of whom is far better off cheating than cooperating, is almost impossible (which is why solving the problem of climate change is so difficult).

My own view is that the answer is not to keep AI out of research. That is not realistic, in the same way that it's not realistic to expect students not to use generative AI. The incentives need to be redesigned, but this will be no easy task. As long as universities, research funders, and publishers reward researchers for quantity, citations, and publication in top-ranked outlets, then we should expect more AI-augmented work, with a narrower scope than society might prefer. If we want AI to expand knowledge rather than simply accelerate competition within narrow foci, then we need institutions that also reward novelty, breadth, and the discovery of new questions. That is the economic challenge we must face up to.

[HT: Marginal Revolution]

Tuesday, 17 March 2026

Seven decades of change in the demographics and research styles of top economics research

Back in 2013, Daniel Hamermesh (University of Texas at Austin) published this article in the Journal of Economic Literature (ungated earlier version here), which summarised changes in the demographics and research styles of top economics research, based on articles published between 1963 and 2011 in three top journals: the American Economic Review (AER), the Journal of Political Economy (JPE), and the Quarterly Journal of Economics (QJE). A new update last year (open access) from Hamermesh extends the analysis to include articles up to 2024.

In terms of demographics, the trends show a continuation and in terms of gender, Hamermesh notes that:

The progression that occurred from the 1960s and 1970s, when only a minute fraction of authors were women, to the early twenty-first century has, if anything, accelerated.

This will be welcome news, given the persistent gender gap in economics (see this post and the links at the end of it). It likely reflects the changing demographics of young economists, with a growing proportion of the young 'stars' in economics being women (and noting that it is young stars who often get published in the top journals that Hamermesh is considering).

In terms of the age structure of authors, Hamermesh reports that:

The changes from 2011 to 2024 continued those that started in the 1980s, but the rate of change has not accelerated. Indeed, most noticeable from 2011 to 2024 was a continuing sharp and statistically significant drop in the representation of the youngest group (and a nearly equal sharp rise among those 36–50)...

...the average age of authorship has increased steadily since 1973. 

Can I change my comment above about the young stars in economics? The increasing median age of authors in top journals seems to be a general trend across academia. Hamermesh then turns to research 'style', documenting a continued dramatic rise in the proportion of articles in those journals that are co-authored:

There were no four-authored papers as recently as 1983; today they account for 17 percent of articles. There were no papers with more than four authors in 2003; today nearly 12 percent of articles have five or more authors (with five articles written by six authors each and one by seven authors). Obversely, sole-authored papers are now quite scarce; and even two-authored papers today only account for slightly more than one-fourth of all articles (compared to a majority as recently as 2003).

Unsurprisingly, the increase in the number of co-authored articles means that the age diversity of author collaborations has increased over time as well. In terms of the types of research, he reports that:

The big changes are the continuing rise in empirical work based on original non-laboratory data and the rapid and even accelerating increase in experimental work. Today these two methods, which both involve collecting original data, account for over half of all published papers, compared to less than 4 percent four decades ago...

These trends are not all unrelated, of course. Experimental research, and the increasing use of large datasets, typically both require larger research teams. They also often require more detailed methods, which may involve both larger teams, and more experienced researchers. Larger teams might be more likely to include female team members. And larger teams often need someone to lead and coordinate all of the team members, and those leaders tend to be more experienced (and older) academics. So, it would not surprise me, if more detailed analysis was conducted, to see that the trends are interconnected.

Now, the interesting thing will be what happens going forward, given the increasing use of generative AI in research (see here, for example). Since generative AI can now do a lot of the work that research assistants and early career researchers previously did, will the trend towards larger research teams be reversed? How will that interact with the gender gap in research (given that the age of female economists skews younger at the moment). And how will it affect the age distribution of researchers (given that men, and younger people, are somewhat more likely to use generative AI). I'll be looking forward to Hamermesh's next update. Hopefully, we don't have to wait another 12 years.

[HT: Marginal Revolution, last year]

Monday, 16 March 2026

Changing their minds could be a good thing for economists

People don't like to change their minds. This may partly be an expression of loss aversion - we really want to avoid losses, including the loss of an idea that we previously thought was true. This leads to status quo bias - we prefer not to change things, and keep them the same, because changing things entails a loss. But what if changing our minds could make us better off? Would we be so reluctant to do so?

This 2025 paper by Matt Knepper (University of Georgia) and Brian Wheaton (UCLA) suggests that economists, at least, should not be afraid to change their minds, because doing so increases the number of citations to their research. Knepper and Wheaton investigate authors who undergo an 'ideological reversal' - previously publishing research that could be considered right-wing, before switching and publishing a paper that draws a left-wing-consistent conclusion, or the reverse (switching from left-wing to right-wing). Their main data source is every economics paper ever published in the top 100 economics journals indexed in Web of Science - some 200,000 articles. They also have a narrower dataset of papers referenced in meta-analyses on policy topics, including:

...the minimum wage, the economics of unions, the taxable income elasticity, the fiscal multiplier, intergenerational transfers, trade and productivity, trade and domestic employment, crowd-out, the gender wage gap, unemployment insurance, disability benefits, universal preschool, childcare and employment, immigration and wages, and more.

Knepper and Wheaton use this narrower dataset to train a machine learning model to categorise the rest of the papers in the dataset, as to how left-wing (or right-wing) the conclusions are. For instance, a paper that concludes that the minimum wage reduces employment is more right-wing, whereas one that concludes that there is no disemployment effect of the minimum wage is more left-wing. Knepper and Wheaton define an author as left-wing if they published more left-wing papers than right-wing ones over the previous five years, and the reverse for right-wing authors. They then use the larger dataset to investigate what happens to each economist who undergoes an 'ideological reversal'. They first outline some descriptive facts based on their dataset, including:

  • Fact #1: The typical author mostly publishes results on one side of the political spectrum.

  • Fact #2: Ideological reversals are not rare; they occur at least once for 40% of authors.

  • Fact #3: Ideological reversals become much more common later in an author’s career, with authors essentially never undergoing a reversal in the first decade of their career.

  • Fact #4: Most ideological reversals do not represent a permanent defection to the other side of the political spectrum, but rather the beginning of repeatedly publishing results on both sides of the spectrum.

  • Fact #5: Ideological reversals occur much more frequently amongst authors who are (initially) classified as right-wing.

That does seem like a surprisingly high proportion of economists who undergo at least one ideological reversal. However, perhaps we should take comfort in that - if the results point in a particular direction, our conclusions should say that, even if that conclusion is inconsistent with our previous conclusions on the same topic.

Do these ideological reversals matter though? Knepper and Wheaton employ a difference-in-differences analysis, comparing the difference in citations (and other metrics) between authors who did, and did not, undergo an ideological reversal, between the time before, and after, the reversal occurred. In other words, they look at whether citation counts rise more for economists who have an ideological reversal than for otherwise similar economists who do not. The results are striking, with:

...a sharp clear increase in citation count following an ideological reversal with essentially no evidence of pre-trends... The citation boost accumulates to approximately 9 over a one-decade period and 30 over a two-decade period.

The results remain consistent when Knepper and Wheaton limit the analysis to papers published before the ideological reversal, and when they limit the analysis to papers in the meta-analysis only (showing that the machine learning approach doesn't drive the results). Knepper and Wheaton also find evidence consistent with no change in the quality of papers before and after the ideological reversal, and that:

Both left-to-right and right-to-left reversals are rewarded by increased citations of roughly the same magnitude. The boost in citations received subsequent to a left-to-right reversal is mostly driven by citations from right-wing authors, and the boost in citations received subsequent to a right-to-left reversal is mostly driven by citations from left-wing authors. Encouragingly, however, the new right-wing (left-wing) audience garnered by a left-to-right (right-to-left) reversal... also engages with and cites the author's previous left-wing (right-wing) papers. This dynamic suggests that ideological reversals help prevent the formation of echo chambers in economics academia and expose authors to opposite ideological findings.

This last result is particularly important, and I believe it allows us to conclude that economists need not fear ideological reversals. In doing so, they can attract a new audience from the other side of the ideological spectrum, bringing the two sides closer together. Hopefully through that, we end up with higher-quality research overall.

[HT: Marginal Revolution, last year]

Wednesday, 4 March 2026

This is not how generative AI should be used in research

I've been using ChatGPT Pro to help with drafting research papers this year, as I noted that I would do in this post from January. It has amped up my productivity a lot, allowing me to finish writing up two papers already, with a third on the way. These were papers where the analysis was already done, but it was the writing that was holding up the process. Having ChatGPT to help with the drafting seems to kickstart my writing, even though I have ended up extensively re-writing everything that ChatGPT produces. I find it a good disciplining tool as much as anything. Several colleagues have asked whether I am disclosing my generative AI use to journal editors when I submit. And I do. I have a standard 'generative AI use statement' that I include in my papers, that notes how it was used, and that I remain responsible for all of the content. You can see an example in this recent working paper.

However, not everyone is as careful with their generative AI use, or as transparent. Consider this example:

That is both infuriating and a sad indictment of the reviewing, editing, and publishing process, not least because, as on Reddit commenter noted, many authors see high-quality work rejected by journals, whereas a paper like this, with obvious flaws, has successfully been published. And it's not an isolated incident. This 2025 article by Artur Strzelecki (University of Economics in Katowice), published in the journal Learned Publishing (open access), catalogues over 1300 instances of likely unacknowledged and frankly stupid use of ChatGPT, up to September 2024.

Strzelecki's approach is to search for text strings that are almost certainly ChatGPT responses to a prompt asking it to generate text. The main example Strzelecki uses, which is in the title of the article, is "as of my last knowledge update". No human author is going to say that in a research paper. Similarly, "as an AI language model", "I don't have access to", and "certainly, here is" are highly indicative of ChatGPT use. There are circumstances where a human might use those phrases in a research paper, but it seems unlikely. Strzelecki screens out papers that mention ChatGPT, and manually checks each paper to ensure the text was not in some way legitimate, and that leaves 1362 articles.

How do these articles get published with this content intact? There are lots of stopping points where this could be caught and corrected (or prevented), but these articles have gotten through all of them. Strzelecki outlines the process. First, perhaps it is only one of the authors (and not all of them) that used ChatGPT. In which case, why didn't the other co-authors pick it up? Next, the paper is submitted to a journal, and often goes through a text review by the publisher. And then the editor or editors (including associate editors) looks at it, and decides whether it should be sent out for peer review. And then the peer reviewers (usually more than one, sometimes four or more) look at the paper in detail and provide comments. Then the editor receives the review reports and makes a decision. The paper may go through more than one round of review and editorial decision. And then, once accepted for publication, the article may be copy-edited. And at any of those stages, this text could be picked up. And yet, for over 1300 articles as of September 2024, the ChatGPT-generated text has not been picked up.

Strzelecki particularly focuses on 89 articles that have been published in journals indexed by Scopus or Web of Science, which should be the most credible journals. Of these:

...as many as 28 of them are in journals with Scopus percentile values of 90 and above. Two journals have a 99th percentile, indicating that they are the top journals in their field...

In total, 64 articles were found in journals considered to be in Q1, top quartile, recognized as the group of the best journals in their respective fields. Twenty-five articles are in the percentile range between 50 and 75, indicating that the journals in which these articles are found belong to Q2.

So, this phenomenon is not limited to low-ranked 'predatory' journals. In fact, looking at the list, there are several journals published by MDPI and Frontiers (for more on those publishers, see here). However, there are a whole lot published by Elsevier and Springer, publishers that we should expect much better of. Although, those are also publishers that publish a lot of journals, and a lot of articles, so perhaps that accounts for their higher numbers within the 89 articles that Strzelecki focuses on. Fortunately, I don't see any reputable journals in economics in the list, but I could be wrong.

Anyway, the takeaway is not so much that generative AI use is widespread in the write-up of research. It is that authors are using generative AI, not being transparent in their use of it, and that the quality control system by journals, even high-ranking journals, is terrible. Strzelecki makes a good point in the conclusion of his article that 89 out of over 2.5 million articles indexed in Scopus is only 0.000035% of the total indexed articles. However, this analysis is only picking up the really, really obvious cases. There will be far more use of generative AI that has not been adequately checked or acknowledged by authors, and not picked up in quality control.

I'm not against using generative AI in the write-up of research. Obviously, because I am doing the same thing. What needs to happen is that researchers need to be transparent and honest when they use generative AI, so that editors, reviewers, and the readers of research can see how it was used. That way, the users of research can evaluate for themselves whether they should believe, discount, or discard research depending on the ways and the extent of generative AI use. Without transparency, that important evaluation step is lost.

[HT: Artur Strzelecki]

Read more:

Saturday, 24 January 2026

The long persistence of retracted 'zombie' papers

When a paper is retracted by a journal, that understandably tends to negatively impact perceptions of the researcher and the quality of their research (see here). However, these 'zombie' papers can maintain an undead existence for some time, continuing to be cited and used, sometimes uncritically, because retractions take time and because publishers are not good at highlighting when an article has been retracted. They may even continue to accrue further citations even after being retracted. In terms of understanding the effect of retractions on the research system, a key question is: how long does it take for a paper to be retracted?

That is essentially the question that this new article by Marc Joëts (University of Lille) and Valérie Mignon (University of Paris Nanterre), published in the journal Research Policy (open access), addresses. Joëts and Mignon draw on a sample of 25,480 retracted research articles over the period from 1923-2023 (taken from the Retraction Watch database), and look at the factors associated with the time to retraction (that is, the time between first publication and when the article is retracted). First, they find that:

...the average time to retraction is approximately 1045 days (nearly 3 years), but there is significant variability, with a standard deviation of 1225 days... However, some extreme cases take much longer, with the longest retraction occurring 81 years after publication.

Joëts and Mignon use several different forms of survival model to evaluate the relationship between the characteristics of an article and the time to retraction. In this analysis, they find that:

Papers in biomedical and life sciences are generally retracted faster than those in social sciences and humanities, and articles published by predatory publishers are withdrawn more promptly than those from reputable journals. Collaboration intensity and type of misconduct also emerge as significant predictors of retraction delays.

The result for predatory journals seems somewhat surprising. However, Joëts and Mignon suggest that:

...predatory journals often publish papers with evident deficiencies that are more easily detectable by external parties, such as watchdog organizations or institutions, leading to quicker retractions when misconduct is identified. Additionally, the lack of formal editorial procedures in predatory journals may result in a less structured and faster retraction process...

Of course, a faster time to retraction doesn't make predatory journals good. It simply makes them less bad, since they almost certainly are a large source of low-quality research that deserves retraction (Joëts and Mignon don't report the proportion of retractions that come from predatory journals).

In terms of collaboration intensity, articles with more co-authors take longer to retract, presumably because more people are involved in the retraction process, or because disputes over who is to blame may take some time to resolve. For types of misconduct, retractions due to 'data issues' take the longest to occur, while those for 'peer review errors' and 'referencing problems' take the least. That likely reflects that it takes some time for data analyses to be replicated and for problems to surface, whereas problems with referencing are more likely to be readily apparent from a simple reading of the article.

Joëts and Mignon also do a lot of modelling of different editorial policy changes and their effects on the distribution of times to retraction, but I don't think we can read too much into that part of the article, as the results are mostly driven by the assumptions on how the policies affect retractions. Nevertheless, this paper provides some insight into why zombie papers can keep shambling through the literature: retractions are slow and the time to retraction depends on discipline, publisher type, collaboration, and the kind of misconduct involved.

Read more:

Wednesday, 7 January 2026

Lessons from Joshua Gans on AI for economics research

I've been increasingly using generative AI (specifically, ChatGPT) to assist with research. I've been quite cautious though, worrying a lot about the quality of AI output, although my worries reduced substantially once ChatGPT started linking to its sources. However, I know many other researchers are using generative AI far more extensively and directly in their research than I am. My approach continues to be to use ChatGPT as an enthusiastic, but not fully polished, research assistant. Given my experience so far, I was interested to read the reflections of Joshua Gans on his year of using generative AI for economics research. His approach was:

I had lots of ideas for papers that I hadn’t developed, so I decided to spend the year working my way down the list. I would also add new ideas as they came to me. My proposed workflow was all about speed. Get papers done and out the door as quickly as possible, where a paper would only be released if I decided I was “satisfied” with the output. So it cut any peer reviews or discussions out during the process of generating research quickly, but I would send those papers to journals for validation. If I produced a paper that I didn’t think could be published (or shouldn’t be), then I would discard it. There were many such papers.

Like Gans, I have a lot of research ideas, and not enough time to pursue all of them. Many of my ideas would go nowhere, even if I did have time to pursue them. But for some others, I have later read research papers that have done something I had thought of earlier but not had time to do myself. There are therefore a lot of missed opportunities, because it isn't possible to perfectly identify the good ideas in advance - you need to try them out before you realise that they are uninteresting or dead ends. Being able to try out more ideas seems like a good thing.

The opportunity cost of spending time pursuing one research idea is not pursuing other ideas. If generative AI allows us to pursue research ideas in less time, then it lowers the opportunity cost of pursuing those ideas. However, as Gans notes:

When you lower the cost of doing something, you do more of it. Normally, the decision whether to continue or abandon a project gives rise to some introspection (or rationalisation) of whether continuing is worthwhile relative to the costs. When the going gets tough, you drop ideas that don’t look as great.

The issue with an AI-first approach is that its benefit, reducing the toughness of going, is also its weak point; you don’t face those decision points of continuing/abandoning as often. That means that you are more likely to end up completing a project. But this lack of decision points means that you end up pursuing more lower-quality ideas to fruition than you would otherwise.

In Gans's experience, AI increases research output, but it also weakens the stopping rule that would otherwise kill bad ideas early, by decreasing the marginal cost of continuing the research. When the marginal cost of continuing is lower, we spend more time on each bad idea before discarding it. And if we spend too long on bad ideas, the opportunity cost of pursuing an idea increases - time spent working on bad ideas is time not spent pursuing ideas that turn out to be better. As a result, the average quality of our research may decrease. That risk needs more careful consideration.

Although he doesn't note the potential increasing opportunity cost, Gans's conclusion seems to point in that direction:

My point is that the experiment — can we do research at high speed without much human input — was a failure. And it wasn’t just a failure because LLMs aren’t yet good enough. I think that even if LLMs improve greatly, the human taste or judgment in research is still incredibly important, and I saw nothing over the course of the year to suggest that LLMs were able to encroach on that advantage. They could be of great help and certainly make research a ton more fun, but there is something in the judgment that comes from research experience, the judgment of my peers and the importance of letting research gestate that seems more immutable to me than ever.

Generative AI has the potential to increase the quality, and the quantity, of research. Gans seems to have seen it in his work, and I've seen it already in my own work too. In fact, my experience so far has been that careful use of generative AI (for example, checking for literature gaps, or exploring econometric methods and robustness checks) has reduced the time wasted on fruitless research that would have gone nowhere. However, it is possible to use too much generative AI in research, just as it is possible to use too little. There is a middle ground, and Gans seems to be finding it from one direction (starting from over-using generative AI), and maybe I am finding it from the other (starting from under-using generative AI). The important thing seems to be ensuring that a human is kept in the loop (as Ethan Mollick noted in his book Co-Intelligence, which I reviewed here). Specifically, we can use generative AI for its strengths (in testing our initial ideas, mapping the literature, exploring alternative methods, or stress-testing assumptions). And we can keep the human in the loop by pausing the research at more key points to consider where it has gotten to and check our intuition, as well as continuing to seek peer review of the draft end-product.

So, Gans might be holding back on the generative AI this year, but I'll be further expanding my use. Starting with something related to this post: writing up some research on using AI tutors in teaching first-year economics, which is research I presented in a brown-bag seminar at Waikato last month (and I will have more on that in a future post).

[HT: Marginal Revolution]

Sunday, 28 December 2025

Specialisation trends in economics, and who’s citing economics

Do research findings in one discipline affect research in other disciplines? This seems like an important question, and one to which we hope the answer is yes. And yet, for a long time disciplines seemed to be becoming more insular, only referencing research within their own narrow fields and sub-fields. More recently, it seems to me that the breadth of citation across disciplines has increased greatly. Call it the 'Google Scholar effect', if you like. Since research has become much more searchable, relevant past research has become easier to find and to cite. If there is a 'Google Scholar effect', we'd expect to see a big change in citation patterns across disciplines starting from 2004, when Google Scholar was introduced.

So, I was interested to read this 2024 article by Sebastian Galiani (University of Maryland), Ramiro Gálvez (Universidad Torcuato Di Tella) and Ian Nachman (Brown University), published in the journal Economic Inquiry (ungated earlier version here). They conduct a citation analysis to look at specialisation trends in economics. Specifically:

...we look for patterns that indicate specialization within a field of economics research, such as a narrowing of the topics it covers in a way that other fields do not, and patterns suggesting a decrease in citations from outside the field...

Their dataset includes:

...all articles published from 1970 up to and including 2016 in the so‐called economics Top 5 journals (The American Economic Review, Journal of Political Economy, The Quarterly Journal of Economics, Econometrica, and The Review of Economic Studies) and for all articles published from 1970 up to and including 2016 in a set of non‐Top 5 general research journals (The Economic Journal, International Economic Review, Economic Inquiry, and The Review of Economics and Statistics)...

Galiani et al. categorise each article into one of four 'fields' of economic research:

...(1) applied, (2) applied theory, (3) econometric methods, and (4) theory... The criteria used to assign a paper to a field were as follows: (1) Applied articles have an empirical or applied motivation. They rely on the use of econometric or statistical methods as a basis for analyzing empirical data, although they may deal with simple models that provide a theoretical framework for the analysis... (2) Applied theory articles construct a theoretical model to explain a fact, with the empirical analysis serving as a supplementary aspect rather than the primary focus of the paper. These articles typically have limited utilization of econometric or statistical analyses, though they may employ simulations (even with empirical data) or refine other techniques to test the implications of the models... (3) Econometric methods articles develop econometric or statistical methodologies... (4) Theory articles lack an empirical section; typically, they approach a topic through modeling and extensive use of formal mathematics and logic. These articles may incorporate a numerical example or a simple model calibration using theoretical data to illustrate the proposed model or to analyze its comparative statics.

The process of assigning articles to a field involves machine learning, which I'm not going to get into here. They also use natural language processing to detect 'latent topics' based on the content of each article, where each article is assigned to exactly ten topics. They use that approach, rather than using keywords or JEL codes because the topics that they identify use similar methods, making them more similar than articles that may have the same keywords but approach the research question in very different ways.

Galiani et al. then look at the long-term trends in the topics, content, and citations (both within and outside of economics) of their large sample of 24,273 articles. The results make for interesting reading, not just in terms of the trends, but in terms of the differences between the different fields of economics:

...theory and econometric methods have shown a narrowing focus on specific research topics since the 1990s, indicating a tendency toward specialization. Theory articles have experienced a significant increase in topics related to formal mathematical proofs and game theory, while econometric methods articles have shown a pronounced rise in topics related to computational statistics, estimators' asymptotic properties, and estimators' bounds. In contrast to applied papers, these fields do not exhibit rising trends in extramural citations (i.e., citations from other disciplines) and in citations from other fields within economics research (in the case of econometric methods, it shows a declining trend). These patterns also indicate a higher degree of specialization.

Applied papers have expanded their coverage to include diverse topics. Applied articles have seen a pronounced rise in topics related to impact analysis, causal analysis, and experimental economics. Over time, applied articles began to receive a higher proportion of citations from external fields, especially from disciplines such as medicine, psychology, law, and to a certain extent, education... Overall, these patterns indicate that applied papers are becoming more multidisciplinary. The case of applied theory articles is less conclusive. While they cover a broader range of topics (similar to applied papers), there has been no significant increase in extramural citations or citations from other fields of economics research (as observed with theory articles).

The observation that applied papers have received an increasing proportion of citations from outside economics over time is particularly interesting. Applied papers are more easily cited outside of economics because they often address research questions that other disciplines are also interested in. Galiani et al. identify that applied articles in economics are often cited in the education, medicine, and psychology fields, more so than articles from other fields of economics. This pattern of extramural citations does seem to become more apparent from the 1990s onwards, which would seem to suggest that Google Scholar (which was introduced in 2004) is not implicated, despite my expectation of a 'Google Scholar effect'.

That economic theory articles are not well cited outside of economics, and even outside of economic theory articles, will come as little surprise. Economic theory articles tend not to cover topics that are of interest outside of economics, and when they do, they are far too technical for those outside of economics to appreciate.

On the other hand, econometric methods should have a greater impact outside of the field, and it is disappointing that those papers do not seem to have an external impact. There are advances in econometrics, particularly in causal inference, that other disciplines should take advantage of. In contrast, economics increasingly involves methods drawn from applied statistics and computer science (including the machine learning and natural language processing that Galiani et al. used in this research).

Overall, the takeaway from this research is that economics is simultaneously becoming more specialised (within economic theory and econometric methods) and having a greater impact outside of the field (for applied papers). It would be interesting to see whether similar results would come from a similar analysis of articles in political science or psychology. If political science or psychology show the same pattern, that would tell us something important about how methods and results diffuse, or how they don't.

[HT: Marginal Revolution, last year]

Thursday, 11 December 2025

Do men and women pitch science proposals differently, and does it matter for funding outcomes?

If male academics and female academics write academic papers and grant proposals differently, does that lead to different outcomes by gender? Past studies have worried about whether grant funding decisions are affected by gender bias (see here, for example), and differences in writing style may contribute to that. However, the article I discussed in this post from earlier this year concluded that there was little evidence of bias in grant funding, at least since 2000 in the US.

Nevertheless, I thought it would be interesting to read this 2020 article by Julian Kolev (Southern Methodist University), Yuly Fuentes-Medel, and Fiona Murray (both MIT), published in the journal AEA Papers and Proceedings (ungated here), because it not only looks at the grant funding decisions, but also at writing style. Kolev et al. focus on grant applications submitted to the Bill and Melinda Gates Foundation and the National Institutes of Health (NIH) over the period from 2008 to 2016. The sample includes 6931 Gates Foundation applications and 12,589 NIH applications.

Kolev et al. first subject the applications to textual analysis to the abstract of each application, evaluating the positivity of the text (and the extent to which the word "novel" is used), the readability (using the Flesch reading ease score), concreteness of the language (as opposed to abstractness), and three measures of how narrow or broad the abstract is. In this textual analysis, they find that:

...female applicants are less likely to present their research using positive vocabulary, they are more likely to write with high readability, and they prefer concrete language. Moving to our final three measures, we find an interesting dichotomy: even as female applicants use fewer broad words and more narrow words in their abstracts, we find that their research is characterized by lower MeSH concentrations, meaning that they cover a wider range of medical subjects in their work, at least within the NIH sample. Effect sizes are relatively small: the impact of gender ranges from approximately 0.04 to 0.08 standard deviations for our significant effects.

So, there are small but statistically significant differences in writing style between male academics and female academics in these funding proposal abstracts. Does that translate into differences in outcome? Kolev et al. test for whether the measures of writing style correlate with funding outcome, while controlling for:

...calendar time and application topic fixed effects, controls for total word count and the count of relevant words for dictionary-based metrics, and applicant publication history and gender.

In this analysis, Kolev et al. find that:

For Gates applicants, high levels of concreteness tend to improve the odds of funding; by contrast, at the NIH, we find a strong positive impact for MeSH concentration and marginal effects for both broad and narrow words.

So, the evidence is weak that writing style matters, or that writing style differences between the genders affect the success of funding applications. However, not so fast. There is a key problem with this analysis. If you read the quote above about the control variables in this second analysis, you may note that they control for gender. That might sound sensible, but if you're wanting to evaluate whether writing style differences between the genders affect funding outcomes, you don't want to control for both writing style and gender. What Kolev et al. have actually tested is whether writing style differences within each gender affect funding application outcomes, finding that they don't. As an example, their analysis doesn't answer the question of whether readability differences between male and female academics affect funding outcomes, it answers the question of whether readability differences matter overall (which it appears they don't), controlling for the average difference in funding outcomes between men and women. Those are quite different questions.

In other words, we probably want to know whether style mediates the effect of gender on funding outcomes, but this analysis doesn't do that. Instead, they should either run the analysis with gender as the main explanatory variable, then add the style variables and see if the coefficient on the gender variable shrinks, or run the analysis with interactions between gender and the style variables.

The results are, on the one hand, surprising. Quality of writing should matter. Stylistic differences should matter less. However, the quality of the proposed research should matter even more than the quality or style of writing. And this study wasn't even evaluating the quality or style of all of the writing (or the quality of the proposed research), only the style of writing in the abstract for the proposal. That, along with the issue with the second analysis above, make this paper of limited use for understanding whether there is a gender difference in funding outcomes (and, if there is, whether writing style differences contribute to the difference). The difference in writing style is an interesting result in itself, but we need to know more.

Read more:

Sunday, 16 November 2025

Andrew Leigh on big data vs. randomised controlled trials

'Big data' has become the catchcry of many data scientists and researchers in recent years. It's also become increasingly used in economics. However, by itself the analysis of big data doesn't provide anything but big data correlations. Even when big datasets are available, there is still a place for randomised controlled trials (RCTs). That is the essence of this new article by Andrew Leigh (Parliament of Australia), published in the journal Australian Economic Review (sorry, I don't see an ungated version online).

It should come as no surprise that Leigh is pro-RCT. After all, he is the author of the book Randomistas (which I reviewed here), which was essentially a tribute to RCTs. Leigh clearly sees the rise of big data, and its increasing use as a substitute for RCTs, as a threat to good research. In the article, he takes great pains to point out instances where big data draws the wrong conclusions, compared with RCTs on the same topic. For example:

Randomised trials have demonstrated a strongly beneficial effect of statins on reducing cardiovascular mortality. Yet when they analysed a database covering the entire Danish population, researchers found that the chance of death from cardiovascular causes was one‐quarter higher among those who took statins than among those who did not. The explanation is straightforward: people who were prescribed statins were at elevated risk of having a heart attack. Yet even when researchers made statistical adjustments, using all the variables available in the database, they were unable to reproduce the well‐known finding that statins have a beneficial effect on cardiovascular mortality.

Analysis of the Danish database also suggested that the relative risk of cancer was 15% lower among patients who took statins, an effect that remained statistically significant even after controlling for other observed factors about the patients. Yet this result is at odds with the evidence from randomised trials. A meta‐analysis of randomised trials, covering more than 10,000 cases of cancer, found no effects of statins on the incidence of cancer, nor on deaths from cancer...

The observational data was doubly wrong. Observational data failed to replicate the well‐known finding that statins improve heart health. And observational data wrongly suggested that statins reduce the risk of cancer. Randomised trials, which were not biased by selection effects, provided the correct answer. 

That is only one example of many in the article. However, while Leigh is pro-RCT, he is not anti-big-data. He notes that:

Large data sets are a valuable complement to randomised trials. But big data is not a substitute for randomisation.

If we take anything away from Leigh's article, it should be that point. Big data is incredibly useful. However, it must be analysed using the tools of causal inference (of which randomised controlled trials are just one example) if we want to move beyond finding correlations. The problem with big data is compounded by a focus on statistical significance (as Ziliak and McCloskey noted in their book The Cult of Statistical Significance, which I reviewed here). Big datasets will find statistically significant correlations even when the size of the relationship is very small. That is an asset when causal methods are applied, but is very much a liability when big data are analysed without consideration of causality. RCTs are one way of disciplining our research approach in order to ensure that the effects we estimate are causal, and as Leigh notes:

While correlations in large data sets do not necessarily indicate causation, administrative data can be enormously helpful in ensuring the precision of estimates from randomised trials.

The article finishes with high-level strategies that policy makers and practitioners can use to ensure that RCTs are embedded within the analysis of public policy:

I advocate five approaches. Encourage curiosity in yourself and those you lead. Seek simple trials, especially at the outset. Ensure experiments are ethically grounded. Foster institutions that push people towards more rigorous evaluation. Collaborate internationally to share best practice and identify evidence gaps.

Those all sound like good approaches. I would add a sixth: Employ analysts with a thorough grounding in causal inference methods generally, if not RCTs specifically. We need more policy analysis that establishes causal evidence of impact.

Wednesday, 25 June 2025

Generative AI is mastering 'metrics

The capabilities of generative AI continue to grow. In the latest example, some enterprising economists have developed an agentic AI that can complete tasks using econometrics (the economist's statistical toolset) - 'mastering 'metrics', as Angrist and Pischke would say (see my review of their excellent econometrics text). The agentic AI approach is outlined in this new working paper by Qiang Chen (Shandong University) and co-authors. As they explain:

We propose and implement a zero-shot learning framework, called Econometrics AI Agent, that enables AI agents to acquire domain knowledge without costly LLM fine-tuning. The framework’s core component is an econometrics “tool library” implementing popular econometric methods, including IV-2SLS, DID, and RDD.

...we augment each econometric tool with detailed “prompts”—comprehensive method descriptions that specify inputs, hyperparameters, and outputs. These prompts are provided alongside corresponding Python implementations, creating a standardized interface between the econometric methods and the AI agent. This design allows the LLM to leverage both its general econometric knowledge and the specifically crafted prompts and tools, enabling it to conduct complex econometric analyses through multi-round interactions with users. The resulting framework empowers Econometrics AI Agent to independently handle applied econometric tasks, delivering comprehensive results that include parameter estimation, inference, and analytical discussions.

The Econometrics AI Agent that Chen et al. created is available here. That site also includes detailed installation instructions, and a helpful demonstration video. Coming back to the paper, Chen et al. show the capabilities of the model by testing it on several real-world problems:

We evaluate the Econometrics AI Agent through two sets of inquiries. The first comprises 18 exercises from the coursework assignments of a doctoral-level course titled “Applied Econometrics” at the University of Hong Kong, with Python-generated standard solutions. These exercises cover OLS & PanelOLS regression, propensity score matching, IV-2SLS regression, Difference-in-Differences (DID) analysis, and Regression Discontinuity Design. The second set consists of test datasets from randomly selected seminal articles in reputable journals, primarily accompanied by Stata-based replication packages.

They compare the performance of their agent, in terms of creating code that works correctly, and in terms of the resulting estimated coefficient of interest, in comparison to three alternatives:

...(i) direct LLM generation in Python code, (ii) direct LLM generation in Stata code, and (iii) baseline general-purpose AI agents without specialized econometric tools and domain knowledge.

So, this is an approach to replication, which is important (for example, see here), and is a point that I will return to later. Overall, in comparison to LLMs and general-purpose AI agents, the Econometrics AI Agent performs much better. In terms of the econometrics coursework assignments, Chen et al. find that:

The Econometrics AI Agent demonstrates superior performance with a 95% directional replication rate and average coefficient value errors below 3%. In contrast, both GPT-generated Python and Stata control groups show incorrect directions in over half of test cases. While the general AI Agent achieves a 78% directional replication rate, its coefficient values frequently deviate significantly from true values.

The rate of 'perfect replication' (which Chen et al. defined as the errors in the coefficient, standard error, and p-value all within 1% of the 'true' value) was 51.85 percent for the Econometrics AI Agent, but less than 30 percent for the other models. Turning to the published paper replications, Chen et al. find that the rate of 'perfect replication' was just 27.41 percent for the Econometrics AI Agent, but that was still far higher than the other models, which all had rates under 18 percent. In relation to those results, Chen et al. note that:

...the Econometrics AI Agent does show room for improvement. For example, its performance declines for complex econometric methods like DID and RDD compared to simpler approaches such as OLS and IV-2SLS. Similarly, results slightly deteriorate when moving from straightforward coursework problems to more sophisticated paper replication tasks. However, these limitations can be addressed through the AI agent’s domain knowledge architecture—specifically by developing customized tools and enhancing prompt instructions to better support complex algorithms and detailed requirements.

Indeed, it is the modular nature of the agent's architecture that may be its key advantage, allowing modules relevant to each econometric task to be added or updated over time. On this point, Chen et al. note that:

Unlike the costly and often infeasible process of fine-tuning an LLM to keep pace with rapid academic advances in developing new techniques, our agent can be updated simply by adding new tool functions and descriptions to the prompt library. This modularity allows the agent’s knowledge base to expand alongside the field’s developments, making the integration of recently published procedures as straightforward as adding new modules.

So, we can expect that the agent's capabilities, and its accuracy, will only improve over time. We may not be very far away from a time when empirical economists spend far less of their time on coding esoteric econometric code in order to extract meaningful results. What will we do with our free time? Maybe we'll be able to turn our attention to a greater variety of research questions.

There is a further positive aspect to these results. The replication crisis is real in many disciplines, including economics. Having AI agents that can automate the steps required to generate econometric results will decrease the time cost of completing replications. That means that we may expect more paper replications in the future (at least, more of the type of replication that Miguel and Christensen call a 'verification'). This move will certainly be a positive, leading to improvements in the quality of research in the future.

[HT: Marginal Revolution]

Read more:

Tuesday, 3 June 2025

How prevalent is large language model use in the write-up of economics research?

Back in January, I poked fun at a paper on students' acceptance of ChatGPT that had parts that were clearly written by generative AI. And I recently read a working paper that was great (and I'll blog on it sometime soon), up until the Conclusion section, which was clearly written by generative AI. But how common is this? Academics worry about how often students are using AI to write assignments or essays, but how often are we doing so?

That is essentially the question addressed in this new article by Maryam Feyzollahi and Nima Rafizadeh (both University of Massachusetts Amherst), published in the journal Economics Letters (sorry, I don't see an ungated version online). Feyzollahi and Rafizadeh investigate the top 25 economics journals over the period from 2001 to 2024, and basically look for word choices that are characteristic of large language models (LLMs). As they explain:

We construct two equally-sized word sets for our analysis: treatment words that are characteristically associated with LLM-assisted writing, and control words that represent traditional academic writing patterns... The treatment words are selected based on two criteria. First, we analyze a large corpus of confirmed LLM-generated academic text to identify words that appear with systematically higher frequency compared to human writing. Second, we cross-reference our selections with existing literature on language model patterns... to validate our choices... The control words are selected based on two criteria. First, these words represent established economic and econometric concepts that have maintained consistent usage patterns in academic writing over our sample period. Second, they are semantically unrelated to our treatment words, ensuring that any potential changes in treatment word frequencies do not spillover to or correlate with control word usage through meaning associations.

I know that you're wondering about the word list, and it is provided in Table 2 from the paper:

That list seems more nuanced (I swear that ChatGPT did not write this sentence!) than the word choices that have previously highlighted as signals of LLM use, like "rich tapestry", "realm", or "mosaic". However, some old favourites like "delve" and "foster" do appear in the list, so clearly LLMs haven't completely evolved to avoid their characteristic phrases.

Feyzollahi and Rafizadeh compare the relative frequency of the treatment and control words. Their results are quite well illustrated in Figure 1 (a) from the paper, which shows how the use of the words "intricate" (a treatment word) and "coefficient" (a control word) have changed over time:

Maybe starting from 2023, research became more intricate, or there were more intricacies in the findings of research? Or more likely, LLMs suddenly started to play an increasing role in the write-up of research. Generalising from that comparison of just two words, Feyzollahi and Rafizadeh use a simple regression model and find:

...compelling evidence of increasing LLM adoption over time. When considering both post-treatment years... the analysis documents a significant increase of 4.76 percentage points in the frequency of LLM-associated terms, with the effect maintaining remarkable stability across all specifications.

And then when comparing 2023 and 2024, Feyzollahi and Rafizadeh find:

 ...an accelerating pattern of LLM adoption. The initial impact in 2023... shows an increase of 2.85 percentage points, while the effect more than doubles to 6.67 percentage points in 2024...

So, LLM use is small, but growing quickly in academic economics. And there are many reasons to believe that these results understate the true use of LLMs in the write-up of research in economics. It takes some time for research to get published, so there will likely be far more papers in the 'publication pipeline' that have used LLMs. Authors can re-write text that was drafted by an LLM in order to mask the LLM's contribution. LLMs may be getting better at writing in an 'academic style' that avoids the use of phrases that signal the use of an LLM (no more delving!).

Overall though, it is clear that LLMs are increasingly being used to write up research for publication. A relevant question to ask is: does LLM use reduce the quality of the underlying research? Personally, when I read a paper where an LLM has clearly been used in the writing, I chuckle to myself. However, I haven't as yet had cause to disbelieve the underlying results of the research. However, my reaction doesn't necessarily reflect the views of academics in general. Regardless, when an LLM is used that use should be transparently disclosed by the authors (indeed, John List reported results of a quick survey of his followers on LinkedIn recently, where only 14 percent of them suggested that the use of Claude in a research paper should not have been disclosed).

We should not be surprised that researchers are using LLMs. LLMs can increase our productivity by helping us to write up our research more quickly. In that sense, we face similar incentives to students, who are trying to complete their assignments and essays more quickly. Both students and researchers, though, should at the very least acknowledge their use of these tools.

Read more:

Monday, 19 May 2025

The impact of Nobel Prizes and MacArthur Fellowships on the winners

What happens to the research productivity of winners of top research awards? On the one hand, a research award like a top fellowship or a Nobel Prize might increase a researcher's impact, as other researchers follow the path they have laid down. On the other hand, maybe there is some 'mean reversion', where a previously high-flying researcher simply returns to a less stellar research trajectory (which would look like a decrease in productivity). Or, perhaps a top research award grants a researcher the freedom to explore new, higher-risk areas of research, which could lead to much higher, or much lower, productivity overall?

The question of what happens to researchers after winning a top research award is addressed in this 2023 article by Andrew Nepomuceno, Hilary Bayer, and John Ioannidis (all Stanford University), published in the journal Royal Society Open Science (open access). They looked at the pre- and post-award citation counts for all 72 winners of the Nobel Prize in chemistry, medicine, or physics over the period from 2004 to 2013, and 119 of the 238 McArthur Fellows (only including those in STEM or social science fields) over the same years. Specifically, they compared publications published in two periods of three years: (1) the two years before the award and the year of the award; and (2) the three years after that. They counted citations for the pre-award period up to 2015, and the post-award period up to 2019 (so that both the pre-award and post-award periods had the same number of observed years of citations).

In their main results, Nepomuceno et al. report that:

Nobel Laureates and MacArthur Fellows received fewer citations for post-award work than for pre-award work... The difference was driven predominantly by Nobel Laureates while there was little difference, on average, for pre- versus post-award citation impact for MacArthur Fellows. The median decrease was 80.5 citations among Nobel Laureates and 2 among MacArthur Fellows. For Nobel Laureates, the decrease reached statistical significance (Wilcoxon signed-rank test p = 0.004), whereas for MacArthur Fellows the decrease was not statistically significant (Wilcoxon signed-rank test p = 0.857)...

Post-award citation impact was lower than the pre-award citation impact for 45 of 72 (62.5%) Nobel Laureates and for 63 of the 119 (52.9%) MacArthur Fellows.

Both Nobel Laureates and MacArthur Fellows suffered a reduction in the citation count per-publication after receiving their award, but for different reasons. The Nobel Laureates published the same number of papers in the period after the award as they did before the award. But their lower citations mean that the citation count per-publication was lower. In contrast, the MacArthur Fellows published more papers after the award than they did before the award, but with no change in total citations (again, meaning that the citation count per-publication was lower).

One major difference between the two groups is age - Nobel Laureates are much older than MacArthur Fellows. So, Nepomuceno et al. conducted further analyses stratified by age (in three groups: under 42 years old, 42-57 years old, and over 57 years old), and found that:

...the declining citations pattern was seen only for researchers who were 42 or older at the time of the award, while an opposite pattern was seen for early career researchers who were given an award (especially MacArthur award) at an age of 41 or younger.

However, looking at Table 2 in the paper, it is clear that the negative impact on total citations is largest for the youngest Nobel Prize winners (those aged under 42 years), but is negative for all three age groups. In contrast, there is a positive impact on citations for the youngest McArthur Fellows, and a negative impact for McArthur Fellows aged over 42 years.

Overall, Nepomuceno et al. conclude that:

Although the MacArthur Fellowship and Nobel Prize selection committees share a stated goal of assisting winners in realizing their potential more fully, in terms of citation counts neither the MacArthur Fellowship nor the Nobel Prize heralded increased research impact for the subsequent work and for Nobel Laureates there was even a significant decline.

It is tempting, then, to conclude that these awards are not a good idea. I'm not so sure. I think the research highlights different impacts of the two awards, and I think we learn something potentially important from this. Nobel Laureates may tend to rest on their laurels (pun intended), or may suffer from mean reversion. Or, perhaps they use the profile accorded by their new status as Nobel Laureates to try and have greater policy or political influence, with an opportunity cost of lower research influence. That suggests that it is better to award Nobel Prizes to end-career academics, lest younger academics be diverted from important and path-breaking research. The recent trend in awarding Nobel Prizes to younger recipients (definitely noticeable in economics) may therefore have a negative unintended consequence. In contrast, because there is a positive citation impact for young recipients, the MacArthur Fellowships should be targeted in greater proportion to younger researchers. There is less to be gained from awarding those Fellowships to end-career academics.

To be fair, that is more-or-less how those two awards have historically been allocated: Nobel Prizes to end-career academics, and MacArthur Fellowships ('genius grants') to young stars. This research suggests that might be an important practice to continue.

[HT: Marginal Revolution, back in 2023]

Wednesday, 14 May 2025

An interesting paper about the first 50 years of Nobel Prize winners in economics

The first Nobel Prize in economics was awarded in 1969, to Ragnar Frisch and Jan Tinbergen. The fiftieth prize was awarded in 2018, to William Nordhaus and Paul Romer. In total up to that point, there had been 91 Nobel laureates in economics. This 2019 article by Allen Sanderson (University of Chicago) and John Siegfried (Vanderbilt University), published in the journal The American Economist (ungated version here), reviews those first fifty awards. In addition to summarising the topics, Sanderson and Siegfried collate a lot of interesting factoids, starting with the origins of the award:

The 1895 will of Swedish scientist Alfred Nobel specified that his estate be used to create annual awards in five categories—physics, chemistry, physiology or medicine, literature, and peace—to recognize individuals whose contributions have conferred “the greatest benefit on mankind.” Nobel Prizes in these five fields were first awarded in 1901...

In 1968, Sweden’s central bank, to celebrate its 300th anniversary and also to champion its independence from the Swedish government and tout the scientific nature of its work, made a donation to the Nobel Foundation to establish a sixth Prize, the Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel...

Sanderson and Siegfried summarise the backgrounds of the laureates, which are mostly unsurprising (a lot of top universities, and a lot of economics and mathematics), although:

Some notable surprises include Middle Tennessee State Teachers College... (James Buchanan) and South Dakota State University (T. W. Schultz).

And apparently, Eugene Fama's undergraduate degree was in romance languages! Sanderson and Siegfried also note that:

Economics joins literature and peace as the Nobel fields that have generated the most controversy. First, as a well-known quip has it, “economics is the only field in which two people can share a Nobel Prize for saying opposing things.” The 1972 Prizes awarded to Myrdal and Hayek spring to mind, as would the 2013 awards to Fama and Shiller...

I have often been tempted to create an assessment for my ECONS102 class to name and justify the best economist (living or dead, but eligible when living) never to have won a Nobel Prize. Sanderson and Siegfried provide their own list of economists who died before the economics Nobel Prize existed (but after Nobel Prizes were first awarded in 1901), which includes Leon Walras (who died in 1910), Vilfredo Pareto (1923), Alfred Marshall (1924), Thorstein Veblen (1929), John Bates Clark (1938), John Commons (1945), John Maynard Keynes (1946), Irving Fisher (1947), Joseph Schumpeter (1950), John von Neumann (1957), Arthur Pigou (1959), and Karl Polanyi (1964). That seems like a reasonable list to me.

Sanderson and Siegfried also provide a further list of economists who died after 1969 but never received a Nobel Prize, but could have done, which includes Frank Knight (died in 1972), Alvin Hansen (1975), Oskar Morgenstern (1977), Joan Robinson (1983), Piero Sraffa (1983), Fischer Black (1995), Amos Tversky (1996), Zvi Griliches (1999), Sherwin Rosen (2001), John Muth (2005), J.K. Galbraith (2006), Anna Schwartz (2012), and Martin Shubik (2018). Sanderson and Siegfried then add:

To this list, one could certainly add more of their contemporaries, for example (in alphabetical order), Anthony Atkinson (2017), William Baumol (2017), Harold Demsetz (2019), Evsey Domar (1997), Rudiger Dornbusch (2002), Henry Roy Forbes Harrod (1978), Harold Hotelling (1973), Nicholas Kaldor (1986), Jacob Mincer (2006), Hyman Minsky (1996), and Ludwig von Mises (1973), among many others.

I would agree with many of those from both lists, especially Robinson, Baumol, Demsetz, and Hotelling. It is worth noting that Fischer Black would almost certainly have shared the 1997 Nobel Prize with Myron Scholes (and Robert Merton), while Amos Tversky would almost certainly have shared the 2002 Nobel Prize with Daniel Kahneman (and Vernon Smith). There have also been surprising near misses in each direction, one of which was William Vickrey, who died three days after the award was announced (and therefore some months before the award ceremony). The other notable near miss was where:

Polish macroeconomist Michal Kalecki was nominated for the Nobel Prize in 1970 but died in April of that year...

Sanderson and Siegfried wisely steered clear of suggesting potential future winners (after 2018 when their sample ends). Nevertheless, the article is a great summary of the first 50 years of the Nobel Prize in economics, and well worth a read. 

Thursday, 8 May 2025

The disturbing lack of impact of comments and replications of economics research

Back in 2018, I wrote a comment on an article that I was the reviewer for. The article (open access) described four types of 'economic citizen'. My comment (open access) argued that one of the four types was not distinct from the other three, but instead captured a different dimension of economic citizenship that might apply to any of the other three types. The comment was published alongside the article, along with a reply (open access) by the original article's authors, in the journal Education Sciences. What has happened since is that, according to Google Scholar, the original article has been cited 23 times. The comment has been cited just one, by the authors' reply, which has itself never been cited.

Now it may be that my critique of the original article is not valid, and so subsequent authors citing the original article don't feel the need to cite the comment. Or, it could be that the critique is valid (I still think it is), but is being ignored for some reason. If this is a common experience for research, then it calls into question whether research is self-correcting or not. If research findings are not easily overturned or challenged, then future researchers may be wasting a lot of time and effort on research that will not advance the field, being based on questionable earlier research.

How big a problem is this? That is the question addressed in this new article by Jörg Ankel‐Peters, Nathan Fiala, and Florian Neubauer (all Leibniz Institute for Economic Research), published in the journal Economic Inquiry (open access). Ankel‐Peters et al. look at all 56 replications published as comments between 2010 and 2020 in the American Economic Review, arguably the overall top journal in economics (and certainly in the top five journals). As they explain:

For the self‐correction claim to hold, we hypothesize that a comment should lead to a strong reaction of the literature, especially for a comment raising substantive concerns about an OP. If it does not respond strongly, we argue, the prior in the literature sustains. We look at two facets of a strong response: (1) Citations of the comment relative to citations to the OP after comment publication (henceforth: citation ratio), and (2) Whether the comment affects the OP's annual citations.

Ankel‐Peters et al. don't conduct formal statistical tests, and instead rely on a descriptive analysis. Nevertheless, the results are compelling:

We find that AER comments do not affect the OP's [Original Paper's] citations and hence their influence on the literature. We observe an average citation ratio of 14%. Comments are cited on average seven times per year since their publication—compared to an average of 74 citations per year for the OP since publication of the comment. Comments are, hence, not cited much in absolute terms, and a lot less than the OP. The latter implies that most OP citations ignore the comment.

They also find very similar results when they focus on citations in the top five economics journals, and similar results when they account (subjectively) for whether the comment 'must be cited' alongside the original paper because of how substantive the concerns raised are. Ankel-Peters et al. conclude that:

We interpret this as evidence for the absence of self-correction mechanisms in economics.

Obviously, that is disappointing. Replications are seen as important to ensuring the integrity and credibility of research (for example, see here). If papers that have failed to replicate, or that have substantive problems with them, are being cited uncritically in the following literature, then economics research could easily be led down some dead-end paths. What can be done to mitigate this problem? Ankel-Peters et al. don't offer much in the way of solutions. However, the bare minimum would be that comments and replications should receive more prominence so that readers of the original research are aware of them. To this end, Ankel-Peters et al. report a very small victory:

In response to a previous version of our paper, the current AER editor has let us know that the journal has changed its policy and now, for new comments, will provide a link on the OP's website. This is a small but perhaps important first step to giving replication work in economics the attention it needs and deserves.

Indeed. Now all journals need to follow AER's lead (which isn't much of a lead - many journals already do that). A better solution would be if top journals required citation of relevant high-quality comments or replications whenever an original paper is cited. And it would be even better if top journals published more high-quality replications and comments, as well as more high-quality systematic reviews and meta-analyses. Then we might advance economics research in a more informed way.