Friday, 4 September 2026

This week in research #142

Here's what caught my eye in research over the past week:

  • Sacerdote, Staiger, and Tine (with ungated earlier version here) find that test score–optional policies harm the likelihood of admission for high-achieving applicants from disadvantaged backgrounds, meaning that the availability of test scores on an application can promote rather than hinder social mobility
  • Armona et al. (with ungated earlier version here) develop and test a model of what is newsworthy to a media outlet
  • Crossin et al. (open access if you set up a free account) find using data from a nationally representative sample that an estimated 71.1 percent of the New Zealand population think politicians should do more to keep people safe from alcohol harm, with majority support across the political spectrum
  • Smit (open access) outlines the drawbacks to remote working that may have prevented greater internal migration from cities to the periphery in the Netherlands

Thursday, 3 September 2026

Jeanna Smialek on why we should remember Maria Edgeworth

I have read several books now on the history of economic thought and, for the most part, women are conspicuously absent from those books. One obvious exception is Edith Kuiper's A Herstory of Economics (which I reviewed here), and there are a few bits in Steven Medema's The Economics Book (which I reviewed here). The latter brought both Harriet Martineau and Jane Marcet to my attention (and my two class AI tutors, Harriet in ECONS101, and Jane in ECONS102, are named after them).

A surprising omission from both books was brought to my attention by this New York Times article (ungated version here) by Jeanna Smialik: Maria Edgeworth. [*] According to Wikipedia, Maria Edgeworth was a novelist and an aunt of the much more famous 19th Century economist, Francis Ysidro Edgeworth. However, Smialek makes a strong case for the importance of Maria Edgeworth as one of the earliest voices in the developing field of economics, with her economics expressed within her fiction. However, Smialek also spends some efforts to explain why Edgeworth has been largely forgotten as an economist:

As the field professionalized, the second- and third-generation economists distanced themselves from the women and the fiction that had once helped to popularize and explain their ideas. Serious science, after all, could not possibly be for everyone. The economist Alfred Marshall referred to the women who tried to simplify economic doctrine as “parasites” in a footnote to his hugely influential textbook, “Principles of Economics,” first published in 1890 and popular throughout the 1900s...

Her literature was lost partly because of its clunkiness — its economics lessons made it harder to digest as lighter and more naturalistic stories became the style.

Her economics was lost for a different reason. Subsequent influential academics disparaged the ways that the early women of economics had presented the field. They bristled at being simplified, and they dismissed the idea that the women could have contributed something meaningful. 

Smialek also notes that Edgeworth provided editorial comments on Jane Marcet's Conversations on Political Economy, arguably the first economics textbook:

Edgeworth was as skilled an editor as she was a writer. She told Marcet that Adam Smith ought to be abridged, because he “needs it much,” whereas the utilitarian philosopher Jeremy Bentham was “absolutely impossible” to put in fewer words, because what he needed was “to be diluted.”

“The more amusing anecdote & illustration you can mix with your solid information the better,” Edgeworth wrote to Marcet. “It should be your object rather to sow seeds than to exhibit full grown plants — Your work should excite curiosity to go further.”

Clearly, there is more for us to learn about the history of economic thought from Maria Edgeworth's life and her writing. And we are in luck! Smialek has a forthcoming book about Edgeworth due to be released next month, titled The Invisible Hand of Maria Edgeworth. I'm definitely looking forward to that one, and you can expect a review of it here in due course (although, with my backlog of reading, that might not be until sometime next year!).

*****

[*] Smialek notes that Maria Edgeworth was brought to her attention by Robert Heilbroner's The Worldly Philosophers (which I reviewed here). Heilbroner's reference to Edgeworth clearly didn't make as great an impression on me, as I don't remember it!

Sunday, 30 August 2026

The impact of a subtle wording change on measurement of worries about climate change

Consider the following statement: "The effects of climate change are too far in the future to really worry me". Would you say that you: strongly agree with that statement; tend to agree; neither agree nor disagree; tend to disagree; or strongly disagree? Now consider this alternative wording of the same statement: "The effects of climate change are too far in the future to really worry about". Would your level of agreement be the same to this second statement as to the first?

There is good reason to believe that the wording of that question matters, as shown in this new article by Lieke Voorintholt (Trier University), Adriaan Soetevent, and Gerard van den Berg (both University of Groningen), published in the Journal of Economic Psychology (ungated earlier version here). They argue that the single-word difference, from "worry me" to "worry about" may shift respondents' focus from themselves individually towards broader society. Therefore, they hypothesise that more survey respondents would report being worried when asked if they worry about the effects, compared to whether the effects worry them individually. They also hypothesise that the treatment effect would be greater for older respondents, since the "about" version prompts thinking beyond the respondent's own life expectancy.

Voorintholt et al. use data from the Innovation Panel of Understanding Society, which its website describes as "a test-bed for innovative ways of collecting data and for developing new areas of research". The survey is run in Great Britain, and their analysis is based on a sample of 2791 survey respondents from Wave 16 of the survey, conducted in 2023. Half of the respondents were randomised to be asked the "me" version of the question, and the other half were asked the "about" question. Voorintholt et al. term the "me" version the control group, because that is the wording that is used in the main Understanding Society survey. So, the treatment is simply the replacement of "me" with "about" in the wording of the statement.

Their main results are summarised in Figure 1 from the article:

When asked the "about" version, more respondents say they strongly disagree with the statement, while fewer respondents say they tend to agree with the statement. These differences can be interpreted as greater worries about climate change expressed by the treatment group than the control group. This is also supported by their quantitative analysis, which shows the average response among treated respondents is 0.18 points higher (on a 1-5 scale from strongly agree to strongly disagree) than control respondents. Voorintholt et al. find some suggestive evidence in support of larger effects for older respondents. The coefficient on an interaction term between treatment and an age dummy variable for those aged 66 and over is positive and relatively large, but is statistically insignificant.

Overall though, this research is a good reminder that even a one-word change can materially affect how survey respondents answer a question, and therefore how we measure important variables. The effect of subtle changes should not be underestimated. If you want to measure the level of concern that people express about climate change, it really matters whether respondents are asked whether climate change “worries me” or whether it is something to “worry about”.

Saturday, 29 August 2026

This week in research #141

This week I attended the 65th Congress of the European Regional Science Association in Sofia, Bulgaria. This is one of my favourite conferences each year, because of the quality and variety of papers that are presented. My own presentation was on the increasing trend towards natural population decrease (more deaths than births) across the territorial authorities and local boards in New Zealand, which will become an increasingly serious policy issue in the future. Here are some of the highlights I found from the conference:

  • Ugo Fratesi (abstract 270, on page 2 of the abstract book) looked at regional resilience to positive shocks in China, flipping on its head the idea of resilience, which is usually only applied to negative shocks (and this generated a bit of discussion about whether resilience is the right word to use in that context)
  • Elisabetta Ottoz (abstract 126, on page 8 of the abstract book) investigated businesses operating in the night-time economy in Italian cities, finding that those businesses were larger and more profitable than recreational businesses that operated during the daytime or evenings
  • Peter Nijkamp gave a very detailed keynote that focused on wellbeing and introduced me to a variety of new terms drawn from disparate parts of the literature, including 'city love' (because love is more enduring than happiness), spatial comfort zones (where people are unwilling to move away, because they have everything they need within some close distance), prosilience (like resilience, but more about making the most out of a negative shock so that the area ends up better rather than simply returning to the previous state), and panarchy (adjustment behaviour is cyclical both temporally and spatially)
  • In his discussion of Nijkamp's keynote, Kingsley Haynes laid down a challenge, noting that regional scientists had done a great job of analysing aggregate data, but not so much with individual-level data (which reminded me of at least two conversations I have had over the years about how difficult or impossible it is to apply spatial models to individual-level data)
  • Zhiwu Wei (working paper here) looked at the relationship between local wealth inequality and protests in the Global South, finding that the positive inequality-protest relationship is very localised, and weakens substantially as you consider a broader definition of what is 'local'
  • Alessandra Faggian (abstract 537, on page 93 of the abstract book) investigated the relationship between fertility and productivity in Italy, using historical wheat suitability as an instrument, and found that higher local productivity leads to significantly higher fertility rates (I'm not sure there is a policy prescription here, because raising productivity seems to be just as difficult as raising fertility!)
  • Simone Piras (abstract 111, on page 130 of the abstract book) presented the results of a discrete choice experiment of Scottish residents’ willingness to move (or not) to places that differ in terms of various attributes, finding that residents would be more willing to move to places with good natural environment and digital connectivity and which are close to services and family, and that an intervention that provided information about Scotland's national policy on balancing the population had only small impacts

Aside from the conference, here's what caught my eye in research over the past week:

  • Gans (open access) shows using a theoretical model that prompt injection to manipulate AI referees of journal article submissions can actually improve the peer review process under certain conditions (I'd want to see some empirical support for this before I believed it)
  • Di Iasio and Wahba (open access) find that stronger anti-immigration attitudes significantly reduce migration inflows to EU destinations, with effects that are larger for intra-EU mobility than for migration from non-EU countries
  • Wang and Liang find that city-level crime rates declined by a statistically significant 5.7 percent following Hukou reform in China, which freed up internal migration
  • Ardito et al. (with ungated earlier version here) find that delayed retirement as a result of pension reform in Italy significantly increased sick leave due to occupational injuries, by 2.6 percent of the pre-reform mean

Friday, 28 August 2026

The unfair competition argument and the push for a 'Temu tax'

There are many arguments put forward for why trade should be restricted. In my ECONS102 class last week, as part of our topic on international trade and globalisation we covered five of the most common, each of which has a little story that goes along with it:

  1. The jobs argument: Trade with other countries will lower prices for goods and services where our country has a comparative disadvantage. This will reduce the quantity that domestic firms produce, and the number of people that they employ.
  2. The national security argument: Some goods and services are vital to national security. A conflict that disrupted trade in those goods and services would have serious negative impacts, so it may be better for our country to produce those goods and services itself, rather than relying on trade.
  3. The infant industry argument: Some industries are likely to be important for the future growth prospects of our country, but right now our firms in those industries are small and can't compete with firms from other countries. It might be best to protect those industries now, giving them a chance to grow, and take advantage of learning curve effects and economies of scale.
  4. The 'protection as a bargaining chip' argument: Our country often has to negotiate with other countries, and having trade restrictions in place now gives us something that we can offer up in order to get a better deal in those negotiations.
  5. The unfair competition argument: Firms in different countries are subject to different laws and regulations (such as consumer protection laws, labour laws, and environmental laws), giving firms from relatively lightly regulated countries a cost advantage over firms from countries that are relatively more heavily regulated.

The unfair competition argument has been playing out in New Zealand recently, in relation to local retailers having to compete with foreign producers such as Temu. As the New Zealand Herald reported back in April:

Carolyn Young, chief executive at Retail NZ, said New Zealand could look at what France and South Africa had done, as models of how a tax or levy could be applied to help local retail.

France is implementing an environmental fee on ultra-fast fashion brands, which will rise to €10 ($20) per item by 2030.

“When you think about a business in New Zealand, they pay New Zealand staffing rates. They comply with the health and safety regulations in New Zealand and their products do as well.

“They have to comply to the Fair Trading Act and the Consumer Guarantees Act. There’s always costs involved in those areas. And anything you get in from offshore, you have no idea what their labour environment is like or what they’re paying their people. The product doesn’t have to meet any health and safety standards and they’re not compliant with New Zealand regulations around fair trading and consumer guarantees.”

She said the Government should impose stronger measures to help level the playing field, such as a levy paid by shoppers.

“If you were buying from offshore, what we would want to see is that there would be a levy that would be applied to that, that would be at a level that would be some sort of equaliser between what New Zealand businesses have to do and comply with.

Notice that is almost exactly the unfair competition argument I outlined earlier. Young also points to the jobs argument as well, saying:

“Will everybody come back from shopping with them? I don’t know, but we have to try because that’s just going to make it much more difficult because as soon as you shop offshore, the money goes offshore.

“It doesn’t stay in New Zealand, doesn’t create jobs in New Zealand, doesn’t, you know, keep businesses open. And at some point, that’s going to really matter.”

She said if everyone would shop in New Zealand, it would help the economy significantly.

The 'help the economy significantly' statement needs some pushback. A tax on products that consumers buy from Temu will mean that the prices consumers pay will be higher. They would pay higher prices on goods they buy from Temu. And because the 'Temu price' that domestic retailers have to compete with would be higher, domestic retailers would face less downward pressure on their prices and consumers would therefore likely be paying a higher price when buying locally as well. When Young says that a 'Temu tax' would "help the economy significantly", what that means is that it would help domestic retailers, who could charge a higher price, and would sell more products to domestic consumers, if the cost of buying from Temu were higher because of the 'Temu tax'. The government would benefit somewhat from the additional revenue from the tax. But those gains to retailers and the government need to be balanced against the losses for domestic consumers.

Each of the arguments against free trade also has one or more counterarguments. In the case of the unfair competition argument, consumers who care about differences in labour standards, environmental protection, or consumer rights can already choose to buy from domestic retailers. To some extent, the fact that many do not suggests that they value the lower prices available from overseas retailers more highly than the additional protections that domestic regulation provides. Of course, that counterargument is weaker if consumers lack information about where or how goods are produced, or when the regulations are addressing external costs that consumers do not themselves bear.

Overall, in the absence of some market failure that the 'Temu tax' is correcting, the gains to domestic retailers (increased 'producer surplus') and the government (increased tax revenue) need to be weighed against the losses to domestic consumers (decreased 'consumer surplus'). In the standard tariff case, those losses tend to outweigh the gains, resulting in lower total welfare overall (see this post, which explains the welfare changes in detail). A 'Temu tax' might be a good idea politically, but it is unlikely to be a good idea economically. It certainly isn't the case that it would "help the economy significantly", unless you mostly ignore the costs it would impose on domestic consumers.

Thursday, 27 August 2026

The economics of pricing AFL Finals tickets

In my ECONS101 class, we have a topic that is devoted to understanding pricing and business strategies that deviate from the ideal 'marginal revenue equal to marginal cost' approach to pricing for firms with market power. In particular, I focus part of the topic on firms that price below the short-run profit-maximising price for strategic reasons.

One example is the NFL, which prices tickets for the Super Bowl too low. This is a surprising example to students, because the face value for the cheapest ticket for Super Bowl LX this year was $950. However, we know that this price is too low because the cheapest tickets on the secondary market were selling for around $6,400. That suggests that the NFL is leaving money on the table - they could earn much more if they set the ticket prices higher. Why would the NFL set the price lower than the short-run profit-maximising price for the Super Bowl? One reason may be that they want to maintain a long-term relationship with NFL fans. They want going to the Super Bowl to be an achievable aspiration for fans. Most fans won't be able to pay US$4,000 (plus the travel and accommodation costs) to attend every year, but at that price it is reasonable for fans to believe that they can attend once during their lifetime. If Super Bowl tickets cost over US$10,000, then that aspiration becomes much less achievable. This long-term strategy is also visible through the allocation of Super Bowl tickets. Each team gets an allocation of tickets, 35% of which must go to fans, and that allotment is typically given to the team's season ticket holders (see here).

The Super Bowl is recognisable to my students. However, thanks to this article in The Conversation by Paul Crosby (Macquarie University), I now have another example that is somewhat closer to home (albeit not necessarily more recognisable than the Super Bowl):

It’s AFL finals time – and in a season on track for record attendance, the league’s decision to freeze the price of entry-level finals tickets for an 11th straight year is smart economics...

So why are entry-level footy tickets staying cheap? It’s all about investing in the future – especially when there’s a lifetime of spending at stake...

Finals matches routinely sell out, so standard economics would suggest raising prices.

Instead, entry-level tickets for the AFL’s qualifying, elimination and semi finals, as well as this weekend’s wildcard round, have been frozen at A$35. Entry-level preliminary finals tickets will be $65, unchanged since 2016.

Crosby explains this pricing strategy as arising from fairness and goodwill towards fans, then notes that:

Cheap tickets can be an investment. A full stadium creates value beyond ticket revenue. Crowds generate the atmosphere that makes live sport attractive to television audiences and, in turn, valuable to broadcasters and sponsors.

An AFL supporter will often remain attached to the same club for decades, buying memberships and merchandise, watching broadcasts and eventually bringing their children along. A fan won over by a $35 ticket is worth far more than the profit on that ticket.

Keeping tickets affordable, and occasionally letting kids in free, helps recruit that next generation.

By keeping the price of finals tickets low, the AFL is foregoing short-run profit maximisation in favour of a long-term strategy that keeps fans engaged, keeps them attending games, and may lead to greater revenue and profits overall in the long run.

Crosby compares the AFL's pricing strategy with the approach adopted by music acts, who are increasingly using 'dynamic pricing' to maximise short-run profits from every concert. However, the situations are different in a meaningful way. The AFL can afford to have a long-term focus because it has a revolving cast of star players, while the teams endure. AFL fans typically follow a team, rather than particular players. In contrast, not every musical act is going to be The Rolling Stones, still touring after 60-plus years. For the most part, a musical act is not a revolving cast of musicians (especially for solo acts). A musical act's window for profiting from their talent is much shorter than the AFL, so we might expect to see short-run profit-maximising behaviour from musical acts than for the AFL. It is all about maximising profits over the appropriate time horizon. For the AFL or the NFL, a price that looks too low for today's ALF finals game or Super Bowl may be exactly the right price for maximising profits over the lifetime of a fan.

Saturday, 22 August 2026

Book review: Economics for Dummies

Since I recently read Econometrics for Dummies (which I reviewed here), it seemed fitting that I follow up with Economics for Dummies, by Sean Flynn. I read the second edition (published in 2011), but I see that there is a third edition now (published in 2019). While my impression was that the econometrics book would be a good companion to a 'traditional' textbook on econometrics, Economics for Dummies is very much a textbook treatment of microeconomics and macroeconomics, which would not be out of place as the required textbook for an undergraduate course in the principles of economics. While Flynn does a good job of keeping things light, as you might expect from a book in Wiley's 'Dummies' series, this is not really a 'pop economics' book.

That disappointed me a little, because I wasn't expecting to read a textbook. Having said that, Flynn does take a refreshing approach to some parts. I really liked the way that he builds up the demand curve from a starting point of marginal utility. I hadn't seen it done that way before. And there were several parts of the macroeconomics section, including a more complete circular flow of income that we often see in textbooks, and a surprisingly thorough discussion of Keynesian economics, that I appreciated. If I was teaching the macroeconomics section of my ECONS101 class, I would certainly be drawing on this book to assist with those aspects.

The book does have several surprising blind spots though. It covers asymmetric information and adverse selection, but never discusses the roles of signalling or screening as solutions to adverse selection problems. It also contains a few errors, including explaining that supply curves slope upwards because of increasing costs (it should be increasing marginal costs), and referring the government spending as part of monetary policy (it is part of fiscal policy). Possibly those bits got fixed up in the more recent edition. I sure hope so.

Overall, I wouldn't strongly recommend this book, but only because there are a generous number of economics principles textbooks around. This book, unfortunately, doesn't really stand out from the crowd. If you had to read it as a textbook, it would be fine, but I wouldn't prefer it over any of the other mainstream options.

Friday, 21 August 2026

This week in research #140

Here's what caught my eye in research over the past week (another quiet week, it seems):

  • Higney et al. (open access) estimate the marginal value of bird song, using a choice experiment with a 'simulated soundscape'
Also new from the Waikato working papers series:
  • Gibson, Boe-Gibson, and Scrimgeour use night-time lights to measure the urban expansion and contraction of urban areas in New Zealand, finding that rapid expansion is concentrated in commuter towns surrounding major cities and amenity-oriented settlements, while some smaller towns experience contraction associated with economic restructuring and changing locational advantages

Thursday, 20 August 2026

John List on critical thinking and teaching economics

When it comes to buzzwords in education, critical thinking probably ranks at or near the very top. The problem is that not everyone agrees what critical thinking is, how students can be taught to think critically, or even whether critical thinking is something that can be explicitly taught at all (as opposed to being tacit knowledge). So, it can be instructive to see how other people, particularly top thinkers in your own discipline, conceptualise critical thinking, and how they propose training students to think critically.

So, I was glad I finally round time this week to read this 2022 article by John List (University of Chicago), published in the Journal of Economic Education (ungated earlier version here). He develops a 'Critical Thinking Hierarchy', and then describes how this can be applied to student learning.

List defines critical thinking as "individual skills that facilitate logical and informed decisions", and notes that:

In my own experiences, these skills can naturally be divided into two complementary pillars:

  • Connecting the dots with empiricism: developing and assimilating empirical evidence and updating of one’s beliefs

  • Connecting the dots with abstract thought: putting the puzzle together with conceptual reasoning; thought experiments

The hierarchy that List develops is then described in Figure 1 from the paper:

The hierarachy starts with modal thinking, characterised by various biases, preconceptions, and prejudices. It then moves to neophyte thinking, where the importance of thinking and empiricism begin to be recognised, but the execution isn't yet there. Adept thinking progresses to critical questioning, recognising blind spots, and understanding causation. Finally, great thinkers understand and correct for their own biases, and constantly re-examine their assumptions. An obvious question, though, is how a teacher moves their students up through the hierarchy. List advocates that teachers should encourage their students to 'slow think' (drawing inspiration from Daniel Kahneman's research, as outlined in his book Thinking, Fast and Slow). Specifically, List recommends that we teach students to apply six basic tenets:

1. state, explain, and clarify the question(s)

2. think through the question(s) from multiple points of view, expressing their own priors using logical thinking

3. gather, organize, assimilate information and data

4. identify assumptions, shortcomings, and implications of the data generation process

5. update priors, both their own priors and consider how other’s views might change

6. explain and apply what they learn, connecting what they just learned to other economic concepts, learnings from another course, and/or their everyday life

He then offers a set of necessary conditions through which those six basic tenets can be embedded to develop critical thinking (CT):

A. Early discussion of simple empirical tools; distinction between correlation and causation; and, provide examples throughout the course that distinguish causation and correlation

B. Theory of mind should be reinforced (I cannot think of a better place than game theory); briefly introduce psychological biases that prevent the student from becoming Adept

C. Because CT development is a social activity, both lab and field experiments should be used as pedagogical devices to promote CT...

D. Connect theory to empiricism and highlight potential shortcomings of both without undoing the major insights

Finally, the punchline is that this approach is part of the textbook that List co-authored with Daron Acemoglu and David Laibson. So, was the whole article an elaborate advertisement for the textbook? Perhaps, but I think there is a greater value in what List has shared. Right now, my colleagues and I are re-designing the economics curriculum at the University of Waikato from the ground up. We haven't explicitly embedded critical thinking in that design and yet, when I look at the four necessary conditions that List has proposed, and the two critical thinking skills that the article starts with, I can definitely see those in what we are developing. In particular, linking economic theory to empirical research and insights is already a particular strength of economics at Waikato.

List wrote his article in 2021, and it was published in 2022, before the release of ChatGPT and the widespread adoption of large language models. It would be reasonable to ask whether generative AI would make a difference to what List proposes. I don't believe that it change the core of List's framework. In fact, it makes the framework more important and changes how we should implement it, because the thoughtful integration of generative AI into higher education may lead to greater opportunities for students to engage in intentional development of the critical thinking skills that List proposes. Think about the application of simple empirical skills. These are skills that students can learn to use alongside generative AI, if their interactions are intentionally designed to promote engagement and learning, rather than cognitive offloading. For example, students might use generative AI to collate data and test an empirical relationship, while thinking about how they might distinguish causation from correlation. A custom GPT would allow for limitation in the range of techniques that could be employed so that novice students don't get overwhelmed. Lab and field experiments can be integrated into classroom learning, and I know that both Steve Tucker and I do that in our classes already. Experiments can be incredibly powerful tools for promoting student learning.

Critical thinking may be an education buzzword, but that doesn't mean that it isn't important. I like List's metaphor of 'connecting the dots'. As generative AI continues to develop and becomes a more useful tool across a wider range of analytical and other activities, connecting the dots alone will not be enough.  Students will need to know which dots matter, whether they genuinely connect, and what conclusions the connected dots allow them to draw.

Tuesday, 18 August 2026

The impact of ransomware attacks on hospitals and patients

In May 2021, the Waikato District Health Board (DHB) was hit with a ransomware attack. It took some four weeks for clinical services to be restored, and in the meantime, surgeries were postponed, and patients and health staff were negatively impacted. Health services have been a common target of these ransomware attacks, and the consequences could be tragic. Fortunately, in the case of the Waikato DHB attack, there is no evidence of patients dying as a result.

That isn't always the case though. This recent article by Hannah Neprash, Claire McGlave, and Sayeh Nikpay (all University of Minnesota), published in the American Economic Journal: Economic Policy (ungated earlier version here) looks at the impact of ransomware attacks on hospitals in the US. They use Medicare administrative claims data, along with data from HackNotice and the Office for Civil Rights Breach Portal on ransomware attacks on hospitals. They find 74 attacks over the period from 2016 to 2021, affecting 160 hospitals.

Neprash et al. then look at the effect of the ransomware attack on hospital volume (separating emergency room, inpatient, and outpatient volume), and hospital revenue from Medicare, as well as patient mortality. They use a difference-in-differences analysis, which involves looking at the difference in each outcome variable between the time before and the time after the ransomware attack, for hospitals that were attacked, and those that were not. They identify control group hospitals (that weren't attacked) as those most similar to the affected hospitals in terms of non-profit status, health system membership, and quartile of Medicare admissions in the year prior to the attack. They also conduct an event study, which allows them to look at how the impact changes over time.

In their main analysis, Neprash et al. find that:

During the initial week of a ransomware attack, hospital volume falls by 17–24 percent in the ER, inpatient, and outpatient settings. Medicare revenue declines by 19–39 percent at ransomware-attacked hospitals. A full recovery to pre-attack volume and revenue occurs within two to three weeks on average. A back-of-the-envelope calculation suggests that the average ransomware attack reduces annual hospital revenue by roughly 1 percent.

These are quite substantial effects (and notice the recovery time is not dissimilar to the case of the Waikato DHB attack). What happens to the patients who are affected by their hospital being attacked? It turns out that many patients were redirected to nearby hospitals, although the extent of redirection differs by type of care, as when Neprash et al. look at the local hospital market rather than the individual hospital, they find that:

...nearby hospitals absorb displaced emergency department patient volume from attacked facilities, such that market-level emergency department volume does not change during attacks. Inpatient and outpatient hospital volume is partially absorbed by nearby hospitals, though not fully, resulting in a market-level volume decrease during the first week of a ransomware attack.

That also means the costs of a ransomware attack spill over to neighbouring hospitals, which have to absorb some of the displaced patients. What about patient outcomes? Here are the most serious impacts, as Neprash et al. find that:

...ransomware attacks increase in-hospital mortality for patients already admitted to ransomware-attacked hospitals when the attack begins, compared to patients whose admissions concluded in the five weeks prior. Our estimates suggest that ransomware attacks resulted in the deaths of between 69 and 76 Medicare patients—representing roughly 1 Medicare death per month due to ransomware over the course of our study period.

Notice the effects fall on patients who were already in hospital care at the time the attack started. Many of those patients may not be easily redirected to other hospitals, and so have little choice but to ‘ride out’ the attack in the affected hospital. Neprash et al. also find larger mortality impacts at smaller or independent hospitals, during particularly severe attacks, and among patients with complex care needs (such as patients in intensive care, or those with multiple chronic conditions). Neprash et al. don't offer much of a policy prescription in their discussion of their results, limiting themselves to recommending:

...a combination of policies designed to reduce the likelihood of any successful ransomware attacks (e.g., minimum cybersecurity standards for hospitals) and policies designed to reduce the severity of ransomware attacks when they do happen (e.g., incident planning requirements).

For me, the key takeaway from this research is that ransomware attacks are not just a cybersecurity issue, they are a health security issue. We can be thankful that the Waikato DHB attack avoided the much worse outcomes that US hospitals have experienced. The response nevertheless required serious efforts by IT professionals and imposed a heavy workload on health and administrative staff. We should treat these results as a warning that hospitals need to treat resilience to cyberattacks as part of their core patient-safety planning, rather than simply as an IT problem.

Monday, 17 August 2026

Discrimination on #EconTwitter

I've written a number of times about correspondence experiments designed to identify discrimination in labour markets. In such an experiment, the researcher applies for a bunch of jobs, using fake job 'applicants' that differ only on the basis of some known characteristics (gender, for example). The difference in callback rates (or some other similar measure) between applicants with different characteristics provides a measure of discrimination on the basis of those characteristics.

A new application of this approach instead looks at discrimination on #EconTwitter (on X). The research is reported in this 2025 article by Nicolás Ajzenman (McGill University), Bruno Ferman (Sao Paulo School of Economics), and Pedro Sant’Anna (MIT), published in the journal American Economic Review: Insights (ungated earlier version here). In their experiment, Ajzenman et al. created 80 'bot' accounts on X between May and August 2022 that mimicked real PhD student accounts, but differed in terms of gender (male or female), race (Black or White), and university affiliation (top-ranked [top ten in the 2017 US News ranking of economics graduate programmes] or lower-ranked [ranked 79-100] university). Each bot account was active for twelve days, during which time it initially randomly retweeted posts from economics journals to establish its credibility, then followed a random selection of 100 members of the #EconTwitter community (from a pool of over 10,000 X users who had tweeted or retweeted content using the #EconTwitter hashtag between January and February of 2022).

Ajzenman et al. then measure the number of times each account is followed back by the user it followed, and look at the difference in follow-back rates for bot accounts with different characteristics (gender, race, and university affiliation). The raw follow-back rates are shown in Figure 1 in the paper:

Bot accounts that presented as Black males from lower-ranked universities received the lowest follow-back rate of 14.4 percent, while bots presenting as White females from top-ranked universities were followed back 23.9 percent of the time. In their regression models, Ajzenman et al. find that:

Users in the #EconTwitter community are 2.1 percentage points (12 percent) more likely to follow White than Black PhD students, 3.5 percentage points (21 percent) more likely to follow students from top-ranked universities than those from lower-ranked ones, and 4.3 percentage points (25 percent) more likely to follow female than male students. We also find that the racial gap in follow-backs remains among students claiming to be from top-ranked universities. This suggests that racial discrimination persists even in the presence of a signal indicative of higher academic potential.

Turning to how the results vary based on the characteristics of the X users that were followed, Ajzenman et al. find no differences by gender or race, or by the reach of the X user (measured by the number of followers they have). They also find no difference between X users who have demonstrated concern about the lack of diversity in economics (by following one or more X accounts devoted to the topic) and those that haven't. However, they also find that:

...subjects who signal concern about the lack of diversity in economics discriminate the most in terms of university affiliation. While both groups of subjects favor students affiliated with top-ranked institutions, the difference in follow-back rates between top- and lower-ranked students is considerably larger among concerned subjects (8.6 percentage points, against 2.5 percentage points for the remaining subjects, p-value of difference < 0.05)...

So, concern about demographic diversity clearly doesn't imply an absence of other forms of discrimination, as these results seem to show that those members of the #EconTwitter community who are most concerned about diversity discriminate more against students from lower-ranked institutions than users who are less concerned.

Given the longstanding gender bias in economics, the higher follow-back rates for bot accounts presenting as female are perhaps the most surprising result. Ajzenman et al. suggest several possible explanations. First, some users may be conscious of the barriers women face in the profession and therefore make a deliberate effort to engage with them. Second, some users may be using X partly to establish social rather than professional relationships, which would suggest a more negative interpretation of the higher follow-back rate. Ajzenman et al. also suggest two more strategic explanations. Researchers may see advantages in collaborating with women because female economists tend to receive less credit for joint work. Alternatively, a woman gaining admission to a highly ranked PhD programme may be providing a stronger signal of ability, given the additional barriers women face in reaching that point. The experiment cannot distinguish among these explanations, so teasing out which mechanisms are more important would require further, more detailed, research.

And no doubt Ajzenman et al. had hoped to conduct some more detailed research than what they reported, but their data collection had to be cut short. As they explain:

We planned to run 30 experimental waves between May and December 2022, which would have given us more than enough power to identify reasonable effects. This was to account for potential problems, such as X blocking some accounts. We stopped earlier because an X user saw some of the accounts and posted about the experiment during the eleventh wave, which compromised the continuity of the experiment.

Sometimes research just doesn't go as planned, and that might explain why this research got published in the more modest AER: Insights journal, rather than the top-five journal the authors probably were initially hoping for.

Sometimes research just doesn't go as planned. However, Ajzenman et al. had already collected enough data to reveal a substantial amount of discrimination on #EconTwitter. Perhaps the most important result is that signalling concern about diversity doesn't necessarily make people immune to other forms of bias, especially in terms of academic prestige.

[HT: Marginal Revolution]

Sunday, 16 August 2026

Computer gaming and binge drinking may be complements, not substitutes

In economics, two goods are substitutes if consumers tend to consume more of one if the price of the other increases. One way of thinking about that is that if the price of Good X increases, consumers switch to purchasing Good Y instead, and the quantity of Good Y demanded increases. Two goods are complements if consumers tend to consume less of one if the price of the other increases. In this case, if the price of Good X increases, consumers buy less of Good X (due to the Law of Demand), but also buy less of Good Y, and the quantity of Good Y demanded decreases.

Whether a pair of goods are substitutes or complements is determined by the cross-price elasticity of demand: the responsiveness of the quantity demanded of one good to a change in the price of the other good. If the cross-price elasticity is positive, the two goods are substitutes. If the cross-price elasticity is negative, the two goods are complements. Another way of thinking about this is that, following a change in the price of one good, ceteris paribus (holding all else constant), we would expect the quantities demanded of substitutes to move in opposite directions, while the quantities demanded of complements would move in the same direction.

There are obvious examples of substitutes and complements. Coke and Pepsi are the iconic example of substitute goods used in almost every introductory economics class. An example of complements that I use in my classes is video game consoles and games. However, it isn't always straightforward to determine whether a pair of goods are substitutes or complements. Sometimes they may be substitutes in one context, but complements in another. So, whether goods are substitutes or complements is an empirical question.

Take the example of computer gaming and binge drinking. When I was growing up, those two 'goods' certainly seemed like complements. My friends and I spent many nights drinking beer or RTDs and playing hotseat turn-based strategy games like Robosport, Warlords II, or Heroes of Might and Magic.[*] That experience made me a little surprised to see the hypothesis in this 2021 article by Torleif Halkjelsvik, Geir Brunborg, and Elin Bye (all Norwegian Institute of Public Health), published in the journal Drug and Alcohol Review (open access), which was that binge drinking and computer gaming are substitutes. Now, modern computer gaming differs in meaningful ways from how it looked when I was young. Nevertheless, I was surprised that Halkjelsvik et al. hypothesised in the direction they did.

Their hypothesis rested on several ideas, and was motivated by the observed increase in gaming and decrease in alcohol consumption by young people over time. First, alcohol and gaming are both outlets for thrill seeking, and are both responses to boredom, so increasing computer gaming might reduce the need for drinking. Second, both drinking and computer gaming are sources of social bonding, so again more computer gaming reduces the need for drinking.

Halkjelsvik et al. test their hypothesis with data from the European School Survey Project on Alcohol and Other Drugs (ESPAD), which surveys 15 and 16-year-old students every four years. They use data from 23 countries over the period from 1995 to 2015 (although noting that not all countries are part of the survey in every year), and look at the correlation between frequency of binge drinking (drinking five or more drinks on an occasion) and frequency of computer gaming, using a multi-level linear probability model. If their hypothesis that gaming displaces drinking is correct, the relationship should be negative. However, Halkjelsvik et al. find that:

...the association between country-level changes in computer gaming and binge drinking was estimated as positive...

So, increases in the average frequency of computer gaming at the country level tended to be associated with increases in the frequency of binge drinking. And, at the individual level:

The between individual-effect was positive, suggesting a four percentage point (±2 percentage points) higher binge drinking prevalence among students who report playing computer games daily.

Of course, the analysis that Halkjelsvik et al. conducted doesn't establish a causal relationship, it only shows correlations. And, importantly, they aren't directly testing whether computer gaming and binge drinking are complements in the economic sense, as that would require looking at how consumption of one responds to changes in the price of the other. However, their results are at least consistent with computer gaming and binge drinking being complements. Rather than moving in opposite directions, as we might expect if gaming displaced drinking (as Halkjelsvik et al. hypothesised), gaming and binge drinking tend to move in the same direction. Which, admittedly on the basis of a rather smaller and less representative sample, my friends and I could have told them.

*****

[*] My kids are bemused at the very idea that there was ever such a thing as hotseat multiplayer games. Sadly, they gradually died out as online games became more widely available in the early 2000s. However, they were really good for multi-tasking with some tabletop gaming at the same time, since only one player played the hotseat game at a time.

Saturday, 15 August 2026

Taking advantage of loss aversion in education

Many years ago (I forget exactly when), I introduced extra credit into my ECON110 class (which is what is now ECONS102). The idea was to provide an incentive for students to attend class, since they could earn extra credit for completing various in-class exercises. A couple of years later, I briefly changed the way that I framed the extra credit, from being "extra marks that would be gained from attending", to "extra marks that would be lost by not attending".

If students were purely rational, the change from 'gain framing' to 'loss framing' the extra credit should have had no impact on student attendance. However, I was looking to exploit the fact that most people are quasi-rational, rather than purely rational. Quasi-rational decision-makers are loss averse, meaning that they value losses more than equivalent gains. For a loss averse person, losing $20 makes them unhappy to a greater extent than winning $20 makes them happy.

Does a change from 'gain framing' to 'loss framing' work? That is the question that this new article by Antal Ertl, Éva Holb (both Eötvös Lóránd Science University), and Barna Bakó (Corvinus University of Budapest), published in the Journal of Economic Behavior and Organization (open access), tries to answer. They use data from a field experiment at Corvinus University of Budapest, where students enrolled in a compulsory macroeconomics course for business students were randomised into one of three conditions: (1) Gain group, which earned points in each of four tests and the final examination as usual; (2) Loss group, which started each test and the final exam with full points, but had points deducted for each incorrect answer; and (3) Hybrid group, which was the same as the Gain group for the tests, but switched to the loss framing for the final examination.

Ertl et al. have a sample of 321 students who consented to be part of the research, completed an initial questionnaire at the start of the term, and earned a non-zero grade. Randomisation was conducted at the level of the tutorial group (so all students in a tutorial were in the same treatment), in such a way that each teacher had groups across more than one treatment. One wrinkle in their analysis is that the best three out of the four tests would count towards a student's grade, meaning that students may end up putting differential effort into each test, depending on how they have performed in the other tests already completed. So, in addition to looking at the effect of treatment on each test mark individually, Ertl et al. look at the effect on the 'best three' tests collectively, as well as the exam mark.

If randomisation were perfect and the treatment groups were balanced, the comparison between the Loss group and the Gain group would demonstrate the overall effect of loss framing on student performance. The comparison between the Loss group and the Hybrid group for the final exam, compared with the same comparison for the best three tests, would demonstrate whether students adjust in such a way that the loss framing has less impact over time (because the Hybrid group would be in their first loss-framed assessment, while the Loss group would be in their fifth such assessment). The treatment groups weren't perfectly balanced, with students sorting into tutorial groups in part based on whether they worked part-time. So, Ertl et al. control for working part-time, the tutorial day and time, and the tutorial group teacher, as well as other demographic and background variables.

In their main analysis, they find support for the positive effects of loss framing:

For the Loss treatment, the effect on the average of the Best 3 Tests is 3.2 percentage points, although the difference is not statistically significant. The treatment effect on the Final Test score, however, shows a large difference of 9.6 percentage points when not controlling for Best 3 Tests’ scores, i.e., how well students did throughout the semester before the Final Test.

After controlling for performance in the best three tests, the effect of the loss framing on performance in the final examination is a statistically significant 7.8 percentage points. Turning to the comparison of the Loss and Hybrid groups, Ertl et al. find that:

...the estimated effect sizes for Loss and Hybrid are essentially the same for the Final Test, once we take into account how well students did perform throughout the semester...

These results are consistent with loss framing leading to better student performance, and there being no novelty effect - the effect of loss framing doesn't appear to decline over time. Ertl et al. go on to show that the effects are similar for both male and female students, but larger for students who did not take advanced mathematics in high school than for those that did. They also show that the treatment did not seem to negatively affect students' perceptions of the course, because the teaching evaluations were similar for the different treatment groups.

Finally, Ertl et al. do provide a note of caution in their conclusion:

previous studies have highlighted possible psychological and motivational costs associated with loss framing... These findings suggest that the mechanism by which loss framing improves performance may, at least in part, operate through heightened tension and concern about avoiding mistakes rather than through enhanced intrinsic motivation. Moreover, in extreme cases, loss-framed grading may even produce adverse effects — for example, low-performing students might become discouraged early in the semester after ‘‘losing’’ too many points. Once it becomes apparent that only a passing grade is attainable at best, the loss-framed structure may make this limitation increasingly salient, potentially exacerbating anxiety and disengagement. Over time, this could have broader implications for students’ well-being and their willingness to enroll in courses or programs that employ such systems.

Ertl et al. don't directly test for these effects, but they should be a concern. We may be able to improve student performance through loss-framing assessments, but that might come at a cost to student mental health and wellbeing.

And that brings me back to the example I started with, from my ECON110 class. When I switched extra credit from gain-framed to loss-framed, student attendance in class did improve slightly. However, the bigger impact seemed to be the number of students who would contact me by email, seeking special consideration for missing the extra credit, offering to provide medical certificates or other evidence to explain their absence, and asking for extra chances to complete the in-class exercises. It turned out to be administratively much more costly for me, and so the change was short-lived (to the extent that I cannot even remember which year I tried this in). Those reactions could suggest a negative psychological effect of the switch from gain framing to loss framing.

So, not all interventions that are effective for promoting student performance should be adopted. We need to carefully consider both the benefits and the costs of the intervention first. Taking advantage of student loss aversion might be worth exploring further, but I would want to see a wider evaluation that included student wellbeing outcomes before adopting it.

Friday, 14 August 2026

This week in research #139

Here's what caught my eye in research over the past week (a quiet week, it seems):

  • List (open access) comments on how to address the generalisability of research
  • Lehner et al. (open access) find that the opening of a Walmart Supercenter is associated with a 2.2 percentage point (18%) increase in poverty, and that the increase is largest for younger and less-educated adults

And the latest paper from my own research (led by my former PhD student Muhammad Irfan, along with Ushan Goonawardane, and Craig Robertson), which was also covered in the New Zealand Herald:

  • Our new article (open access if you register for free) in the New Zealand Medical Journal performs a comparison of methamphetamine contamination of 423 properties across New Zealand, before and after the requirement to test for methamphetamine was eased in May 2018, and finds a significant increase in methamphetamine contamination

Thursday, 13 August 2026

Customers shouldn't pay less when they use a self-checkout, they should pay more

The New Zealand Herald reported last week:

State representative Nikki Lucas has introduced a bill that would require retail businesses selling food in the state to offer a 10% discount to those who used the self-checkout lane.

“Retail businesses increasingly rely on self-checkout systems to reduce staffing and operational costs by shifting responsibilities traditionally performed by employees onto consumers,” she wrote...

Consumer NZ head of advocacy Gemma Rasmussen said her organisation thought there was validity to the argument in New Zealand, too.

Call me radical, but I think that Lucas and Rasmussen have this backwards. Customers shouldn't pay less when they use a self-checkout, they should pay more. To see why, I'm going to rely on the concept of price discrimination - where the seller sells the same good or service to different groups of consumers for different prices.

Consider two groups of consumers (impatient, and patient), and two options (self-checkout, and regular checkout). The first group of consumers is impatient, and they want to get out of the store as soon as possible, and for that reason they prefer to use self-checkout. This group can be said to have a short time horizon for their purchases. This short time horizon makes their demand for goods less elastic (less sensitive to price). The second group of consumers is more patient, and they are willing to wait. This group can be said to have a longer time horizon for their purchases, which makes their demand for goods more elastic (more sensitive to price).

If supermarkets want to price differently for each group, which group should pay the higher price? The answer to that question is shown in the two diagrams below. Both diagrams show a firm with market power (a supermarket), and each diagram corresponds to one of the sub-markets. The sub-market on the left represents the patient buyers, who have more elastic demand - notice that the demand curve D1 is relatively flat (which means that a change in price will have a big effect on the quantity that these consumers demand). The sub-market on the right represents the impatient buyers, who have less elastic demand - notice that the demand curve D2 is relatively steep (which means that the same change in price would have a smaller effect on the quantity that these consumers demand, than it would for the patient consumers). The marginal cost (MC) is the same in both sub-markets - it doesn't cost the supermarket any more to sell a product to an impatient buyer than what it costs them to sell that same product to a patient buyer. [*]

The supermarket will maximise profits by selling the quantity where marginal revenue (MR) is equal to marginal cost (MC) - this is the standard short-run profit-maximising condition (as I discussed in this post). In the impatient sub-market, the profit-maximising quantity occurs where MR2=MC, which is Q2. In order to sell that quantity in the impatient sub-market, the supermarket should set the price equal to P2. The problem with that high price P2 is that in the patient sub-market, no consumers would be willing to buy the good at all. The supermarket can increase profits if it charges a different price in the patient sub-market from the price it charges in the impatient sub-market. In the patient sub-market, the profit-maximising quantity occurs where MR1=MC, which is Q1. To sell that quantity in the patient sub-market, the supermarket should set the price equal to P1. In other words, the supermarket should charge a higher price to the impatient consumers, and a lower price to the patient consumers.

The problem here is that supermarkets don't know (for sure) which group (impatient or patient) any particular consumer belongs to. But by offering different checkout options, the customers can sort themselves into the impatient (less elastic demand) group and the patient (more elastic demand) group, because the impatient consumers use the self-checkout. In other words, the supermarket should charge a higher price to the users of the self-checkout.

This is an example of menu pricing (or second-degree price discrimination) - where the consumers are presented with a menu of options, and they select the one they prefer.  Crucially, the seller knows that some menu options appeal to consumers with more elastic demand, and other options appeal to consumers with less elastic demand. In this case, there are two menu options - self-checkout, or regular checkout, and the supermarket knows that the self-checkout appeals to the impatient consumers who should be charged a higher price.

So, customers who use a self-checkout right now shouldn't be arguing to lower prices. They should think themselves lucky that supermarkets aren't optimising, because if they were, the prices at self-checkouts would be higher than at regular checkouts.

*****

[*] You could argue that it doesn't cost the same to offer purchase through regular checkouts and self-checkouts. However, how big is the cost difference, really? Let's say that it takes two minutes to scan your items, but would take three minutes through the regular checkout, because the payment process tends to take a bit longer at a regular checkout. With self-checkout, the supermarket would save three minutes of labour. Say that the supermarket pays their checkout staff $30 per hour (somewhat more than the minimum wage). By using the self-checkout, you've saved the supermarket $1.50 of labour in this example (3/60 * $30). Except, that calculation doesn't take into account that the self-checkout is not a zero-labour option. There is usually a checkout person who has to watch over the consumers using the self-checkout. So, the saving is actually a bit less than that. It almost certainly isn't close to the 10 percent discount that Lucas is arguing for. Most of the cost of the items that you buy at the supermarket is the wholesale cost that the supermarkets pay, not the checkout labour cost.

Tuesday, 11 August 2026

Generative AI, cognitive offloading, and escaping the 'illusion of competence'

Last week, for maybe the first time, I found myself telling a student not to use AI. That might sound extraordinary. After all, generative AI hit the big time with the release of ChatGPT in November 2022, and increasing numbers of students have been using it ever since. Many lecturers immediately freaked out, and many institutions initially reacted by banning or restricting AI use, before realising that they were fighting a losing battle against the incoming tide of generative AI and trying to impose 'guardrails'.

I've never asked my students not to use generative AI. In fact, I've encouraged it. I even have custom AI tutors set up for each of the papers I teach, that use a knowledge base of materials from the paper to give students targeted assistance. I have little to fear from generative AI, because the vast majority of assessment in my papers is in-person and invigilated (that's one of the beautiful things about teaching first-year papers - I can argue that basic concepts and applications can be authentically assessed in an exam environment).

Anyway, back to the story. The student was using our class AI tutor during class, to give them a solution to a problem we were working on during class. I pointed out that it defeated the purpose of doing the problem in class, if they used Jane (our ECONS102 AI tutor is named after Jane Marcet, the author of the 19th-Century popular economics book, Conversations on Political Economy) to solve it for them. The problem wasn't the use of Jane per se (after all, I encourage them to use her). It was that by using Jane to solve the problem for them, they were missing out on a key learning opportunity.

Probably, I was a little hard on the student. After all, they were using the tools available to them, and engaging in cognitive offloading - reducing the demand or mental load that they face by offloading a task onto generative AI. And increasingly, students are engaging in this cognitive offloading, sometimes in helpful ways, but often in ways that are detrimental. That is one of the conclusions from this 2026 report (with non-technical summary on The Conversation) by Jason Lodge and Leslie Loble (both University of Technology Sydney).

The report has a lot of quotable quotes. For instance, they note that:

It is not possible to engage in critical thinking when one has nothing to think critically about. A person does not simply think critically in a vacuum. A scientist thinks critically about a flawed methodology by drawing on a vast store of knowledge about experimental design. A historian thinks critically about a primary source by drawing on their knowledge of the document’s social, political, and historical context...

This, to me, highlights the key challenge that education faces with generative AI. In order for students to be well prepared for engaging with generative AI, they need to be able to evaluate AI output. And without a thorough grounding in disciplinary knowledge, their evaluations would at best be superficial. And that is why I hold the line on having assessment in my papers that explicitly excludes the use of generative AI. My papers build the foundation on which students' later use of generative AI, and their evaluation of AI outputs, can build.

Lodge and Loble note that:

Every task or learning activity is essentially now a group activity. It just so happens that the other member or members of the group are machines that have practically all human knowledge at their fingertips (in their databases/algorithmic weights). Like any other group activity, students can benefit from that collaboration or get the smart kid to do all the work for them.

That is absolutely what is happening. Students' learning activities are now mostly group activities, even when they are the only human in their group. In group work, how the work is shared is important, and that is where cognitive offloading comes in. Lodge and Loble distinguish between two forms of offloading:

  • Beneficial offloading occurs when AI is used to manage extraneous cognitive load (e.g., checking grammar), freeing a learner’s limited working memory to focus on essential, intrinsic tasks.

  • Detrimental offloading (outsourcing) occurs when a learner uses AI to bypass this intrinsic cognitive effort (the desirable difficulties) required to build long-term knowledge schemas. This offloading also seems to extend to vital metacognitive and self-regulated learning capabilities, compounding the negative impact of outsourcing on learning.

Importantly, Lodge and Loble note that students typically don't understand metacognition. They haven't intentionally engaged in thinking about their own thinking and understanding how they learn or managing that process. In my experience, many students tend to have been passive recipients of learning approaches, without really engaging with the process themselves. And even those that do engage usually haven't thought deeply about how they learn. And so, when they use a tool that gives them ready answers, it may seem to students that they are learning more efficiently. Lodge and Loble label this an 'illusion of competence', noting that:

Research has long shown that fluent learning materials, such as high-quality videos, can lead people to greatly overestimate how much they have learned by mistaking the ease of processing (fluency) for the depth of learning...

Lodge and Loble's report doesn't stop at the point of diagnosing the problem though. They present three main solutions, that involve:

  • shifting generative AI use towards beneficial offloading, where students free up cognitive resources to focus on intrinsic learning. Lodge and Loble offer the example that "AI can be used to provide scaffolding, structured practice, and feedback, all aimed at managing the cognitive burden on the learner and enabling progressive independence...";
  • deliberately designing AI interactions to include metacognitive responsibilities, so that students must pause, reflect, and assess their own understanding; or
  • shifting the fundamental role of AI from being an 'answer oracle' to a tool that provokes intrinsic load. Lodge and Loble offer examples such as asking students to teach the AI (which plays the role of a confused student), setting up AI as a Socratic tutor, or asking students to independently verify AI outputs.

Those solutions have implications for how I design and use AI tutors in my papers. In order to limit students from engaging in detrimental cognitive offloading, the AI tutor shouldn't simply act as an answer machine. Instead, they should encourage students to attempt problems themselves, offer hints or scaffolding when they get stuck, and ask them to explain or justify their reasoning. And, importantly, the AI tutor could also prompt students to reflect on what they understand, what they don't understand, and whether they could solve the problem on their own. The master prompt for my AI tutors does instruct them to take a Socratic approach, but they don't adhere to it strictly. I'll certainly be putting more thought into the master prompt to see if I can dissuade them from being answer machines and to incorporate more of the metacognitive elements in the future.

The solutions provided by Lodge and Loble are useful, and hopefully they prompt other lecturers to think intentionally about students' (and possibly their own) engagement with generative AI. I especially like the second option, because I believe that we all (and not just students) can benefit from better understanding our learning process, and recognising when we are engaged in genuine and effortful learning. This is not the first time I've encountered concerns about cognitive offloading in the context of generative AI and education (see here). And it is interesting that the same, or a similar, set of solutions keep being presented. In particular, integrating technology as a complement, rather than a substitute, for thinking is important. Generative AI has dramatically lowered the cost of getting answers. It hasn't lowered the cost of learning. A better understanding of metacognition is therefore important too.

Perhaps, then, the lesson from my interaction with the student isn't that they shouldn't have been using generative AI in class. It's that they need to understand when using generative AI in class supports their learning, and when it substitutes for the thinking that learning requires.

I think most universities are now considering explicitly including generative AI in the curriculum, in order to better prepare students for future careers that will no doubt involve substantial interactions with generative AI. Perhaps we should also be considering explicitly including metacognition in the curriculum?

Read more:

Monday, 10 August 2026

The impact of using the CORE textbook in Uruguay

We introduced the CORE textbook The Economy at the University of Waikato when we recoded the compulsory economics paper in our management degree from ECON100 to ECONS101 (see here). We were early adopters, as the CORE textbook was only released in 2017. It was a big change, and largely a positive one. The CORE textbook was free, substantially lowering the cost for students to access an important learning resource. Because the textbook was online, it could be constantly updated. And I really liked the way that it turned the traditional approach to the teaching of microeconomics on its head. Instead of starting with perfect competition and the supply and demand model, and then teaching imperfect competition as an exception, the CORE textbook started with monopolistic competition (where firms sell products that are differentiated from those of their competitors) and teaches perfect competition as an exception. Since many firms operate in monopolistically competitive markets, the approach that CORE adopted seems more attuned to the real world that students see.

I've often wondered whether the CORE textbook improved students' learning though. So, I was interested to read this recent article by Federico Araya (Universidad de la República, Uruguay) and co-authors, published in the journal Economica (sorry, I don't see an ungated version online [*]). They evaluate the impact of adopting the CORE textbook for the introductory economics course at the Faculty of Economic Sciences and Administration (FCEA) at the Universidad de la República, the largest university in Uruguay.

Although the CORE textbook was introduced at FCEA in 2020, Araya et al. start their analysis from 2021, to avoid the impacts on online teaching during the pandemic. FCEA offered two introductory microeconomics courses, one of which used CORE and the other continued to use their traditional textbook. Students were randomly assigned to either course based on the last number of their identification document, with 30 percent of students assigned to the course that used CORE.

However, the textbook was not the only difference between the two courses. As Araya et al. explain:

Although attendance is optional in both courses, in 2022 and 2023, the CORE course introduced a modification to its evaluation system, assigning 10% of the total grade to group activities conducted during class sessions. This change may have created an incentive for higher attendance.

So, their evaluation is not a clean comparison of the same course taught with two different textbooks, but will compare two different pedagogies, one which uses the CORE textbook and in-class group activities that are worth grade points, and one that uses a traditional textbook without the in-class group activities.

Araya et al. then compare the two groups in terms of whether students passed the introductory microeconomics course, as well as whether students passed an introductory calculus course and whether they passed the intermediate microeconomics course that follows on from the introductory course, while controlling for a range of demographic and socioeconomic variables for each student. They find:

...no statistically significant differences in pass rates between CORE and the conventional course, with the exception of the 2022 cohort.

I was initially surprised that they decided to evaluate each cohort separately, rather than pooling them. However, the 2021 cohort is different because that year the CORE course didn't have the in-class group activities, whereas it did for the 2022 and 2023 cohorts. Combining the 2022 and 2023 results would give us a better sense of the overall effect (of the combined CORE textbook plus group activities intervention). Instead, we are shown a statistically significant positive effect in 2022, but no statistically significant effect in 2023. That doesn't tell us whether the effects in those two years were actually different from each other, or whether the combined intervention had a positive overall effect across the two cohorts. One further problem here that muddies the comparison is that the CORE and traditional courses didn't use the same assessment, and so passing one course may be different from passing the other. And that might also explain the different cohort-specific results (if the 2022 traditional course had more difficult assessments than the CORE course, for example).

That won't be a problem for comparisons in terms of student performance in introductory calculus and intermediate microeconomics. For those courses, Araya et al. also find no statistically significant effects on passing.

So, at least there is no evidence from this study that the CORE textbook (with or without in-class group activities) made students worse off. Although equally, there is no evidence that it made them better off either. That allows me to raise an issue that is general to much of the similar research on educational interventions (including my own research on the impact of AI tutors in ECONS101). We might expect to see no significant effect on student performance even from a successful intervention. That's because a successful intervention may make studying easier for students, freeing up time that they can then devote to other activities. That might be studying for their other courses (although notice that in this case any reallocation of study effort doesn't appear to have affected the probability of passing introductory calculus), or something entirely different (maybe working more, or having more leisure time). So, I'm not surprised to see no effect of CORE on student performance in this study.

We continue to use the CORE textbook in my ECONS101 class (although last year we moved to the new edition, The Economy 2.0). We don't follow the text very closely, at least not in the microeconomics section of the paper that I teach. Nevertheless, it continues to provide the base material for a lot of what we teach. And it's good to know that at least one study says that there is not evidence that continuing to use a 'non-traditional' text is doing harm to students.

*****

[*] It's kind of ironic that a paper evaluating the impact of an open-access teaching resource is not itself published open-access.

Read more:

Saturday, 8 August 2026

How important is the apprenticeship model to entering a research career?

Those of us working in research careers can invariably share stories about working as a research assistant, cleaning datasets, coding, running models, doing literature reviews, and many of the other less-glamorous tasks that make up the research process. That was a key aspect of our apprenticeship into the world of research, and a gateway into a research career. But how important really is research assistance as an entry point into a research career?

That is the question addressed in this 2025 NBER Working Paper (ungated version here) by Ina Ganguli (University of Massachusetts, Amherst) and Raviv Murciano-Goroff (Boston University). They look at the impact of working in a university lab (an important subset of research assistance work) on subsequently pursuing a scientific career. Interestingly, Ganguli and Murciano-Goroff use changes in the local minimum wage as an exogenous source of variation in lab employment, so this paper also indirectly contributes to the literature on the employment effects of the minimum wage, in a context (research labs) that is not often the focus of that literature.

Their data comes from UMETRICS, which collates data on research grants across universities, and their dataset covers 32 universities over the period from 2000 to 2019 (although the dataset has coverage up to 2022, including those additional years would mean having to account for the COVID-19 pandemic).

First, using a dataset collated at the lab level, Ganguli and Murciano-Goroff use a staggered difference-in-differences approach to look at the impact of minimum wage changes on employment of undergraduates in the labs. This analysis essentially compares the change in undergraduate employment between the time before and the time after an increase in the minimum wage, between university labs that were affected by the minimum wage increase and those that were not. In that analysis, they find that:

...following minimum wage increases, labs decrease the employment of undergraduates by 7.4% on average...

So, not dissimilar to the literature on minimum wage effects on employment, when focused on young people in exposed occupations. Ganguli and Murciano-Goroff then use the minimum wage change as an instrument for students' exposure to laboratory research as an undergraduate. The key assumption is that minimum wages while an undergraduate affect later scientific careers only through their effect on lab employment opportunities.

Using a dataset of over 28,000 undergraduates and their subsequent career paths, Ganguli and Murciano-Goroff find that:

...decreased exposure to scientific work translates into significantly lower rates of undergraduate research assistants pursuing doctoral degrees or working in the life sciences sector after graduation. We find that working one fewer quarter in a lab during an undergrad student’s college years translates into between a 7.0 and 10.3 percentage point decrease in the rate of enrolling in a doctoral-level program. Given our sample of 28,283 students, this implies that if all students had experienced a minimum wage increase, roughly 500 fewer undergraduates in our sample would have pursued these advanced degrees.

So, the results imply that one fewer quarter of lab experience as an undergraduate reduces enrolment in a doctoral programme by 7.0 to 10.3 percentage points. That is a fairly large effect and should make graduate research programmes take notice of the importance of undergraduate research assistance opportunities for the pipeline into graduate research.

So, working in a lab as an undergraduate is a key pathway towards a career in the life sciences, and when those opportunities are restricted, students are less likely to embark on such a career. These results won't be too surprising to those of us who have been through an apprenticeship as a researcher. It is likely that I would be doing something very different right now if I hadn't been pulled into a lot of research projects towards the end of my undergraduate studies. I almost certainly wouldn't have pursued a PhD, or become an academic. Research assistant jobs are more than just jobs. They provide mentoring, information, skills, networks, references, and a chance to try out being a researcher in a safe setting.

These results also highlight two other things for me. First, they suggest that minimum wage increases may have a negative impact on the career pipeline into science. I'm unsure that this negative impact of the minimum wage has been identified before. However, we should be cautious because Ganguli and Murciano-Goroff are primarily interested in using minimum wage changes to identify the effect of undergraduate research experience on later careers, rather than estimating the overall long-run consequences of minimum wage increases for the scientific workforce. Nevertheless, their results suggest that this may be an important unintended consequence of higher minimum wages, and one that would be worth further research.

Second, given that research assistants do a lot of the 'drudge work' in research, if generative AI is also able to do a lot of that work, that further suggests a negative impact on the career pipeline into science through the rise of generative AI. I know others have written on this before (see here), so the challenge here is not a new idea.

On the plus side, that suggests another method that could be used to further test for the impacts of undergraduate (and graduate) research assistance on future academic careers, beyond focusing on life sciences (as Ganguli and Murciano-Goroff do). By comparing academic fields that are more (or less) exposed to generative AI (especially in the early days of generative AI), we might be able to tease out how important research assistance is to future academic careers across many fields.

Research apprenticeship is important, and these results give us a clear indication of how important it can be. Cleaning data, running models, searching the literature, and doing all the other seemingly mundane tasks of a research assistant are not just cheap ways for senior researchers to get research done. They are also how the next generation of researchers learns what research is, discovers whether they enjoy doing it, and gets started on a research career. If these opportunities are reduced, whether through higher minimum wages or through substitution by generative AI, we may save on some of the drudge work today, but at the cost of having fewer researchers tomorrow.

[HT: Marginal Revolution]