Showing posts with label Artificial intelligence. Show all posts
Showing posts with label Artificial intelligence. Show all posts

Saturday, 18 July 2026

Will generative AI mean the end of rational ignorance?

In this Substack post back in March, Andy Hall made the case for generative AI to create 'political superintelligence':

The more I work with and study AI, the more I believe it can give every human being on the planet access to a sort of political superintelligence, if we shape it right. And that intelligence, in turn, can make governments smarter and more effective, representatives more faithful, and institutions more responsive than anything we’ve built in over 2,000 years of experimenting with democracy.

Hall's post is worth reading in its entirety, but I want to explore a related point - will generative AI mean the end of rational ignorance for voters? Rational ignorance is the idea that it may be better for voters to not know what decisions policymakers are making on their behalf. That's because it's costly (in terms of time and effort) for voters to keep track of how the decisions that policymakers (and politicians) make on their behalf will affect them (economists call those monitoring costs). The benefit that a voter would receive by becoming informed of what policymakers (and politicians) are doing is relatively small, because their ability to change an election (and therefore policy) is very small. When the monitoring costs are greater than the benefits of being better informed, then voters would be better off not paying the monitoring costs. That is, voters would be better off not paying attention to what the policymakers (and politicians) are doing - the voters would be better off remaining rationally ignorant. This theory of rational ignorance was introduced in the 1950s by the late economist Anthony Downs.

Where does generative AI fit into this? Generative AI could meaningfully lower the monitoring costs for voters, as it gives the opportunity for voters to ask for quick summaries of policy proposals that may affect them. This will be even more effective as generative AI understands more about users' preferences. Moreover, agentic AI offers voters even greater opportunity to investigate what policymakers (and politicians) are doing, at relatively low cost.

When the monitoring costs decrease, then the rationale for voters to remain rationally ignorant weakens. We might expect voters to become more engaged with what the government is doing on their behalf, and to be more active in engaging with government to make their preferences known. Or, at least, maybe voters will delegate these activities to their favourite agentic AI model.

There are, of course, some reasons for caution. Generative AI might reduce the cost of obtaining political information without reducing the cost of checking whether that information is accurate or unbiased. Moreover, an overly sycophantic generative AI that knows the voter's preferences might reinforce the voter's existing views rather than challenging them. So, perhaps generative AI simply moves the monitoring costs from monitoring the government to monitoring the generative AI?

Hall makes the point that political superintelligence has the potential to increase the quality of governance. If generative AI enables voters to become better informed at low cost, it could strengthen political accountability. Policymakers (and politicians) who know that voters can easily scrutinise their decisions may be less willing to act against voters’ interests, or may face greater consequences when they do.

We may not have political superintelligence yet, and large numbers of voters may still be rationally ignorant. However, it may not be long before we start to see some substantive changes in the political process, driven in part by the emergence of generative AI.

[HT: Marginal Revolution for the Andy Hall post]

Monday, 13 July 2026

Generative AI and grade inflation

If generative AI can produce assessed work that earns higher marks than a student would earn by themselves, widespread use of generative AI should increase measured grades even in the absence of equivalent learning gains. In other words, we might expect grade inflation to occur as a result of students using generative AI, and that grade inflation to not reflect improved learning. To what extent is this generative AI-induced grade inflation occurring? That is the question addressed by this recent working paper by Igor Chirikov (UC Berkeley). 

Specifically, Chirikov looks at the change in the grade distribution across 319 courses (and over 500,000 enrolments) at "a large, selective public research university in Texas", covering the period from 2018 to 2025. He follows the labour economics literature by measuring 'task exposure' to AI, using the share of required tasks in each course's Fall 2022 syllabus that involved writing or coding - areas in which generative AI is particularly capable and could be used as a substitute for students' own efforts. Using a difference-in-differences (DID) approach, Chirikov compares the difference in the share of A grades and GPA between years before and including 2022 (when ChatGPT was released) and more recent years up to 2025, between courses more or less exposed to generative AI. He finds that:

...grades rose substantially in high-exposure courses after 2022: the share of A grades increased by 13 percentage points (about 30% relative to the 2022 baseline) and GPA by 0.12 points, accompanied by compression of the grade distribution.

Looking at the shares of other grades, Chirikov finds that:

The share of A- grades fell by 4 percentage points, the share of B+ grades by 3 percentage points, and the shares of B and below show smaller and mostly insignificant changes.

So, courses that were more exposed to generative AI have seen a greater increase in the share of A grades and a greater decrease in the share of lower grades than courses less exposed to generative AI. Chirikov then turns to exploring the mechanism that underlies the change, by extending the DID approach to also compare courses that place greater or lesser assessment weight on homework tasks (in what is called a 'triple differences' approach). In that analysis, he finds that:

...above-median homework courses show an additional 16 percentage point increase in the share of A grades relative to below-median courses with the same level of AI exposure. The effect on GPA follows a similar pattern, with an additional 0.13 point increase in high-homework courses, though this estimate is less precisely estimated...

In above-median homework courses, the share of other grades declines significantly with AI exposure relative to below-median homework courses.

So, the observed pattern of grade inflation, with grades near the top of the distribution shifting towards As to a greater extent in courses that have greater homework weight, provides strong evidence consistent with generative AI contributing to the observed grade inflation, principally by substituting for student effort on unsupervised assessment tasks.

We should be cautious about over-interpreting the results from this study. It is based on results from a single university, and it measures course-level task exposure to generative AI, rather than students' actual use of generative AI. Nevertheless, the results are consistent with what we would expect, and are likely to hold in other contexts.

Given that, how should universities adapt to solve this issue? The obvious response is to move to more invigilated, in-person assessment. However, Chirikov cautions that:

Not all skills can be meaningfully evaluated under exam conditions: the ability to produce a well-researched essay, develop a software project, or conduct an empirical analysis requires sustained engagement that timed in-person assessments cannot capture. Restricting assessment to formats that are AI-proof risks measuring a narrower set of capabilities than the ones courses are designed to develop, potentially undermining the learning goals that graded work is meant to serve. A more promising direction is to redesign assessments so that AI use is either structurally constrained by the task or purposefully incorporated into it, for example, by requiring students to document their process, justify their choices, or demonstrate understanding through follow-up interaction.

In my view, the optimal response for universities is a combination of invigilated in-person assessment and authentic assessments in which AI use is permitted (or even encouraged) and evaluated. The appropriate balance depends on the specific learning outcomes for each course. However, as I have argued before, one approach is to scaffold students through their studies, with lower-level courses that rely more on developing core knowledge, relying less on generative AI and therefore using more secure assessment, while higher-level courses increasingly incorporate generative AI use explicitly into the assessment. This also requires scaffolding students through learning how best to apply generative AI at each level.

It almost goes without saying that we cannot simply continue to assess students as we always have done. Generative AI has broken key elements of the assessment toolkit that we previously used. This is not a new observation. However, evidence that generative AI may be accelerating grade inflation makes it even more imperative that assessment practices are updated to better reflect the availability of generative AI.

Read more:

Wednesday, 10 June 2026

Is it working from home, and not generative AI, that is harming the prospects of young workers?

There is growing evidence that the labour market for young workers is challenging. Graduates are finding it more difficult to get jobs after graduation. Several research papers have noted that generative AI may be to blame (see this post, for example), with one research paper referring to the changes in the labour market as seniority-biased technological change (see this post).

But the challenge with trying to attribute changes in the labour market to the rise of generative AI is that there are other contemporaneous changes affecting the labour market as well. One of those changes is the rise of working from home (as I noted in yesterday's post). Working from home may reduce the prospects for junior workers in part because it costs more to supervise and monitor them when they are working from home. Junior workers also benefit from on-the-job learning when they work with other people, and that on-the-job learning is less effective when they work from home. Combining those two effects, working from home reduces the incentive for employers to hire junior workers.

This new working paper by Peter Lambert (University of Warwick) and Yannick Schindler (Ellison Institute of Technology, Oxford) tries to disentangle the effects of generative AI and working from home on employment of younger workers. They use data from Revelio Labs that is made up of monthly matched employer-employee records collected from résumés (predominantly from LinkedIn) to construct a measure of the junior share of all new hires. They also use data from Lightcast on the near-universe of online job postings across thousands of online job sites and other websites. They use the Lightcast data to construct a measure of the share of job postings that require three or fewer years of experience. Their data from both sources covers the period from 2017 to 2025, and includes four countries: the US, the UK, Canada, and Australia.

Lambert and Schindler then use that data, along with measures of 'exposure to generative AI' and 'exposure to working from home' at the occupation level, in a difference-in-differences strategy. That means that they essentially compare the change in the share of junior job hires (or job postings) between occupations that are more or less exposed to generative AI (or working from home). Their main results are neatly summarised in Figure 3 from the paper:

Panel (a) shows that the junior share of new hires decreases significantly in jobs that are more exposed to working from home, from 2023 onwards (the black line). When they also control for exposure to generative AI (the red line), the effect of working from home barely changes. In contrast, Panel (b) shows that the junior share of new hires also decreases significantly in jobs that are more exposed to generative AI, from 2023 onwards (the black line). However, when they also control for exposure to working from home (the blue line), the effect of generative AI becomes much smaller and statistically insignificant. The results are similar for the share of job postings requiring three or fewer years' experience, as shown in Panels (c) and (d) of the figure.

The size of the effects are quite large too. A one-standard-deviation increase in exposure to working from home reduces the junior share of new hires by about two percentage points, and the share of job postings requiring three or fewer years' experience by 1.5 percentage points.

Lambert and Schindler conclude that, based on their results, working from home is a better predictor of the decline in junior hiring than generative AI. Given potential benefits of working from home, they are reluctant to recommend policies against working from home, instead noting that:

...micro-level adjustments may be required to help firms adapt their organizational practices, so as to enjoy the benefits of WFH [work from home] arrangements while simultaneously managing the development of early-career talent.

Seen alongside the negative mental health impacts of working from home (as noted in yesterday's post), this should give us further pause for thought. However, it is worth noting that even if working from home is a better predictor of reductions in junior hiring than generative AI within their model, that doesn't let generative AI off the hook entirely. Since both trends are happening at the same time, reducing working from home might not eliminate the negative impacts on junior hiring, but instead make generative AI appear more important as an explanation. Lambert and Schindler note early in their paper that it is often the same occupations (white-collar occupations) that are most exposed to both working from home and generative AI. Given that, perhaps Lambert and Schindler's recommendation for micro-level changes in organisational practice may be the best mitigation strategy available to us.

[HT: Marginal Revolution]

Read more:

Monday, 25 May 2026

Does the future of higher education look more like a mentoring pyramid scheme?

In response to my recent post about the future of higher education and one-on-one mentoring, one of my students from last year, Yunze, got in touch via email to offer a potential solution:

...I wonder whether it is possible to set a clear academic threshold within each discipline. If students who reach this threshold could mentor upper‑middle‑level students, while professors spend only a small amount of time supervising the overall direction, the system might become more sustainable. However, I suspect this could harm the interests of the top students, since they might otherwise use that time to further advance their own academic achievements, and If [sic] they fail to successfully train students with real research ability, it would likely damage both the university’s reputation and the professor’s own reputation.

You know, I think Yunze is right on the money here. Consider the problems I outlined in the earlier post: (1) the signalling value of education is falling due to generative AI; (2) a one-on-one mentoring approach may be a solution; but (3) one-on-one mentoring doesn't scale due to limited faculty time. If one-on-one mentoring is not conducted between faculty and students, but works more like a pyramid mentoring model, then this might actually work, not just for students, but for faculty and for universities as well.

So, let's think it through. But first, remember that the mentoring model I introduced in the earlier post is not simply a model of small classes, where senior students perform limited teaching roles, such as tutoring. This is a model of genuine mentoring, where the mentor encourages the mentee to become a builder, in the words of Auren Hoffman. A builder creates things, and it is the act of building, and the learning alongside that, which will be a durable signal to future employers. In relation to mentoring, I said in that post that mentors should do the following for their mentees:

Teach them to be builders. Encourage them to create things. Work with them and chart a path forward for their success.

If faculty provide one-on-one mentoring to a small number of senior students, then that makes better use of faculty time than them mentoring hundreds of first-year students. The senior students can then each mentor several second-year students, who in turn can then mentor several first-year students. [*] In this model, faculty time is targeted at the senior students, where the impact of faculty on student employment outcomes may be greatest.

Students benefit from helping junior colleagues to become builders, where the signalling value may remain even in the face of generative AI. Even better, mentoring provides student mentors with an opportunity to build - they may be able to point employers to the success of their mentees as an example of their building, talking also about what went wrong in the mentoring relationship, and what they learned from the experience.

In this mentoring pyramid model, universities retain a key role, but that role becomes very different. Universities essentially become a platform, connecting students with mentors - first-year students with second-year mentors, second-year students with senior student mentors, and senior students with faculty mentors. In the terms of my earlier post, the university runs their own OnlyStudents platform.

Of course, this platform role creates a new problem for universities. If mentoring works mainly as a way of matching students with mentors, then the market may not need eight OnlyStudents platforms in New Zealand, or thousands of OnlyStudents platforms worldwide. A small number of large platforms could have a big advantage in that case - more students attract more mentors, more mentors improve the quality of matching, and better matching attracts still more students. Those network effects could create a winner-take-all dynamic, in which universities would struggle to differentiate themselves simply by running their own mentoring platforms, and where a single surviving OnlyStudents platform might be the ultimate outcome. However, that conclusion depends on the strength of the network effects. If effective mentoring also depends on institutional trust, disciplinary reputation, local employer connections, pastoral care, or an in-person community, then universities may retain some defensible advantages. Geography alone probably won’t be enough, especially if online mentoring is close to being as effective as in-person mentoring, but local connections might still matter. So the question for universities is not just whether they can build their own OnlyStudents, but whether they can attach that platform to something that a larger, more generic OnlyStudents cannot easily replicate.

Universities may also retain a role in the initial and ongoing training of mentors. Since each student, and each faculty member, will need to be a mentor to one or more others lower down in the pyramid, they will need to understand how to mentor. That means universities would not simply be matching students with mentors. They would also need to train mentors, monitor the quality of mentoring relationships, and intervene when mentor-mentee relationships are not working well. Moreover, the adoption of a mentoring pyramid model is likely going to change who the most successful students (and faculty members) are. The top students do not necessarily make the best mentors (or the best tutors, as I have learnt across years of coordinating tutors in my first-year papers). Good mentoring requires a specific skill set, but it is those skills that may also demonstrate the quality of the student as a builder - a signal of high quality for employers.

A further point about the pyramid mentoring model is that it likely requires a strong filtering effect to be financially viable. Since each faculty member can only mentor a limited number of senior students, and each of those senior students can only mentor a limited number of second-year students, who in turn can only mentor a limited number of first-year students, each level of the pyramid probably needs to be somewhat wider than the levels above it. To achieve that, student progression needs a strong filter, limiting the number of students who progress from first-year to second-year, and from second-year to senior.

Let's consider some simple numerical examples that illustrate why filtering is needed. If each faculty member mentors five senior students, and each senior student mentors five second-year students, who each mentor five first-year students, then the pyramid contains 155 students per faculty member - five senior students, 25 second-year students, and 125 first-year students. A model where each faculty member’s salary is covered by fees from 155 students, setting aside any contribution to central university costs, seems likely to be financially viable to me. However, in this model only one-fifth of students could be allowed to progress each year. That means the model would also need some form of orderly exit for students who are filtered out - perhaps an exit qualification, or a pathway into a non-mentored track. The problem is that both options may provide negative signals about the student who is filtered out.

If all students were to progress, then that would require each student to mentor at most one student at the level below. Keeping five senior students mentored by faculty, then the pyramid would contain 15 students per faculty member - five senior students, five second-year students, and five first-year students. That system would be much less likely to cover the cost of faculty time. So, it's unlikely that the pyramid mentoring model would be viable to run without some form of filtering - perhaps not as extreme as only one-fifth of students progressing each year, but clearly not all students could progress every year.

So, to return to my conclusion from the previous post, the current mass higher education model still looks increasingly fragile, but perhaps one or a few universities might be able to navigate their way through. However, the survivors are likely to be first-movers or fast followers in developing a platform market strategy that leverages a pyramid mentoring model. This model is still going to cost students a lot, and the filtering effect would make higher education more elitist as well.

And thanks to Yunze for inspiring this post with his perceptive email comments.

*****

[*] For simplicity, I'm assuming a three-year higher education degree structure, as we have in New Zealand. For a four-year degree structure, you would of course need to add an additional level.

Read more:

Saturday, 23 May 2026

Does the future of higher education look more like one-on-one mentoring?

It almost seems a cliche to say that generative AI is both the greatest threat, and the greatest opportunity, for higher education. That doesn't make the statement any less true. And there are many commentators who are trying to work out what happens next for higher education. I think one of the best is Hollis Robbins, whose Substack Anecdotal Value is well worth subscribing to.

Robbins was recently interviewed by Jay Caspian Kang of The New Yorker (paywalled, but you can find an ungated version here) about the future of higher education. The whole interview is worth reading, but I want to highlight this bit in particular, where Robbins says:

I was in Austin, Texas, a couple of times in March with a bunch of twenty-five-year-old billionaires. This is what they’re looking at. Instead of having the credential from the institution, why not have the credential from the professor? If you have a Hollis Robbins education, what would that signal? What would that credential mean as opposed to a degree from a university? There was some conversation about what that would look like, and one guy at the end of the dinner said, “Instead of OnlyFans, it’s like OnlyProfessors.”

The correct analogy here would be that it would be OnlyStudents, not OnlyProfessors (since the name contains the audience, not the performer). However, Robbins makes a good point. Higher education is, in part, an exercise in signalling (as I've noted before here and here). In fact, Bryan Caplan argued in his book The Case Against Education (which I reviewed here) that one third of the benefit of higher education is signalling (the other two-thirds is made up of genuine learning, socialisation, and transferable skills). 

The diploma that a student receives at the end of their higher education journey is a signal to employers of the student's quality as a future employee. The signal is credible to the employer because it is costly for the student to obtain, and costly in such a way that low-quality students wouldn't attempt the signal. University reputation matters here, because it is an assurance of the second of those conditions - low-quality students wouldn't attempt the signal of a Harvard degree, firstly because they wouldn't gain admission to Harvard in the first place. So, part of the signal comes from getting into Harvard. Second, low-quality students wouldn't attempt the signal of a Harvard degree because passing courses at Harvard is hard (or, at least, harder than at many other universities). The quality of the education at Harvard has traditionally been higher than elsewhere.

In the interview, Robbins makes a further important point, which is that we've spent the last several decades making higher education a commodity. Students studying a degree in a particular subject learn the same things, often using the same teaching materials, the same textbook, and the same style of assessment, regardless of which university they go to. That means that the quality of signal that arises from the education itself has reduced over time, meaning that most of the value of the signal arises from the admission process. Once a student has been admitted to Harvard, their signal is in place, and the further signalling from their education is lower than for comparable students in years past.

Robbins argues that generative AI is accelerating and expanding this commodification of education. Since generative AI has access, through its training corpus, to a large store of human knowledge, to add value over and above generative AI a professor must be a true expert in a very specific subfield. For students, learning a subject from a true expert still retains value, because generative AI cannot as easily replicate the learning that would occur from the expert. At that point, the university becomes less important as a mediator of education.

In other words, Robbins is arguing that the university could be disintermediated, with individual professors issuing credentials instead of universities. Who needs a Harvard degree, when you could have a Hollis Robbins degree? And since Robbins is an expert (in African American sonnet tradition, she says), the signal retains high value. However, that is true only to the extent that employers find the Hollis Robbins degree a credible signal.

That brings me to this Substack post by Auren Hoffman. Hoffman also argues that generative AI has diminished the value of the higher education signal. However, Hoffman argues for a different solution, this time from the perspective of the graduating student:

you have to show you can add value. that is it. that is the only thing.

the test the smart hiring manager applies in 2026 is simple. can you learn something on your own? can you finish what you started? can you do what you said you would do? these were always the skills that mattered. the difference is they are now the ONLY skills that matter, because the credential stopped doing the screening.

if you have a few years of experience, your resume can show you have these skills. if you are a new grad, you have to show what you have created and built.

Hoffman is overstating things a little, as university degrees are unlikely to disappear overnight. However, his critique is still important, as is the implication he draws. Hoffman argues that graduates (or young people, generally, since the solution doesn't depend on a student attending a university or completing a degree) need to become builders:

the most valuable thing a 22 year old can do in 2026 is create something. an app. a screenplay. a side business. an internal tool you wrote for a club you were in. a dinner series. a script that automates something annoying. a website. a chrome extension. a sculpture. a dance party. a discord bot. a substack with 32 readers and a real point of view. anything that moved from idea to working.

will the thing make money? probably not. that is not the point.

the point is that you taught yourself something. you finished it. you can describe what you learned, what broke, what you fixed, why you made the calls you made. that story is the new resume.

every hiring manager would rather interview a 22 year old with a launched app and a github full of weird side projects than a 22 year old with a 3.9 GPA from a top 50 school. it is not close. when one candidate has tangible evidence of what they can ship and the other has a transcript, the transcript will lose every time.

According to Hoffman, the best signal for young people to be sending in future is that they can build. Our students will ultimately be more successful if they can show off their skills (both technical and transferable skills) by building something. That is a signal that is costly, and costly in such a way that low-quality students will not attempt it. The signal is credible, and for the most part it retains currency even in the face of generative AI. Indeed, if the student builds something while effectively leveraging generative AI, then the signalling value to employers may be even greater. The key distinction is between the student using generative AI as a tool and the student using it as a substitute for doing the work. A student who can explain what they built, what broke, what they learned, and why they made the choices they made, is still sending a costly signal. A student who simply lets generative AI build for them is not. Employers would likely see through the latter pretty quickly.

Hoffman's idea isn't exactly new. When I think about some of my best students over the past two decades, they tend to be those that built something either while they were studying, or immediately after. For some, this was their own small business or entrepreneurial activity. For others, it was writing and publishing a research paper. Those were challenging tasks that set them apart from other students - a clear signal of quality. What Hoffman is essentially saying is that, with the signal from higher education itself being removed, the only remaining signal of quality is the signal from being a builder.

Where does that leave higher education staff? I think we can combine Robbins's and Hoffman's ideas, and chart a path forward. We don't need to start an OnlyStudents, and issue our own degrees, but we do need to cultivate closer relationships with our best students. Teach them to be builders. Encourage them to create things. Work with them and chart a path forward for their success. In other words, be a mentor.

Universities are absolutely going to hate this. Mentoring is not an activity that can be offered at scale. For example, there are simply not enough hours in the day for me to individually mentor all of the 350-plus students in my first-year economics class. Nor is mentoring easy to timetable, measure, standardise, or reward under current academic workload models. Universities are built around papers, credit points, learning outcomes, assessment rubrics, and student evaluations, all of which can be offered at scale. One-on-one mentoring doesn't work so well in that system. For example, postgraduate supervision already sits awkwardly within the system, with it being unclear whether it counts as teaching, or research. Nevertheless, mentoring may be one of the few ways that higher education can continue to offer something that is both valuable and difficult for generative AI to replicate.

If generative AI significantly reduces the signalling value of university education, and if students increasingly use it to avoid genuine learning, then the current mass higher education model looks increasingly fragile. Moreover, what remains is going to cost students a lot more. If having a high-quality mentor who can encourage a student to build is the path to their future employment, then it may be worth it. After all, high-quality signals are costly. That aspect of education, at least, won't have changed.

[HT: Marginal Revolution, for both the Hollis Robbins interview and the Auren Hoffman post]

Read more:

Tuesday, 24 March 2026

Evidence that artificial intelligence is increasing the impact, but narrowing the scope, of research

There is growing evidence of positive impacts of generative artificial intelligence on productivity. This includes productivity in research (see this post, for example), including my own. However, some have questioned whether increasing research productivity comes at a cost of narrowing the scope of research.

So, I was interested to read this article by Qianyue Hao (Tsinghua University) and co-authors, published in the prestigious journal Nature (ungated earlier version here) late last year. They look at the impact of AI tools (not limited to generative AI) on the productivity of researchers and the quality of research. Specifically, they look at authors publishing in six representative fields: biology, medicine, chemistry, physics, materials science, and geology, across three 'eras': (1) the 'machine learning era ' (from 1980 to 2014), the 'deep learning era' (from 2015 to 2022), and the 'generative AI era' (from 2023 onwards). Hao et al. compare authors who publish 'AI augmented papers' with those who do not. An 'AI augmented paper' is one that uses methods such as:

...support vector machines and principal component analysis from the machine learning era, and convolutional neural networks and generative adversarial networks from the deep learning era. Large language models, which have emerged in recent years, also rank among the most frequently used methods...

Using a dataset that includes over 27 million papers with complete records that were published between 1980 and 2025, of which about 310,000 were 'AI augmented', Hao et al. find that:

...annual citations to AI papers are 98.70% higher than those to non-AI papers on average...

So, AI augmented research gathers more citations, which suggests that authors using AI in their research achieve greater impact. This is reinforced by evidence that AI augmented papers are published in higher quality journals (with Q1 journals being the highest ranked). Hao et al. report that:

...the proportion of AI papers in Q1 journals is 18.60% higher than that of non-AI papers in all journals; in Q2 journals, the AI proportion is 1.59% higher; whereas Q3 and Q4 journals hold a relatively lower proportion of papers with AI... These results indicate a heterogeneous distribution of AI-augmented papers across journals, with a higher prevalence in high-impact journals.

And AI appears to make authors more productive, as:

On average, researchers adopting AI annually publish 3.02 times more papers... and garner 4.84 times more citations... than those not adopting AI, with consistency.

All of these results seem to hold across all of the disciplines that Hao et al. consider. However, it is not all good news. Hao et al. use machine learning to create a measure of the 'breadth of scholarly attention'. Using that measure, they find that:

Compared with conventional research, AI research is associated with a 4.63% contracted median collective knowledge extent across science, which is consistent across all six disciplines... Moreover, when dividing these disciplines into more than two hundred sub-fields, the contraction of knowledge extent can be observed in more than 70% of them...

Of course, some of the differences here may be due to selection, as the types of researchers, and the types of research, involving AI use may be meaningfully different from those that don't. However, putting the selection issues aside, Hao et al. note that there is a tension between the individual researcher's incentive to produce a greater quantity of research that has higher impact, which would suggest greater use of AI, and the social incentive to produce a greater breadth of research.

So, the takeaway from this paper is that we need to consider researcher incentives, not just productivity. Specifically, this research suggests that the use of AI in research is leading to a 'prisoners' dilemma' outcome: each individual researcher acting in their own best interests (and using AI in their research) leads to an outcome that is worse for society overall (less breadth of research and more incremental gains).

Hao et al. conclude that:

The substantial academic benefits of AI use may be a driving force behind its accelerated rate of adoption; however, we also find unintended consequences from the increased prevalence of AI-augmented research. In all fields, AI-augmented research focuses on a narrower scope of scientific topics and reduces the scientific engagement of follow-on research, leading to more overlapping research work that slows the expansion of knowledge. Further, with a greater concentration of collective attention to the same AI papers, the adoption of AI seems to induce authors to converge on the same solutions to known problems rather than create new ones.

So, what is the solution here? Society probably wants research to be higher quality and have a broad scope. But individual researchers' incentives to use AI in their research appears inconsistent with that outcome. The traditional prisoners' dilemma is a repeated game (see here or here, for example), and the players of that game can avoid the worst outcome by cooperating. In this case, the researchers could cooperate by agreeing not to use AI in their research. The problem is that every researcher has an incentive to cheat on that agreement, since if they use AI, then that will be good for their career. This prisoners' dilemma is more difficult to ensure cooperation in than the traditional game, because there are not just two players who need to cooperate, but thousands (or millions). Ensuring cooperation in a prisoners' dilemma game with many players, each of whom is far better off cheating than cooperating, is almost impossible (which is why solving the problem of climate change is so difficult).

My own view is that the answer is not to keep AI out of research. That is not realistic, in the same way that it's not realistic to expect students not to use generative AI. The incentives need to be redesigned, but this will be no easy task. As long as universities, research funders, and publishers reward researchers for quantity, citations, and publication in top-ranked outlets, then we should expect more AI-augmented work, with a narrower scope than society might prefer. If we want AI to expand knowledge rather than simply accelerate competition within narrow foci, then we need institutions that also reward novelty, breadth, and the discovery of new questions. That is the economic challenge we must face up to.

[HT: Marginal Revolution]

Wednesday, 4 March 2026

This is not how generative AI should be used in research

I've been using ChatGPT Pro to help with drafting research papers this year, as I noted that I would do in this post from January. It has amped up my productivity a lot, allowing me to finish writing up two papers already, with a third on the way. These were papers where the analysis was already done, but it was the writing that was holding up the process. Having ChatGPT to help with the drafting seems to kickstart my writing, even though I have ended up extensively re-writing everything that ChatGPT produces. I find it a good disciplining tool as much as anything. Several colleagues have asked whether I am disclosing my generative AI use to journal editors when I submit. And I do. I have a standard 'generative AI use statement' that I include in my papers, that notes how it was used, and that I remain responsible for all of the content. You can see an example in this recent working paper.

However, not everyone is as careful with their generative AI use, or as transparent. Consider this example:

That is both infuriating and a sad indictment of the reviewing, editing, and publishing process, not least because, as on Reddit commenter noted, many authors see high-quality work rejected by journals, whereas a paper like this, with obvious flaws, has successfully been published. And it's not an isolated incident. This 2025 article by Artur Strzelecki (University of Economics in Katowice), published in the journal Learned Publishing (open access), catalogues over 1300 instances of likely unacknowledged and frankly stupid use of ChatGPT, up to September 2024.

Strzelecki's approach is to search for text strings that are almost certainly ChatGPT responses to a prompt asking it to generate text. The main example Strzelecki uses, which is in the title of the article, is "as of my last knowledge update". No human author is going to say that in a research paper. Similarly, "as an AI language model", "I don't have access to", and "certainly, here is" are highly indicative of ChatGPT use. There are circumstances where a human might use those phrases in a research paper, but it seems unlikely. Strzelecki screens out papers that mention ChatGPT, and manually checks each paper to ensure the text was not in some way legitimate, and that leaves 1362 articles.

How do these articles get published with this content intact? There are lots of stopping points where this could be caught and corrected (or prevented), but these articles have gotten through all of them. Strzelecki outlines the process. First, perhaps it is only one of the authors (and not all of them) that used ChatGPT. In which case, why didn't the other co-authors pick it up? Next, the paper is submitted to a journal, and often goes through a text review by the publisher. And then the editor or editors (including associate editors) looks at it, and decides whether it should be sent out for peer review. And then the peer reviewers (usually more than one, sometimes four or more) look at the paper in detail and provide comments. Then the editor receives the review reports and makes a decision. The paper may go through more than one round of review and editorial decision. And then, once accepted for publication, the article may be copy-edited. And at any of those stages, this text could be picked up. And yet, for over 1300 articles as of September 2024, the ChatGPT-generated text has not been picked up.

Strzelecki particularly focuses on 89 articles that have been published in journals indexed by Scopus or Web of Science, which should be the most credible journals. Of these:

...as many as 28 of them are in journals with Scopus percentile values of 90 and above. Two journals have a 99th percentile, indicating that they are the top journals in their field...

In total, 64 articles were found in journals considered to be in Q1, top quartile, recognized as the group of the best journals in their respective fields. Twenty-five articles are in the percentile range between 50 and 75, indicating that the journals in which these articles are found belong to Q2.

So, this phenomenon is not limited to low-ranked 'predatory' journals. In fact, looking at the list, there are several journals published by MDPI and Frontiers (for more on those publishers, see here). However, there are a whole lot published by Elsevier and Springer, publishers that we should expect much better of. Although, those are also publishers that publish a lot of journals, and a lot of articles, so perhaps that accounts for their higher numbers within the 89 articles that Strzelecki focuses on. Fortunately, I don't see any reputable journals in economics in the list, but I could be wrong.

Anyway, the takeaway is not so much that generative AI use is widespread in the write-up of research. It is that authors are using generative AI, not being transparent in their use of it, and that the quality control system by journals, even high-ranking journals, is terrible. Strzelecki makes a good point in the conclusion of his article that 89 out of over 2.5 million articles indexed in Scopus is only 0.000035% of the total indexed articles. However, this analysis is only picking up the really, really obvious cases. There will be far more use of generative AI that has not been adequately checked or acknowledged by authors, and not picked up in quality control.

I'm not against using generative AI in the write-up of research. Obviously, because I am doing the same thing. What needs to happen is that researchers need to be transparent and honest when they use generative AI, so that editors, reviewers, and the readers of research can see how it was used. That way, the users of research can evaluate for themselves whether they should believe, discount, or discard research depending on the ways and the extent of generative AI use. Without transparency, that important evaluation step is lost.

[HT: Artur Strzelecki]

Read more:

Wednesday, 11 February 2026

Did employers value an AI-related qualification in 2021?

Many universities are rapidly adapting to education in the age of generative AI by trying to develop AI skills in their students. There is an assumption that employers want graduates with AI skills across all disciplines, but is there evidence to support that? This recent discussion paper by Teo Firpo (Humboldt-Universität zu Berlin), Lukas Niemann (Tanso Technologies), and Anastasia Danilov (Humboldt-Universität zu Berlin) provides an early answer. I say it's an early answer because their data come from 2021, before the wave of generative AI innovation that became ubiquitous following the release of ChatGPT at the end of 2022. The research also focuses on AI-related qualifications, rather than the more general AI skills, but it's a start.

Firpo et al. conduct a correspondence experiment, where they:

...sent 1,185 applications to open vacancies identified on major UK online job platforms... including Indeed.co.uk, Monster.co.uk, and Reed.co.uk. We restrict applications to entry-level positions requiring at most one year of professional experience, and exclude postings that demand rare or highly specialized skills...

Each identified job posting is randomly assigned to one of two experimental conditions: a "treatment group", which receives a résumé that includes additional AI-related qualifications and a "control group", which receives an otherwise identical résumé without mentioning such qualifications.

Correspondence experiments are relatively common in the labour economics literature (see here, for example), and involve the researcher making job applications with CVs (and sometimes cover letters) that differ in known characteristics. In this case, the applications differed by whether the CV included an AI-related qualification or not. Firpo et al. then focus on differences in callback rates, and they differentiate between 'strict callbacks' (invitations to interview), and 'broad callbacks' (any positive employer response, including requests for further information). Comparing callback rates between CVs with and without AI-related qualifications, they find:

...no statistically significant difference between treatment and control groups for either outcome measure...

However, when they disaggregate their results by job function, they find that:

In both Marketing and Engineering, résumés listing AI-related qualifications receive higher callback rates compared to those in the control group. In Marketing, strict callback rates are 16.00% for AI résumés compared to 7.00% for the control group (p-value = 0.075...), while broad callback rates are 24.00% versus 12.00% (p-value = 0.043...). In Engineering, strict callback rates are 10.00% for AI résumés compared to 4.00% for the control group (p-value = 0.163...), while broad callback rates are 20.00% versus 8.00% (p-value = 0.024...).

For the other job functions (Finance, HR, IT, and Logistics) there was no statistically significant effect of AI qualifications on either measure of callback rates. Firpo et al. then estimate a regression model and show that:

...including AI-related qualifications increases the probability of receiving an interview invitation for marketing roles by approximately 9 percentage points and a broader callback by 12 percentage points. Similarly, the interaction between the treatment dummy and the Engineering job function dummy in the LPM models is positive and statistically significant, but only for broad callbacks. AI-related qualifications increase the probability of a broad callback by at least 11 percentage points...

The results from the econometric model are only weakly statistically significant, but they are fairly large in size. However, I wouldn't over-interpret them because of the multiple-comparison problem (around five percent of results would show up as statistically significant just by chance). At best, the evidence that employers valued AI-related qualifications in 2021 is pretty limited, based on this research.

Firpo et al. were worried that employers might not have noticed the AI qualifications in the CVs, so they conducted an online survey of over 700 professionals with hiring experience and domain knowledge, but that survey instead shows that the AI-related qualification was salient and a signal of greater technical skills, but lower social skills. These conflicting signals are interesting, and suggestive that employers are looking for both technical skills and social skills in entry-level applicants. Does this, alongside the earlier results for different job functions, imply that technical skills are weighted more heavily than social skills for Engineering and Marketing jobs? I could believe that for Engineering, but for Marketing I have my doubts, because interpersonal skills are likely to be important in Marketing. Again though, it's probably best not to over-interpret the results.

Firpo et al. conclude that:

...our findings challenge the assumption that AI-related qualifications unambiguously enhance employability in early-career recruitment. While such skills might be valued in abstract or strategic terms, they do not automatically translate into interview opportunities, at least not in the entry-level labor market in job functions such as HR, Finance, Marketing, Engineering, IT and Logistics.

Of course, these results need to be considered in the context of their time. In 2021, AI-related skills might not have been much in demand by employers. That is unlikely to hold true now, given that generative AI use has become so widespread. It would be interesting to see what a more up-to-date correspondence experiment would find.

[HT: Marginal Revolution]

Read more:

  • ChatGPT and the labour market
  • More on ChatGPT and the labour market
  • The impact of generative AI on contact centre work
  • Some good news for human accountants in the face of generative AI
  • Good news, bad news, and students' views about the impact of ChatGPT on their labour market outcomes
  • Swiss workers are worried about the risk of automation
  • How people use ChatGPT, for work and not
  • Generative AI and entry-level employment
  • Survey evidence on the labour market impacts of generative AI
  • Tuesday, 10 February 2026

    Who on earth has been using generative AI?

    Who are the world's generative AI users? That is the question addressed in this recent article by Yan Liu and He Wang (both World Bank), published in the journal World Development (ungated earlier version here). They use website traffic data from Semrush, alongside Google Trends data, to document worldwide generative AI use up to March 2024 (so, it's a bit dated now, as this is a fast-moving area, but it does provide an interesting snapshot up to that point). In particular, Liu and Wang focus on geographical heterogeneity in generative AI use (measured as visits to generative AI websites, predominantly, or in some of their analyses, entirely ChatGPT), and they explore how that relates to country-level differences in institutions, infrastructure, and other variables.

    Some of the results are fairly banal, such as the rapid increase in website traffic to AI chatbot websites, a corresponding decline in traffic to sites such as Google, and Stack Overflow, and that the users skew younger, more educated, and male. Those demographic differences will likely become less dramatic over time as user numbers increase. However, the geographic differences are important and could be more persistent. Liu and Wang show that:

    As of March 2024, the top five economies for ChatGPT traffic are the US, India, Brazil, the Philippines, and Indonesia. The US share of ChatGPT traffic dropped from 70 % to 25 % within one month of ChatGPT’s debut. Middle-income economies now contribute over 50 % of traffic, showing disproportionately high adoption of generative AI relative to their GDP, electricity consumption, and search engine traffic. Low-income economies, however, represent less than 1 % of global ChatGPT traffic.

    So, as of 2024, most generative AI use was in middle-income countries, but remember that those are also high-population countries (like India). Generative AI users are disproportionately from high-income countries once income and internet use (proxied by search engine traffic) are accounted for. Figure 12 in the paper illustrates this nicely, showing generative AI use, measured as visits per internet user:

    Notice that the darker-coloured countries, where a higher proportion of internet users used ChatGPT, are predominantly in North America, western Europe, and Australia and New Zealand. On that measure, Liu and Wang rank New Zealand 20th (compared with Singapore first, and Australia eighth). There are a few interesting outliers like Suriname (sixth) and Panama (17th), but the vast majority of the top twenty countries are high-income countries.

    What accounts for generative AI use at the country level? Using a cross-country panel regression model, Liu and Wang find that:

    Higher income levels, a higher share of youth population, bet-ter digital infrastructure, and stronger human capital are key predictors of higher generative AI uptake. Services’ share of GDP and English fluency are strongly associated with higher chatbot usage.

    Now, those results simply demonstrate correlation, and are not causal. And website traffic could be biased due to use of VPNs, etc., not to mention that it doesn't account very well for traffic from China or Russia (and Liu and Wang are very upfront about that limitation). Nevertheless, it does provide a bit more information about how countries with high generative AI use differ from those with low generative AI use. Generative AI has the potential to level the playing field somewhat for lower-productivity workers, and lower-income countries. However, that can only happen if lower-income countries access generative AI. And it appears as if, up to March 2024 at least, they are instead falling behind. As Liu and Wang conclude, any catch-up potential from generative AI:

    ...depends on further development as well as targeted policy interventions to improve digital infrastructure, language accessibility, and foundational skills.

    To be fair, that sounds like a general prescription for development policy in any case.

    Read more:

    Monday, 9 February 2026

    The promise of a personalised, AI-augmented textbook, and beyond

    In the 1980s, the educational psychologist Benjamin Bloom introduced the 'two-sigma problem' - that students who were tutored one-on-one using a mastery approach performed on average two standard-deviations (two-sigma) better than students educated in a more 'traditional' classroom setting. That research is often taken as a benchmark for how good an educational intervention might be (relative to a traditional classroom baseline). The problem, of course, is that one-on-one tutoring is not scalable. It simply isn't feasible for every student to have their own personal tutor. Until now.

    Generative AI makes it possible for every student to have a personalised tutor, available 24/7 to assist with their learning. As I noted in yesterday's post though, it becomes crucial how that AI tutor is set up, as it needs to ensure that students engage meaningfully in a way that promotes their own learning, rather than simply being a tool to 'cognitively offload' difficult learning tasks.

    One promising approach is to create customised generative AI tools, that are specifically designed to act as tutors or coaches, rather than simple 'answer-bots'. This new working paper by the LearnLM team at Google (and a long list of co-authors) provides one example. They describe an 'AI-augmented textbook', which they call the 'Learn Your Way' experience, which:

    ...provides the learner with a personalized and engaging learning experience, while also allowing them to choose from different modalities in order to enhance understanding.

    Basically, this initially involves taking some source material, which in their case is a textbook, but could just as easily be lecture slides, transcripts, and related materials from a class. It then personalises those materials to the interests of the students, adapting the examples and exercises to fit a context that the students find more engaging. For example, if the student is an avid football fan, they might see examples drawn from football. And if the student is into Labubu toys, they might see examples based on that.

    The working paper describes the approach, reports a pedagogical evaluation performed by experts, and finally reports on a randomised controlled trial (RCT) evaluating the impact of the approach on student learning. The experts rated the Learn Your Way experience across a range of criteria, and the results were highly positive. The only criterion where scores were notably low was for visual illustrations. That accords with my experience so far with AI tutors, which are not good at drawing economics graphs, in particular (and is an ongoing source of some frustration!).

    The RCT involved sixty high-school students in Chicago area schools, who studied this chapter on brain development of adolescents. Half of the students were assigned to Learn Your Way, and half to a standard digital PDF reader. As the LearnLM Team et al. explain:

    Participants then used the assigned tool to study the material. Learning time was set to a minimum of 20 minutes and a maximum of 40 minutes. After this time, each participant had 15 minutes to complete the Immediate Assessment via a Qualtrics link.

    They then did a further assessment three days later (a 'Retention Assessment'). In terms of the impact of Learn Your Way:

    The students who used Learn Your Way received higher scores than those who used the Digital Reader, in both the immediate (p = 0.03) and retention (p = 0.03) assessments.

    The difference in test outcomes was 77 percent vs. 68 percent in the Immediate Assessment, and 78 percent vs. 67 percent in the Retention Assessment. So, the AI-augmented textbook increased student learning and retention by about 10 percentage points in both immediate learning and in the short term (three days). Of course, this was just a single study with a relatively small sample size of 60 students in a single setting, but it does offer some promise for the approach.

    I really like this idea of dynamically adjusting content to suit students' interests, which is a topic I have published on before. However, using generative AI in this way allows material to be customised for every student, creating a far more personalised approach to learning than any teacher could offer. I doubt that even one-on-one tutoring could match the level of customisation that generative AI could offer.

    This paper has gotten me thinking about the possibilities for personalised learning. Over the years, I have seen graduate students with specific interests left disappointed by what we are able to offer in terms of empirical papers. For example, I can recall students highly interested in economic history, the economics of education, and health economics in recent years. Generative AI offers the opportunity to provide a much more tailored education to students who have specific interests.

    This year, I'll be teaching a graduate paper for the first time in about a decade. My aim is to allow students to tailor that paper to their interests, by embarking on a series of conversations about research papers based on their interests. The direction that leads will be almost entirely up to the student (although with some guidance from me, where needed). Students might adopt a narrow focus on a particular research method, a particular research question, or a particular field or sub-field of economics. Assisted by a custom generative AI tool, they can read and discuss papers, try out replication packages, and/or develop their own ideas. Their only limits will be how much time they want to put into it. Of course, some students will require more direction than others, but that is what our in-class discussion time will be for.

    I am excited by the prospects of this approach, and while it will be a radical change to how our graduate papers have been taught in the past, it might offer a window to the future. And best of all, I have received the blessing of my Head of School to go ahead with this as a pilot project that might be an exemplar for wider rollout across other papers. Anyway, I look forward to sharing more on that later (as I will turn it into a research project, of course!).

    The ultimate question is whether we can use generative AI in a way that moves us closer to Bloom’s two-sigma benefit of one-on-one tutoring. The trick will be designing it so that students still do the cognitive work. My hope (and, it seems, the LearnLM team’s) is that personalisation increases students' engagement with learning rather than replacing it. If it works, this approach could be both effective and scalable in a way that human one-on-one tutoring simply can’t match.

    [HT: Marginal Revolution, for the AI-augmented textbook paper]

    Sunday, 8 February 2026

    Neuroscientific insights into learning and pedagogy, especially in the age of generative AI

    In May last year, my university's Centre for Tertiary Teaching and Learning organised a seminar by Barbara Oakley of Oakland University, with the grand title 'The Science of Learning'. It was a fascinating seminar about the neuroscience of learning, and in my mind, it justified several of my teaching and learning practices, such as continuing to have lectures, to emphasise students' learning basic knowledge in economics, and retrieval practice and spaced repetition as learning tools.

    Now, I've finally read the associated working paper by Oakley and co-authors (apparently forthcoming as a book chapter), and I've been able to pull out further insights that I want to share here. The core of their argument is in the Introduction to the paper. First:

    Emerging research on learning and memory reveals that relying heavily on external aids can hinder deep understanding. Equally problematic, however, are the pedagogical approaches used in tandem with reliance on external aids—that is, constructivist, often coupled with student-centered approaches where the student is expected to discover the insights to be learned... The familiar platitude advises teachers to be a guide on the side rather than a sage on the stage, but this oversimplifies reality: explicit teaching—clear, structured explanations and thoughtfully guided practice—is often essential to make progress in difficult subjects. Sometimes the sage on the stage is invaluable.

    I have resisted the urge to move away from lectures as a pedagogical tool, although I'd like to think that my lectures are more than simply information dissemination. I actively incorporate opportunities for students to have their first attempts at integrating and applying the economic concepts and models they are learning - the first step in an explicit retrieval practice approach. Oakley et al. note the importance of both components, because:

    ...mastering culturally important academic subjects—such as reading, mathematics, or science (biologically secondary knowledge)—generally requires deliberate instruction... Our brains simply aren’t wired to effortlessly internalize this kind of secondary knowledge—in other words, formally taught academic skills and content—without deliberate practice and repeated retrieval.

    The paper goes into some detail about the neuroscience underlying this approach, but again it is summarised in the Introduction:

    At the heart of effective learning are our brain's dual memory systems: one for explicit facts and concepts we consciously recall (declarative memory), and another for skills and routines that become second nature (procedural memory). Building genuine expertise often involves moving knowledge from the declarative system to the procedural system—practicing a fact or skill until it embeds deeply in the subconscious circuits that support intuition and fluent thinking...

    Internalized networks form mental structures called schemata, (the plural of “schema”) which organize knowledge and facilitate complex thinking... Schemata gradually develop through active engagement and practice, with each recall strengthening these mental frameworks. Metaphors can enrich schemata by linking unfamiliar concepts to familiar experiences... However, excessive reliance on external memory aids can prevent this process. Constantly looking things up instead of internalizing them results in shallow schemata, limiting deep understanding and cross-domain thinking.

    This last point, about the shallowness of learning when students rely on 'looking things up' instead of relying on their own memory of key facts (and concepts and models, in the case of economics), leads explicitly to worries about learning in the context of generative AI. When students rely on external aids (known as 'cognitive offloading'), then learning becomes shallow, because:

    ...deep learning is a matter of training the brain as much as informing the brain. If we neglect that training by continually outsourcing, we risk shallow competence.

    Even worse, there is a feedback loop embedded in learning, which exacerbates the negative effects of cognitive offloading:

    Without internally stored knowledge, our brain's natural learning mechanisms remain largely unused. Every effective learning technique—whether retrieval practice, spaced repetition, or deliberate practice—works precisely because it engages this prediction-error system. When we outsource memory to devices rather than building internal knowledge, we're not just changing where information is stored; we're bypassing the very neural mechanisms that evolved to help us learn.

    In short, internalized knowledge creates the mental frameworks our brains need to spot mistakes quickly and learn from them effectively. These error signals do double-duty: they not only help us correct mistakes but also train our attention toward what's important in different contexts, helping build the schemata we need for quick thinking. Each prediction error, each moment of surprise, thus becomes an opportunity for cognitive growth—but only if our minds are equipped with clear expectations formed through practice and memorization...

    Learning works through making mistakes, recognising those mistakes, and adapting to reduce those mistakes in future. Ironically, this is analogous to how generative AI models are trained (through 'reinforcement learning'). When students offload learning tasks to generative AI, they don't get an opportunity to develop the underlying internalised knowledge that allows them to recognise mistakes and learn from them. Thus, it is important for significant components of student learning to happen without resorting to generative AI (or other tools that allow students to cognitively offload tasks).

    Now, in order to encourage learning, teachers must provide students with the opportunity to make, and learn from, mistakes. Oakley et al. note that:

    ...cognitive scientists refer to challenges that feel difficult in the moment but facilitate deeper, lasting understanding as “desirable difficulties... Unlike deliberate practice, which systematically targets specific skills through structured feedback, desirable difficulties leverage cognitive struggle to deepen comprehension and enhance retention...

    Learning is not supposed to be easy. It is supposed to require effort. This is a point that I have made in many discussions with students. When they find a paper relatively easy, it is likely that they aren't learning much. And tools that make learning easier can hinder, rather than help, the learning process. In this context, generative AI becomes potentially problematic for learning for some (but not all) students. Oakley et al. note that:

    Individuals with well-developed internal schemas—often those educated before AI became ubiquitous—can use these tools effectively. Their solid knowledge base allows them to evaluate AI output critically, refine prompts, integrate suggestions meaningfully, and detect inaccuracies. For these users, AI acts as a cognitive amplifier, extending their capabilities.

    In contrast, learners still building foundational knowledge face a significant risk: mistaking AI fluency for their own. Without a robust internal framework for comparison, they may readily accept plausible-sounding output without realizing what’s missing or incorrect. This bypasses the mental effort—retrieval, error detection, integration—that neuroscience shows is essential for forming lasting memory engrams and flexible schemas. The result is a false sense of understanding: the learner feels accomplished, but the underlying cognitive work hasn’t been done.

    The group that benefits from AI as a complement for studying is not just those who were educated before AI became ubiquitous, but also those who learn in an environment where generative AI is explicitly available as a complement to learning (rather than a substitute). To a large extent, it depends on how generative AI is used as a learning tool. Oakley et al. do provide some good examples (and I have linked to some in past blog posts). I'd also like to think the AI tutors I have created for my ECONS101 and ECONS102 students assist with, rather than hamper, learning (and I have some empirical evidence that seems to support this, which I have already promised to blog about in the future).

    Oakley et al. conclude that:

    Effective education should balance the use of external tools with opportunities for students to internalize key knowledge and develop rich, interconnected schemata. This balance ensures that technology enhances learning rather than creating dependence and cognitive weakness.

    Finally, they provide some evidence-based strategies for enhancing learning (bolding is mine):

    • Embrace desirable difficulty—within limits: Encourage learners to generate answers and grapple with problems before turning to help... In classroom practice, this means carefully calibrating when to provide guidance—not immediately offering solutions, but also not leaving students floundering with tasks far beyond their current capabilities...
    • Assign foundational knowledge for memorization and practice: Rather than viewing factual knowledge as rote trivia, recognize it as the glue for higher-level thinking...
    • Use procedural training to build intuition: Allocate class time for practicing skills without external aids. For instance, mental math exercises, handwriting notes, reciting important passages or proofs from memory, and so on. Such practices, once considered old-fashioned, actually cultivate the procedural fluency that frees the mind for deeper insight...
    • Intentionally integrate technology as a supplement, not a substitute: When using AI tutors or search tools, structure their use so that the student remains cognitively active...
    • Promote internal knowledge structures: Help students build robust mental frameworks by ensuring connections happen inside their brains, not just on paper... guide students to identify relationships between concepts through active questioning ("How does this principle relate to what we learned last week?") and guided reflection...
    • Educate about metacognition and the illusion of knowledge: Help students recognize that knowing where to find information is fundamentally different from truly knowing it. Information that exists "out there" doesn't automatically translate to knowledge we can access and apply when needed.

    I really like those strategies as a prescription for learning. However, I am understandably biased, because many of the things I currently do in my day-to-day teaching practice are encompassed within (or similar to) those suggested strategies. I'll work on making 'guided reflection' a little more interactive in my classes this year, as I have traditionally made the links explicit for the students, rather than inviting them to make those links for themselves. We have been getting our ECONS101 students to reflect more on learning, and we'll be revising that activity (which happens in the first tutorial) this year to embrace more of a focus on metacognition.

    Learning is something that happens (often) in the brain. It should be no surprise that neuroscience has some insights to share on learning, and what that means for pedagogical practice. Oakley et al. take aim at some of the big names in educational theory (including Bloom, Dewey, Piaget, and Vygotsky), so I expect that their work is not going to be accepted by everyone. However, I personally found a lot to vindicate my pedagogical approach, which has developed over two decades of observational and experimental practice. I also learned that there are neuroscientific foundations for many aspects of my approach. And, I learned that there are things I can do to potentially further improve student learning in my classes.

    Wednesday, 14 January 2026

    David Deming on generative AI and commitment to learning, and the impact of generative AI on signalling in education

    When I was writing yesterday's post on generative AI and the economics major, I really wished I had read this post by David Deming on generative AI and learning, and then I could have linked the two together. Instead, I'll use this post to draw on Deming's ideas and flesh out why I think that generative AI makes signalling in education harder, and why that is a problem (in contrast with Matthew Kahn, who as noted in yesterday's post thinks that generative AI reduces problems of information asymmetry).

    First, Deming writes about the tension in education between students' desire to learn, and their desire to make life easier (the 'divided self', drawing on the example of Odysseus:

    A vivid illustration of our divided self comes from a famous behavioral economics paper called “Tying Odysseus to the Mast: Evidence from a Commitment Savings Product in the Philippines”. They found that customers flocked to and greatly benefited from a bank product that prevented them from accessing their own savings in the future. Just like when Odysseus tied himself to the mast of his ship so that he would not be tempted by the alluring song of the Sirens...

    The Sirens offer Odysseus the promise of unlimited knowledge and wisdom without effort. He survives not by resisting his curiosity, but by restricting its scope and constraining his own ability to operate. The Sirens possess all the knowledge that Odysseus seeks, but he realizes he must earn it. There are no shortcuts. This is the perfect metaphor for learning in the age of superintelligence.

    The analogy to generative AI is obvious. Generative AI is a tool that offers unlimited knowledge without effort, but using that tool means that the effort necessary for genuine learning is not expended. As Deming concludes:

    Learning is hard work. And there is now lots of evidence that people will offload it if given the chance, even if it isn’t in their long-run interest. After nearly two decades of teaching, I’ve realized that my classroom is more than just a place where knowledge is transmitted. It’s also a community where we tie ourselves to the mast together to overcome the suffering of learning hard things.

    How does this relate to the quality of signalling? It is worth reviewing the role of signalling in education, as I discussed in this post:

    On the other hand, education provides a signal to employers about the quality of the job applicant. Signalling is necessary because there is an adverse selection problem in the labour market. Job applicants know whether they are high quality or not, but employers do not know. The 'quality' of a job applicant is private information. High-quality (intelligent, hard-working, etc.) job applicants want to reveal to employers that they are hard-working. To do this, they need a signal - a way of credibly revealing their quality to prospective employers.

    In order for a signal to be effective, it must be costly (otherwise everyone, even those who are lower quality job applicants, would provide the signal), and it must be costly in a way that makes it unattractive for the lower quality job applicants to attempt (such as being more costly for them to engage in).

    Qualifications (degrees, diplomas, etc.) provide an effective signal (they are costly, and more costly for lower quality applicants who may have to attempt papers multiple times in order to pass, or work much harder in order to pass). So by engaging in university-level study, students are providing a signal of their quality to future employers. The qualification signals to the employer that the student is high quality, since a low-quality applicant wouldn't have put in the hard work required to get the qualification.

    What does generative AI like ChatGPT do to this signalling? When students can outsource much of the effort required to complete assessments, then not-so-good students no longer need to spend more time or effort to complete their qualification than do good students. Take-home assignments, essays, or written reports might be completed to a passing standard with little effort from the student at all. Completing a qualification is no longer costly in a way that makes it unattractive for lower quality job applicants to attempt. That means that employers would no longer be able to infer a job applicant's quality from whether they completed a qualification or not.

    A solution suggested by Deming's post is for students to find some way of committing themselves to not using generative AI in assessment. For this to solve the signalling problem, the commitment has to be credible (believable), such as being verifiable by potential employers later. While students could commit themselves to not using generative AI, and maintaining effortful learning, it is difficult to see how students who do so could credibly reveal that they have done so. They require some way of ensuring that potential employers could verify that the student didn't use generative AI. This is where universities could step in. If universities can certify that particular qualifications were 'AI-resistant', such as where assessment includes substantial supervised, in-person components (for example, tests or examinations), then that would help maintain the quality of the education signal. There are other options of course, including oral examinations, group or individual presentations, or supervised practice assessments that make learning harder to fake. However, anything that falls short of being AI-resistant in the eyes of employers is unlikely to work. However, limiting assessment styles in order to certify effortful learning doesn't come without a trade-off. AI-resistant assessment is likely to be less accessible, less flexible, less authentic, and potentially more likely to promote anxiety in students.

    Kahn suggested in his post that "AI-proctored assessments and virtual tutors suddenly make effort and mastery visible in real time". That could work. However, AI proctoring by itself is not a solution. In order to retain its status as a signal of quality for students, assessments need to require more effort to complete well for not-so-good students than for good students. Having an assessment where an AI proctors while a student uses a generative AI avatar to make an AI-generated presentation is not going to work. I'm sure that's not what Kahn was envisaging. Proctoring of online assessment (either by humans or by AI) is not as easy as it sounds. Last year I was part of a group tasked with evaluating online proctoring tools, to be rolled out for our new graduate medical school, and I was left thoroughly underwhelmed. All of the tools that we evaluated seemed to have simple workarounds that moderately tech-savvy students could easily employ. The solution that was offered (when the demonstrators could even offer a solution) was to have students complete assessments on-site, which more or less defeats the purpose of online proctoring.

    Anyway, the point is that generative AI reduces the signalling value of education. There are solutions where that signalling value can be retained, but that requires students to commit to effortful learning, and universities to certify that effort in a way that students who don’t expend it cannot mimic.

    [HT: Marginal Revolution]

    Read more:

    Tuesday, 13 January 2026

    Matthew Kahn on generative AI and the economics major

    There doesn't appear to be much of a consensus on how to adapt higher education to generative AI. I have my own thoughts, which I have shared here several times already (see the links at the end of this post). However, I am open to the ideas of others. So, I was interested to read this new paper by Matthew Kahn (University of Southern California), where he discusses his views on the future of the economics major. Specifically:

    I present an optimistic outlook on the evolution of our economics major over the coming decade, centered on the possibility of highly tailored, student-specific training that fully acknowledges the rich diversity of our students’ abilities, interests, and educational goals.

    Kahn is correct in laying out the challenge that we face:

    Faculty now face a steeper challenge in helping students see the value of investing sustained effort in a demanding subject like economics, especially when AI tools can produce quick answers and when attention is pulled in countless directions by social media, short-form video, gaming, and other digital platforms...

    If students are not prepared for rigorous material, then the easy path for them to follow is to rely on the AI as a crutch. AI creates a moral hazard effect. In recent years, I have stopped assigning class papers because it was obvious to me that the well written papers were being written by the AI. Each economics professor faces the challenge of how to use the incentives we control to nudge students to make AI a complement (not a substitute) for their own time investment in their studies.

    The challenge of making AI a complement rather than a substitute for learning has been a common theme in my writing on generative AI in education. Kahn's proposed solutions are not dissimilar from mine too. For instance, in introductory economics:

    Large language models can now go much further, acting as tireless, patient coaches that deliver truly adaptive “batting practice.” The AI begins with simple exercises and progressively escalates in difficulty, adjusting in real time to the student’s performance. This is exactly the repetitive, low-stakes practice every introductory economics student needs to build intuition. Going forward, I expect that we will see a growing number of economics educators introducing specialized AI economics tools...

    And that is exactly what I have done in my ECONS101 and ECONS102 classes this year. Both classes had AI tutors that were pre-trained with a knowledge base of the lecture material, and students could chat with the tutors, ask them questions, develop study guides, practice multiple choice questions, and probably a dozen other use cases I haven't considered. The flexibility of these AI tutors, both for myself and for the students, made them a huge contributor to students' learning this year (at least, that's what students said in their course evaluations at the end of each paper).

    Unfortunately, Kahn's prescription for changes at higher levels of the economics major are much weaker. For instance, for intermediate microeconomics he advocates for making use of short skills videos, then:

    AI will help here. Students can take the written transcripts from these video presentations and feed these to AI and ask for more examples to make it more intuitive for them. Students can explain their logic to AI and allow the AI to patiently tutor them. Students can ask the robot to generate likely exam questions for them to practice on.

    That isn't much of an advance on what he advocates at the introductory level, because it is still simply content plus discussion with an AI tutor. I think there is much more potential value at the intermediate level of getting students to engage in more back-and-forth exploratory discussions with generative AI, and making those discussions a small part of the assessment. That works in theory-based courses (intermediate microeconomics) and econometrics. Kahn could have thought deeper here about the possibilities. However, for intermediate macroeconomics, I really like this suggestion:

    AI tools make it possible to immerse students in the real-time decisions faced by figures such as Ben Bernanke in 2008. What information was available at each moment? What nightmare scenarios kept policymakers awake? Interactive simulations can let students experience economic policymaking “on the fly,” combining partial scientific knowledge with radical uncertainty. Such exercises tend to be far more memorable and engaging than static diagrams.

    Some 'scripted' AI tools, built on top of ChatGPT (like my AI tutors are) would be wonderful tools for simulation. The AI could be instructed to maintain certain relationships through the simulation, introduce particular shocks, and help the students to evaluate different monetary and fiscal policy responses (or, evaluate the impact of fiscal policy changes). This would be a much more tailored approach than the simulation modelling that Brian Silverstone used when I studied intermediate macroeconomics some twenty years ago. Kahn also has great suggestions for field classes:

    Professors teaching field classes often assign a textbook. Such a textbook offers both the professor and the students a linear progression structure but this teaching approach can feel dated as the professor delegates the course structure to a stranger who does not have experience teaching at that specific university. Textbooks are not often updated and the material (such as specific box examples) can quickly feel dated. AI addresses this staleness challenge...

    In recent months, I have experimented with loading many interesting readings to a shared Google LM Notebook website and encouraging my students to ask the AI for summaries about these writings and to ask their own questions...

    This year, I'll be teaching graduate development economics, for the first time in about a decade, and Kahn has pre-empted almost exactly the approach I was intending to adopt, with students engaged in conversation with a generative AI model (I wasn't sure if I would use NotebookLM or ChatGPT for this purpose), then expanding on that conversation within class. I'm also considering the feasibility of getting students in that class to work with generative AI on a short research project - collating and analysing data to answer some particular research questions, or to replicate some specific study. The paper is in the B Trimester, so I still have time to flesh out the details.

    Kahn then goes on to discuss the impacts of generative AI on research assistant and teaching assistant opportunities. I think he is a bit too pessimistic though, since he concludes that human research assistants will only be useful for developing new (spatial) datasets. I think there are many more use cases for human research assistants still, and not just for data collection or data cleaning. Finally, Kahn addresses information asymmetry, noting that:

    For far too long, students have been choosing majors in the dark—picking “prestigious” fields without really knowing what the degree will do for them, while universities have been able to hide behind vague reputations and opaque classrooms. Parents write enormous checks with almost no idea what they’re buying, employers wonder if the diploma still means anything, and everyone quietly suspects a lot of the game is just expensive signaling.

    AI changes that. Cheap, frequent, AI-proctored assessments and virtual tutors suddenly make effort and mastery visible in real time. Professors discover whether students are actually learning the material. Parents can peek at meaningful progress dashboards instead of just getting billing statements. Employers can ask for verifiable records of real skills instead of trusting a transcript that could have been gamed.

    I'm not sold on AI proctoring as a solution. In fact, I worry that it will simply lead to an 'arms race' of student AI tools vs. faculty AI tools. The advent of AI avatars and agentic AI simply makes this even more likely across a wider range of assessment types. However, I do agree with Kahn that a lot of education is signalling to employers, and that generative AI is going to change the dynamics of education away from signalling. Kahn seems to think that is a good thing. I worry the opposite! Without signalling, it is difficult for good students to distinguish themselves, and that limits the value proposition of higher education. Kahn wants "verifiable records of real skills instead of... a transcript that could have been gamed". However, generative AI makes it much easier for students to game the record of real skills, rendering those records less reliable.

    There isn't a consensus on the best path forward. Kahn's paper is a work in progress, and he is inviting others to share their thoughts. I have offered a few of mine in this post, and I look forward to sharing more of my explorations of generative AI in teaching as we go through this year.

    [HT: Marginal Revolution]

    Read more: