Showing posts with label Economics of Education. Show all posts
Showing posts with label Economics of Education. Show all posts

Tuesday, 28 July 2026

Could student-run social media groups reduce university student dropout?

Universities are quite focused on student retention, and as I noted in this 2018 post, if we can identify at-risk students, perhaps we can help to find ways to ensure they succeed. However, around that time I had a summer research scholarship student looking into the reasons that students drop out, and it turned out that each student dropped out for quite different, and difficult to predict, reasons. I call this the Anna Karenina principle of dropout: 'all students who persist are alike; each student who drops out does so in their own way'.

What if there were a simpler way of reducing student dropout, that did not require universities to identify at-risk students in advance, but instead reduced the risk of dropout from the outset? That would seem to be an attractive proposition.

So, I was interested to read this 2021 article by Lucio Masserini (University of Pisa) and Matilde Bini (European University of Rome), published in the journal Socio-Economic Planning Sciences (ungated version here), which evaluates the impact of student-created social media groups, such as Facebook pages, on student dropout. Masserini and Bini use survey data from 1879 first-year students from a major university in Central Italy.

Why would joining social media groups reduce dropout? Masserini and Bini suggest that these groups may help students form social connections and feel greater 'belonging' within the university community, while also providing a way for students to share information about courses, assessments, and study materials.

The key challenge in the analysis is that students are not randomly assigned to join social media groups - they choose whether or not to do so. And students who join these groups may differ from those who don't in ways that would bias a simple comparison of the students who joined social media groups and those who didn't. For example, more engaged students, who are less likely to drop out, might also be more inclined to join university-related social media groups run by other students. Masserini and Bini deal with this using propensity score matching - which involves identifying 'control' students who didn't join a social media group but who are most similar to each 'treated' student who did join a social media group. Then, comparing their matched control and treated students deals with any observable differences between the students who did, and did not, join social media groups.

Masserini and Bini then report a range of results of the estimated impact of social media groups on dropout, based on different assumptions used to do the matching, and:

...with the exception of k=1 nearest-neighbour, all the estimates indicated that students joining groups or Facebook pages had, on average, a lower probability to dropout, compared with those who were not part of such groups. The results also showed that the extent of the difference between the treated and control groups was not negligible, as it varied from 0.081 to 0.113, depending on the matching algorithm.

So, the results suggest that joining student-run social media groups or Facebook pages reduces the probability of a student dropping out by between 8.1 and 11.3 percentage points. Now, I should note that I don't in general find propensity score matching to be terribly convincing as a way of dealing with selection bias. 

Now, I should note that I do not find propensity-score matching entirely convincing as a way of dealing with selection bias. Although matching can make the treatment and control groups similar on observed characteristics, there is still something that is different about the treated and control students that leads the treated students to choose to join social media groups and the control students to choose not to join. That something is an omitted variable in the propensity score matching approach, and it is unclear how big the omitted variable bias will be. If, for example, joiners are more motivated or feel more connected to university life, then some of the apparent effect of joining the group on the probability of dropping out may instead reflect those underlying differences. Masserini and Bini's results are robust across several matching methods and sensitivity checks, which is reassuring, but robustness checks cannot establish that there isn't some omitted variable bias in the matching.

Having said that, if we take these results at face value, then there may be some merit in having student-run social media groups that university students can join. We must bear in mind that these results come from a survey in 2016, and they may not have aged well. But social media groups still exist, and students still participate in them. It could be worth exploring whether these effects still hold, given that increasing student retention remains a key focus for universities.

Having said that, student-run social media groups would probably be a relatively inexpensive way for universities to reduce dropout. Now, these results come from students surveyed in 2016, and both social-media use and the university environment have changed considerably since then. Nevertheless, the basic idea remains plausible. Universities could support the creation of student-run groups and randomly encourage or 'nudge' some students to join, then compare their subsequent retention with that of students who were not encouraged. That experimental approach would provide more contemporary and causal evidence of whether the groups reduce dropout, rather than merely attracting students who were already less likely to drop out.

Read more:

Monday, 20 July 2026

How career stereotypes shape students' choice of major

What jobs do accounting majors get? How about psychology majors? Or economics majors? If you answered, respectively, 'accountant', 'psychologist', and 'economist', you're probably far from alone. When most people think about particular fields of study, they have stereotypical jobs in mind, and they're far more likely to believe that majors get the stereotypical job than any other job. Even in the case of economics, where very few graduates will go into a job with the title 'economist' (many will go into a job with some sort of 'analyst' title, like a business analyst, market analyst, or financial analyst, etc.).

A new article by John Conlon (Ohio State University) and Dev Patel (Brown University), published in the Quarterly Journal of Economics (open access) demonstrates the extent of this stereotyping. They also show using a simple survey experiment that students' stereotypical views can be changed, affecting their choice of major.

The first part of the paper compares students’ beliefs about the careers associated with different majors with actual major-career combinations in the 2017-2019 American Community Survey. Conlon and Patel then use the CIRP Freshman Survey, covering more than nine million first-year students between 1976 and 2015, to compare students’ expected careers with the occupations subsequently observed among graduates from the same cohorts. In that, they find:

...large, systematic, and persistent differences between the careers that freshmen expect to attain and the actual occupations they go on to have... We see that twice as many students expect to become artists, counselors, and lawyers (about 5% each) than actually do (2%–3% each). Four times as many students expect to become writers and doctors (2.7% and 11.1%) than do (0.7% and 2.8%).

Interestingly, the occupations that students most overestimate themselves as having are those that are rare and representative of particular majors (like writers or artists), while those that are most underestimated are those that are common alternative occupations that many majors may later hold (like teaching or business). Conlon and Patel also show using implicit association tests that people:

...strongly and systematically associate majors with their representative careers: implicit associations are 0.30–0.36 standard deviations higher for representative major–career pairs than for nonrepresentative pairs...

This supports Conlon and Patel’s interpretation that stereotyping contributes to students’ exaggerated beliefs about the connection between majors and the corresponding representative careers. Conlon and Patel then show using a theoretical model that stereotyping can increase misallocation of labour. What that means is that, given the occupation in which a graduate eventually works, that person might have been better off studying a different field. They also present suggestive evidence that links greater stereotyping with job dissatisfaction and regrets about the chosen field of study.

The welfare loss, job dissatisfaction, and regret, then motivated a survey experiment where Conlon and Patel attempt to correct for the stereotyping. The survey experiment proceeded as follows, using students from Ohio State University:

The survey began by asking students the percent chance that they would graduate with the two majors they selected as being most likely to pursue (their “top-ranked” and “second-ranked” majors). It then asked their self and population beliefs about the likelihood of each career group conditional on these two majors. Students were randomly sorted into a control group and a treatment group. Those in the control arm answered questions about their classes so far that semester and how they had (or had not) contributed to their major and career plans...

In the treatment arm, information modules provided students with the actual distribution of careers conditional on each of their top two majors according to data from the [American Community Survey]. For each major, it told them several headline numbers about the frequency of the careers they had listed as their most likely jobs if they graduated with that major...

Conlon and Patel then looked at whether assigning a student to the treatment group, where they received accurate information about the chance their top-ranked majors would lead to particular careers, affected students' beliefs about their own chances of working in each major's 'representative career'. They found that the information partially (but not fully) corrects students' misbeliefs about the chance of attaining the representative career, and that:

...reducing students’ self beliefs about their chance of having their top-ranked major’s representative job by 10 percentage points decreases intentions toward that major by 0.11 standard deviations (about 3.5 percentage points, p < .05).

So, when accurate information lowers students' beliefs about their chances of obtaining the representative career associated with their top-ranked major, their intentions towards that major, and their enrolment in that major, decline. Interestingly though, correcting beliefs about a student’s second-ranked major can make that alternative more attractive. The response therefore depends on both the information that students receive and how much they value the different careers.

That would seem to be good news for some university majors, where the 'representative occupations' are relatively common (e.g. accounting or marketing) and bad news for other majors where the 'representative occupations' are less common (e.g. journalism or film studies).

Where does economics fit in as a major? Sadly, Conlon and Patel don't answer that question, as they combine economics with accounting, finance, marketing, and other business fields. In that broad group, 73 percent of prospective majors expected to enter a business career, compared with 47 percent of graduates who actually did so. So economics may not escape stereotyping. It may simply be stereotyped as a route into 'business', rather than more narrowly as a route to a job called 'economist'. However, we don't know for sure from these results.

Would better information help prospective economics students to make better choices? On the one hand, it might deter students who see economics as a guaranteed path to one particular career, but it might attract others by showing that economics is not tied to a single occupation. Perhaps the problem for economics is not a lack of possible career destinations, but that its many potential destinations are less vivid than the single job title of 'economist'.

[HT: Marginal Revolution, back in 2022 when it was Conlon's job market paper]

Monday, 13 July 2026

Generative AI and grade inflation

If generative AI can produce assessed work that earns higher marks than a student would earn by themselves, widespread use of generative AI should increase measured grades even in the absence of equivalent learning gains. In other words, we might expect grade inflation to occur as a result of students using generative AI, and that grade inflation to not reflect improved learning. To what extent is this generative AI-induced grade inflation occurring? That is the question addressed by this recent working paper by Igor Chirikov (UC Berkeley). 

Specifically, Chirikov looks at the change in the grade distribution across 319 courses (and over 500,000 enrolments) at "a large, selective public research university in Texas", covering the period from 2018 to 2025. He follows the labour economics literature by measuring 'task exposure' to AI, using the share of required tasks in each course's Fall 2022 syllabus that involved writing or coding - areas in which generative AI is particularly capable and could be used as a substitute for students' own efforts. Using a difference-in-differences (DID) approach, Chirikov compares the difference in the share of A grades and GPA between years before and including 2022 (when ChatGPT was released) and more recent years up to 2025, between courses more or less exposed to generative AI. He finds that:

...grades rose substantially in high-exposure courses after 2022: the share of A grades increased by 13 percentage points (about 30% relative to the 2022 baseline) and GPA by 0.12 points, accompanied by compression of the grade distribution.

Looking at the shares of other grades, Chirikov finds that:

The share of A- grades fell by 4 percentage points, the share of B+ grades by 3 percentage points, and the shares of B and below show smaller and mostly insignificant changes.

So, courses that were more exposed to generative AI have seen a greater increase in the share of A grades and a greater decrease in the share of lower grades than courses less exposed to generative AI. Chirikov then turns to exploring the mechanism that underlies the change, by extending the DID approach to also compare courses that place greater or lesser assessment weight on homework tasks (in what is called a 'triple differences' approach). In that analysis, he finds that:

...above-median homework courses show an additional 16 percentage point increase in the share of A grades relative to below-median courses with the same level of AI exposure. The effect on GPA follows a similar pattern, with an additional 0.13 point increase in high-homework courses, though this estimate is less precisely estimated...

In above-median homework courses, the share of other grades declines significantly with AI exposure relative to below-median homework courses.

So, the observed pattern of grade inflation, with grades near the top of the distribution shifting towards As to a greater extent in courses that have greater homework weight, provides strong evidence consistent with generative AI contributing to the observed grade inflation, principally by substituting for student effort on unsupervised assessment tasks.

We should be cautious about over-interpreting the results from this study. It is based on results from a single university, and it measures course-level task exposure to generative AI, rather than students' actual use of generative AI. Nevertheless, the results are consistent with what we would expect, and are likely to hold in other contexts.

Given that, how should universities adapt to solve this issue? The obvious response is to move to more invigilated, in-person assessment. However, Chirikov cautions that:

Not all skills can be meaningfully evaluated under exam conditions: the ability to produce a well-researched essay, develop a software project, or conduct an empirical analysis requires sustained engagement that timed in-person assessments cannot capture. Restricting assessment to formats that are AI-proof risks measuring a narrower set of capabilities than the ones courses are designed to develop, potentially undermining the learning goals that graded work is meant to serve. A more promising direction is to redesign assessments so that AI use is either structurally constrained by the task or purposefully incorporated into it, for example, by requiring students to document their process, justify their choices, or demonstrate understanding through follow-up interaction.

In my view, the optimal response for universities is a combination of invigilated in-person assessment and authentic assessments in which AI use is permitted (or even encouraged) and evaluated. The appropriate balance depends on the specific learning outcomes for each course. However, as I have argued before, one approach is to scaffold students through their studies, with lower-level courses that rely more on developing core knowledge, relying less on generative AI and therefore using more secure assessment, while higher-level courses increasingly incorporate generative AI use explicitly into the assessment. This also requires scaffolding students through learning how best to apply generative AI at each level.

It almost goes without saying that we cannot simply continue to assess students as we always have done. Generative AI has broken key elements of the assessment toolkit that we previously used. This is not a new observation. However, evidence that generative AI may be accelerating grade inflation makes it even more imperative that assessment practices are updated to better reflect the availability of generative AI.

Read more:

Wednesday, 27 May 2026

Is it better to have a more educated mayor?

It seems somewhat self-evident that having a more educated mayor would be better than having a less educated mayor. However, whether education is a positive attribute for a mayor really depends on whether, and to what extent, more educated mayors act differently than less educated mayors. Do they spend more, or less? How do they spend the public budget?

This new article by Alessio Mitra (University of Kent), published in the European Journal of Political Economy (ungated earlier version here) directly addresses the second question - how does mayoral education affect public finance? Mitra uses data from municipal elections in Italy over the period from 2000 to 2015, focusing on municipalities with a population of less than 15,000 (because larger municipalities use different electoral rules). He defines a more educated mayoral candidate as one with a university degree, and a less educated mayoral candidate as one without a degree.

Mitra applies a regression discontinuity design (RDD), which involves comparing municipalities that narrowly elected a more educated mayoral candidate over a less educated candidate with similar municipalities where the more educated candidate lost to the less educated candidate. In very close elections, the identity of the winner is plausibly as-good-as random, provided there is no manipulation around the threshold related to the education of the candidates. In other words, since the difference between getting 50.01 percent of the vote and getting 49.99 percent of the vote is essentially random, the education of the winning mayoral candidate is basically determined randomly in these close elections between candidates with different education levels. With that assumption in mind, observed differences between the municipalities where a more educated candidate won with those where they lost can be attributed to the difference in mayoral education.

Mitra's dataset includes more than 18,000 mayoral elections, of which 1211 have a margin of victory of less than five percent (which he defines as a close election, and includes in the analysis). He looks at the differences in public expenditure, initially focusing on changes in the share of spending devoted to operational expenses (or 'current expenditure' as he terms it) or public investment. In this, Mitra finds that:

When an educated mayor is elected by chance, public investment rises by 3 percentage points of total expenditure compared to a less educated counterpart.

Digging down into the allocation of that public investment, he finds that:

...educated mayors allocate an additional 1 percentage point of total expenditure to education investment, accounting for one-third of the overall increase in public investment.

And going a bit deeper than that:

Among education investments, immovable assets dedicated to nurseries receive the largest increase in resources.

Consistent with Italy’s balanced budget requirement on municipalities, there is no significant change in fiscal deficit. That means that the additional spending devoted to public investment must mean a corresponding reduction in operational expenditure. Mitra doesn't really dig into that at all.

What we take away from this paper is that more educated mayors devote more spending to education. In the Waikato Economics Discussion Group today, we discussed what mechanisms might underlie this difference, which is something that Mitra didn't explore. Perhaps more educated mayors see more value in education. After all, they invested more in their own education than a less educated mayor did. However, that's not entirely consistent with spending more on public investment in early childhood education.

A second possibility is that more educated mayors have a lower intrinsic discount rate, increasing their willingness to make long-term investments, both in their own education and in the education of their citizens. This is more consistent with devoting spending to public investment in early childhood education.

A third possibility is that more educated mayors may be better at the administration of public investment, such as project approvals, capital budgeting, grants, or procurement. This means that they have greater capacity for public investment projects. However, that greater capacity wouldn't necessarily be more apparent for public investment in education, or early childhood investment.

However, an intriguing but speculative fourth possibility is that more educated mayors understand that public investment can be used strategically to affect demographics. Many municipalities in Italy are facing extreme population ageing and/or declining populations. Mitigating (but probably not reversing) those population changes may be possible through creative policy. If the municipality invests in early childhood education, that may make the municipality more attractive for young parents to relocate to, and may reduce cost pressures that hamper fertility. The problem with this as an explanation is that it isn't clear that these trends and policy solution would be more apparent to a more educated mayor than to a less educated one.

The second possibility seems to me like the most promising. However, exploring the reasons why more educated mayors spend more on public investment, particularly in education, is a promising exercise for future research.

One last point is that the effects are actually quite modest. The total budget for a municipality of 15,000 population would be around €15-25 million per year (based in part on this and this, both in Italian, but see also here for public finance data for all Italian municipalities). A reallocation of three percentage points to public investment represents up to an additional €750,000 per year. And if one-third of that is spent on public investment in education, that is an additional €250,000 per year. It's not nothing, but it's certainly not building multiple new schools. Maybe it's an additional small school building per year.

So, is it better to have a more educated mayor? This research suggests yes, but that relies on a normative view that more spending on public investment, particularly in education, is overall a good thing. However, the size of the effect doesn't suggest transformational change, and we don't really know what the trade-offs are in terms of what categories of operational spending were reduced. A university degree does not necessarily make someone a better mayor, and this paper cannot tell us whether more educated mayors have better preferences, longer time horizons, or simply greater administrative capacity. What it does show is that who gets elected can change not just how much is spent, but what kind of future a municipality chooses to invest in.

Monday, 25 May 2026

Does the future of higher education look more like a mentoring pyramid scheme?

In response to my recent post about the future of higher education and one-on-one mentoring, one of my students from last year, Yunze, got in touch via email to offer a potential solution:

...I wonder whether it is possible to set a clear academic threshold within each discipline. If students who reach this threshold could mentor upper‑middle‑level students, while professors spend only a small amount of time supervising the overall direction, the system might become more sustainable. However, I suspect this could harm the interests of the top students, since they might otherwise use that time to further advance their own academic achievements, and If [sic] they fail to successfully train students with real research ability, it would likely damage both the university’s reputation and the professor’s own reputation.

You know, I think Yunze is right on the money here. Consider the problems I outlined in the earlier post: (1) the signalling value of education is falling due to generative AI; (2) a one-on-one mentoring approach may be a solution; but (3) one-on-one mentoring doesn't scale due to limited faculty time. If one-on-one mentoring is not conducted between faculty and students, but works more like a pyramid mentoring model, then this might actually work, not just for students, but for faculty and for universities as well.

So, let's think it through. But first, remember that the mentoring model I introduced in the earlier post is not simply a model of small classes, where senior students perform limited teaching roles, such as tutoring. This is a model of genuine mentoring, where the mentor encourages the mentee to become a builder, in the words of Auren Hoffman. A builder creates things, and it is the act of building, and the learning alongside that, which will be a durable signal to future employers. In relation to mentoring, I said in that post that mentors should do the following for their mentees:

Teach them to be builders. Encourage them to create things. Work with them and chart a path forward for their success.

If faculty provide one-on-one mentoring to a small number of senior students, then that makes better use of faculty time than them mentoring hundreds of first-year students. The senior students can then each mentor several second-year students, who in turn can then mentor several first-year students. [*] In this model, faculty time is targeted at the senior students, where the impact of faculty on student employment outcomes may be greatest.

Students benefit from helping junior colleagues to become builders, where the signalling value may remain even in the face of generative AI. Even better, mentoring provides student mentors with an opportunity to build - they may be able to point employers to the success of their mentees as an example of their building, talking also about what went wrong in the mentoring relationship, and what they learned from the experience.

In this mentoring pyramid model, universities retain a key role, but that role becomes very different. Universities essentially become a platform, connecting students with mentors - first-year students with second-year mentors, second-year students with senior student mentors, and senior students with faculty mentors. In the terms of my earlier post, the university runs their own OnlyStudents platform.

Of course, this platform role creates a new problem for universities. If mentoring works mainly as a way of matching students with mentors, then the market may not need eight OnlyStudents platforms in New Zealand, or thousands of OnlyStudents platforms worldwide. A small number of large platforms could have a big advantage in that case - more students attract more mentors, more mentors improve the quality of matching, and better matching attracts still more students. Those network effects could create a winner-take-all dynamic, in which universities would struggle to differentiate themselves simply by running their own mentoring platforms, and where a single surviving OnlyStudents platform might be the ultimate outcome. However, that conclusion depends on the strength of the network effects. If effective mentoring also depends on institutional trust, disciplinary reputation, local employer connections, pastoral care, or an in-person community, then universities may retain some defensible advantages. Geography alone probably won’t be enough, especially if online mentoring is close to being as effective as in-person mentoring, but local connections might still matter. So the question for universities is not just whether they can build their own OnlyStudents, but whether they can attach that platform to something that a larger, more generic OnlyStudents cannot easily replicate.

Universities may also retain a role in the initial and ongoing training of mentors. Since each student, and each faculty member, will need to be a mentor to one or more others lower down in the pyramid, they will need to understand how to mentor. That means universities would not simply be matching students with mentors. They would also need to train mentors, monitor the quality of mentoring relationships, and intervene when mentor-mentee relationships are not working well. Moreover, the adoption of a mentoring pyramid model is likely going to change who the most successful students (and faculty members) are. The top students do not necessarily make the best mentors (or the best tutors, as I have learnt across years of coordinating tutors in my first-year papers). Good mentoring requires a specific skill set, but it is those skills that may also demonstrate the quality of the student as a builder - a signal of high quality for employers.

A further point about the pyramid mentoring model is that it likely requires a strong filtering effect to be financially viable. Since each faculty member can only mentor a limited number of senior students, and each of those senior students can only mentor a limited number of second-year students, who in turn can only mentor a limited number of first-year students, each level of the pyramid probably needs to be somewhat wider than the levels above it. To achieve that, student progression needs a strong filter, limiting the number of students who progress from first-year to second-year, and from second-year to senior.

Let's consider some simple numerical examples that illustrate why filtering is needed. If each faculty member mentors five senior students, and each senior student mentors five second-year students, who each mentor five first-year students, then the pyramid contains 155 students per faculty member - five senior students, 25 second-year students, and 125 first-year students. A model where each faculty member’s salary is covered by fees from 155 students, setting aside any contribution to central university costs, seems likely to be financially viable to me. However, in this model only one-fifth of students could be allowed to progress each year. That means the model would also need some form of orderly exit for students who are filtered out - perhaps an exit qualification, or a pathway into a non-mentored track. The problem is that both options may provide negative signals about the student who is filtered out.

If all students were to progress, then that would require each student to mentor at most one student at the level below. Keeping five senior students mentored by faculty, then the pyramid would contain 15 students per faculty member - five senior students, five second-year students, and five first-year students. That system would be much less likely to cover the cost of faculty time. So, it's unlikely that the pyramid mentoring model would be viable to run without some form of filtering - perhaps not as extreme as only one-fifth of students progressing each year, but clearly not all students could progress every year.

So, to return to my conclusion from the previous post, the current mass higher education model still looks increasingly fragile, but perhaps one or a few universities might be able to navigate their way through. However, the survivors are likely to be first-movers or fast followers in developing a platform market strategy that leverages a pyramid mentoring model. This model is still going to cost students a lot, and the filtering effect would make higher education more elitist as well.

And thanks to Yunze for inspiring this post with his perceptive email comments.

*****

[*] For simplicity, I'm assuming a three-year higher education degree structure, as we have in New Zealand. For a four-year degree structure, you would of course need to add an additional level.

Read more:

Saturday, 23 May 2026

Does the future of higher education look more like one-on-one mentoring?

It almost seems a cliche to say that generative AI is both the greatest threat, and the greatest opportunity, for higher education. That doesn't make the statement any less true. And there are many commentators who are trying to work out what happens next for higher education. I think one of the best is Hollis Robbins, whose Substack Anecdotal Value is well worth subscribing to.

Robbins was recently interviewed by Jay Caspian Kang of The New Yorker (paywalled, but you can find an ungated version here) about the future of higher education. The whole interview is worth reading, but I want to highlight this bit in particular, where Robbins says:

I was in Austin, Texas, a couple of times in March with a bunch of twenty-five-year-old billionaires. This is what they’re looking at. Instead of having the credential from the institution, why not have the credential from the professor? If you have a Hollis Robbins education, what would that signal? What would that credential mean as opposed to a degree from a university? There was some conversation about what that would look like, and one guy at the end of the dinner said, “Instead of OnlyFans, it’s like OnlyProfessors.”

The correct analogy here would be that it would be OnlyStudents, not OnlyProfessors (since the name contains the audience, not the performer). However, Robbins makes a good point. Higher education is, in part, an exercise in signalling (as I've noted before here and here). In fact, Bryan Caplan argued in his book The Case Against Education (which I reviewed here) that one third of the benefit of higher education is signalling (the other two-thirds is made up of genuine learning, socialisation, and transferable skills). 

The diploma that a student receives at the end of their higher education journey is a signal to employers of the student's quality as a future employee. The signal is credible to the employer because it is costly for the student to obtain, and costly in such a way that low-quality students wouldn't attempt the signal. University reputation matters here, because it is an assurance of the second of those conditions - low-quality students wouldn't attempt the signal of a Harvard degree, firstly because they wouldn't gain admission to Harvard in the first place. So, part of the signal comes from getting into Harvard. Second, low-quality students wouldn't attempt the signal of a Harvard degree because passing courses at Harvard is hard (or, at least, harder than at many other universities). The quality of the education at Harvard has traditionally been higher than elsewhere.

In the interview, Robbins makes a further important point, which is that we've spent the last several decades making higher education a commodity. Students studying a degree in a particular subject learn the same things, often using the same teaching materials, the same textbook, and the same style of assessment, regardless of which university they go to. That means that the quality of signal that arises from the education itself has reduced over time, meaning that most of the value of the signal arises from the admission process. Once a student has been admitted to Harvard, their signal is in place, and the further signalling from their education is lower than for comparable students in years past.

Robbins argues that generative AI is accelerating and expanding this commodification of education. Since generative AI has access, through its training corpus, to a large store of human knowledge, to add value over and above generative AI a professor must be a true expert in a very specific subfield. For students, learning a subject from a true expert still retains value, because generative AI cannot as easily replicate the learning that would occur from the expert. At that point, the university becomes less important as a mediator of education.

In other words, Robbins is arguing that the university could be disintermediated, with individual professors issuing credentials instead of universities. Who needs a Harvard degree, when you could have a Hollis Robbins degree? And since Robbins is an expert (in African American sonnet tradition, she says), the signal retains high value. However, that is true only to the extent that employers find the Hollis Robbins degree a credible signal.

That brings me to this Substack post by Auren Hoffman. Hoffman also argues that generative AI has diminished the value of the higher education signal. However, Hoffman argues for a different solution, this time from the perspective of the graduating student:

you have to show you can add value. that is it. that is the only thing.

the test the smart hiring manager applies in 2026 is simple. can you learn something on your own? can you finish what you started? can you do what you said you would do? these were always the skills that mattered. the difference is they are now the ONLY skills that matter, because the credential stopped doing the screening.

if you have a few years of experience, your resume can show you have these skills. if you are a new grad, you have to show what you have created and built.

Hoffman is overstating things a little, as university degrees are unlikely to disappear overnight. However, his critique is still important, as is the implication he draws. Hoffman argues that graduates (or young people, generally, since the solution doesn't depend on a student attending a university or completing a degree) need to become builders:

the most valuable thing a 22 year old can do in 2026 is create something. an app. a screenplay. a side business. an internal tool you wrote for a club you were in. a dinner series. a script that automates something annoying. a website. a chrome extension. a sculpture. a dance party. a discord bot. a substack with 32 readers and a real point of view. anything that moved from idea to working.

will the thing make money? probably not. that is not the point.

the point is that you taught yourself something. you finished it. you can describe what you learned, what broke, what you fixed, why you made the calls you made. that story is the new resume.

every hiring manager would rather interview a 22 year old with a launched app and a github full of weird side projects than a 22 year old with a 3.9 GPA from a top 50 school. it is not close. when one candidate has tangible evidence of what they can ship and the other has a transcript, the transcript will lose every time.

According to Hoffman, the best signal for young people to be sending in future is that they can build. Our students will ultimately be more successful if they can show off their skills (both technical and transferable skills) by building something. That is a signal that is costly, and costly in such a way that low-quality students will not attempt it. The signal is credible, and for the most part it retains currency even in the face of generative AI. Indeed, if the student builds something while effectively leveraging generative AI, then the signalling value to employers may be even greater. The key distinction is between the student using generative AI as a tool and the student using it as a substitute for doing the work. A student who can explain what they built, what broke, what they learned, and why they made the choices they made, is still sending a costly signal. A student who simply lets generative AI build for them is not. Employers would likely see through the latter pretty quickly.

Hoffman's idea isn't exactly new. When I think about some of my best students over the past two decades, they tend to be those that built something either while they were studying, or immediately after. For some, this was their own small business or entrepreneurial activity. For others, it was writing and publishing a research paper. Those were challenging tasks that set them apart from other students - a clear signal of quality. What Hoffman is essentially saying is that, with the signal from higher education itself being removed, the only remaining signal of quality is the signal from being a builder.

Where does that leave higher education staff? I think we can combine Robbins's and Hoffman's ideas, and chart a path forward. We don't need to start an OnlyStudents, and issue our own degrees, but we do need to cultivate closer relationships with our best students. Teach them to be builders. Encourage them to create things. Work with them and chart a path forward for their success. In other words, be a mentor.

Universities are absolutely going to hate this. Mentoring is not an activity that can be offered at scale. For example, there are simply not enough hours in the day for me to individually mentor all of the 350-plus students in my first-year economics class. Nor is mentoring easy to timetable, measure, standardise, or reward under current academic workload models. Universities are built around papers, credit points, learning outcomes, assessment rubrics, and student evaluations, all of which can be offered at scale. One-on-one mentoring doesn't work so well in that system. For example, postgraduate supervision already sits awkwardly within the system, with it being unclear whether it counts as teaching, or research. Nevertheless, mentoring may be one of the few ways that higher education can continue to offer something that is both valuable and difficult for generative AI to replicate.

If generative AI significantly reduces the signalling value of university education, and if students increasingly use it to avoid genuine learning, then the current mass higher education model looks increasingly fragile. Moreover, what remains is going to cost students a lot more. If having a high-quality mentor who can encourage a student to build is the path to their future employment, then it may be worth it. After all, high-quality signals are costly. That aspect of education, at least, won't have changed.

[HT: Marginal Revolution, for both the Hollis Robbins interview and the Auren Hoffman post]

Read more:

Monday, 27 April 2026

Who is morally responsible for grade inflation?

This article in The Conversation last year by Ciprian N. Radavoi, Carol Quadrelli, and Pauline Collins (all University of Southern Queensland) pointed me to their interesting article published in the Journal of Academic Ethics (open access) on who is morally responsible for grade inflation. The article attracted my attention because it engages moral philosophy in the task of identifying responsibility for grade inflation, a problem that I have blogged about before.

Radavoi et al. start by noting that grade inflation is unethical, relying on three main normative ethics theories: (1) deontology; (2) virtue ethics; and (3) consequentialism. As they explain:

In a deontological perspective, the problem with grade inflation is that it is a dishonest action...

A virtue ethics approach further substantiates grade inflation as unethical... As for what counts as virtues (that is, excellent traits of character), there have been numerous lists proposed in the history of philosophy, and courage, integrity and justice feature in most... regardless of the reason one inflates grades, that person does not act as a just person... It is easy to see how the act of equally rewarding with top marks hardworking and lazy students is unequal treatment that shows social irresponsibility, and fails to show leadership...

As for consequentialism, Radavoi et al. note all of the parties that are potentially harmed by grade inflation, including students whose grades are inflated (because they are "disincentivised from studying, shielded from the educative experience of failure, and instilled with a false sense of success"), employers (because grades become less valuable as a signal of the quality of job applicants), universities (whose reputations are damaged when they become known to be merely 'diploma mills'), and society generally. On the latter, Radavoi et al. note that:

If academic teachers do not fulfill their gatekeeping role, society will end up with incompetent doctors who may damage someone’s health, incompetent lawyers may ruin someone’s wealth or liberty, and so on. Also, if grades lose their quality of correctly indicating achievement, society will lose its trust in higher education and in universities as places of learning and excellence.

Having established grade inflation as unethical, Radavoi et al. then present the qualitative results from a survey of Australian academics, in the form of quotes from open-ended questions from the survey. This highlights the role of management coercion, and student evaluations of teaching , with the latter often being the mechanism through which management pressure is applied to academics. Radavoi et al. then unpack whether academics are manipulated into grade inflation, or coerced, concluding that:

...at least for casual academics, it seems safe to say they inflate grades under coercion, and this mitigates their moral blameworthiness: indeed, the coercive pressure on them, exercised via SETs, is insurmountable given the insecurity of their position and the overwhelming desire to secure another contract.

In contrast, in Radavoi et al.'s view, academics with continuing employment are generally not coerced into grade inflation, as those academics have greater agency to take a stand against grade inflation. Of course, context matters, and I suspect that there are not a lot of academics who feel genuinely secure enough in their employment to take a stand against university management on principle. I know from personal experience that even when a senior academic is willing to take a stand on behalf of a larger constituency of academics, those other academics will not necessarily voice their support in a public forum (and yet, simultaneously, be very willing to thank the senior academic in private). Taking a stand against grade inflation is only one example. Top-down one-size-fits-all rules imposed on teachers and their subjects, and enforcing onerous levels of flexibility that suit students but impose high costs on staff in terms of workload and administration, are other examples where pressure from university management has been applied, and yet academics have not effectively resisted. However, I am getting off topic.

Radavoi et al. do highlight that moral responsibility for grade inflation cannot be attributed solely to the academics responsible for grading. The institutional context, and management pressure (whether manipulation or coercion) are important too. However, Radavoi et al. do not consider the extent to which students are also implicated in this situation, albeit in a different way from academics or university management. As higher education has increasingly treated students as consumers of education services, successive cohorts of students have increasingly been encouraged to adopt the ideal of the ‘sovereign customer’, able to demand that education providers deliver particular outcomes for them. It is hardly surprising, then, that some students come to see higher grades not only as something to be earned, but as part of the education services that they have paid for. This does not make students primarily responsible for grade inflation, but it does mean that student expectations can reinforce the pressures placed on academics and university management.

So, perhaps moral responsibility cannot be attributed solely or largely to university management either. All three parties (academics, university management, and students) each bear some responsibility for grade inflation. However, that responsibility is not necessarily shared equally. The greatest responsibility should fall on those with the greatest power to change the incentives that make grade inflation attractive to students, convenient for managers, and sometimes the least risky option for academics. Which party bears the greatest responsibility will depend crucially on context.

Read more:

Thursday, 16 April 2026

Blame it on the rain (on open day), or campus tour weather and university choice

The University of Waikato Open Day is coming up next month. We'll have thousands of prospective students on campus for most of the day, learning about their study options, attending mini-lectures, talking to current staff and students, and collecting lots of free stuff that we give away. Many people question the value of these open days. Do they make a difference to students' choice of university? Undoubtedly, for some students at the margin they will make a difference. And from my experience, open days can affect some students' subject choice.

One thing that open days provide prospective students is a 'vibe' for their potential study location. This is where I might criticise open day, because it really is almost nothing at all like a 'normal' university day, so it doesn't give prospective students any idea what university is really like. The 'vibe' might also be affected by elements beyond the university's control, like the weather. How important is the weather? This recent NBER Working Paper by Olivia Feldman, Joshua Hyman, and Matthew McGann (all Amherst College), provides some idea. They look at the effect of weather on the day that a student undertakes a campus tour at an unnamed "institute of higher education" (IHE) on whether the student subsequently applies and/or enrols at that IHE, finding that weather affects applications but not enrolment. The campus tour is a more limited version of our open day:

Tours are typically given by current students and involve the guide walking the participants around campus for about an hour while sharing information about the institution, academics, student life, campus dormitories, academic buildings, dining halls, and sports and recreational facilities.

Feldman et al. use administrative data on all campus tours between summer 2016 and fall 2024, along with hourly weather data from The Weather Channel, and data on where each student subsequently enrolled from the National Student Clearinghouse. They report that overall:

28.8 percent of participants apply to the institution, and 2.2 percent ultimately enroll.

In their main analysis, Feldman et al. apply a simple OLS regression model, with application (or enrolment) at the focal IHE as the dependent variable, and weather variables (cloudy, rainy, and several temperature ranges to capture hot and cold days) as the explanatory variables of interest. They find that there is:

...a 1.7 percentage point (5.9%) lower application rate when the tour is cold, a 2.3 point (8.0%) lower rate when the tour is warmer, and a 2.9 point (10.1%) lower rate when the tour is hot. Further, cloudy tours reduce the application rate by 1.4 percentage points (4.9%), and tours with precipitation reduce it by 2.4 points (8.3%).

Those effects on applications are quite large in context. However, when Feldman et al. look at the effect on enrolment (rather than just application) they find statistically insignificant effects (albeit using several different composite variables for 'bad weather' as the explanatory variable of interest, rather than individual weather variables as in the earlier analysis).

My takeaway from this paper is that students’ choice of university is fairly resilient to the effects of weather on the day of their campus tour. While poor weather may reduce the chances that a student applies to a particular university, it doesn’t seem to have much effect on whether they ultimately enrol there. Of course, this is evidence from a single US institution, and may not easily translate to the New Zealand context. Still, extending these results to open days suggests that while the ‘vibe’ on the day might affect whether a student applies to the University of Waikato, and the weather contributes to that vibe, it probably isn’t an effect that we should worry too much about.

University enrolments fluctuate from year to year, and there are lots of variables that affect them. One thing this study suggests is that, while rain on open day might dampen spirits, it probably isn't the major cause of low enrolments. So, if the numbers are down, we needn’t blame it on the rain (on open day).

[HT: Marginal Revolution]

Thursday, 26 February 2026

Tuition fees, incentives, and 'ghost students'

When the New Zealand government introduced 'first-year fees free' in 2018, the universities expected a big uptick in student numbers. It didn't happen (as I discussed in this 2023 post). As the figure below (source) shows, the mild downward trend in domestic student numbers (equivalent full-time students, or EFTS) continued for at least a couple of years past 2018:

My colleagues were worried that we would see an increase in the number of students who enrol, and then do nothing at all (what we call 'ghost students'). My impression was that this didn't happen, but until now I never looked intentionally at the numbers. However, the figure below shows the proportion of each of my A Trimester ECON100 classes (up to 2017) or ECONS101 classes (for 2018 onwards) that were ghost students (I didn't teach the class in 2022, which is why there is no observation for that year). Here, I define a 'ghost student' as any student who didn't attempt any of the tests or exams (although they may have attended some classes during the trimester). In each trimester, the class had between 250-350 enrolments in total. [*]

As the figure shows, there was a big jump in 'ghost students' in 2021, but that is attributable to the COVID pandemic and the weirdness of that whole time period, rather than anything to do with fees-free. In most years, somewhere between three and five percent of students are 'ghosts'. In 2025, the government shifted from first-year fees free to final-year fees free. There's no evidence that change affected the proportion of 'ghost students' either. Or it's too early to tell - the proportion in 2025 was lower than either of the previous two years.

Why might we expect the changes in fees to affect the number of 'ghost students'? It comes down to incentives. As my ECONS101 students will hear next week, when the cost of something decreases, we tend to do more of it. First-year fees free decreased the cost of being a 'ghost student', so ceteris paribus (holding all else constant), we would expect to see more 'ghost students'. Final-year fees free (with first-year fees reintroduced) increased the cost of being a 'ghost student', so ceteris paribus, we would expect to see fewer 'ghost students'. The fact that didn't happen is interesting, and we'll come back to that a bit later.

To see why the New Zealand effect might be negligible, it helps to compare with a setting where student status comes with larger immediate benefits. To do that, I want to discuss this recent article by Johannes Berens (RH Köln), Leandro Henao, and Kerstin Schneider (both University of Wuppertal), published in the journal Labour Economics (ungated earlier version here). They look at the impact of the removal of tuition fees in North Rhine-Westphalia in Germany in 2011. Tuition fees were a very modest EUR500 per year (for every year of study), and Berens et al. essentially compare students who were more or less affected by the policy (depending on how many years they didn't have to pay fees for), looking at a range of academic outcomes including exam registrations and withdrawals, credit points earned, grades, and dropout probabilities, as well as the number of 'ghost students'.

Their data come from a single university, with over 11,000 students who first enrolled between 2008 and 2011. The students in the 2008 cohort would have graduated before the fees were removed, while those in the 2011 cohort would not have faced any fees at all. The other cohorts would have had fees in their later year/s, but not earlier year/s. Applying a difference-in-differences approach, Berens et al. find that:

...abolishing tuition fees significantly affected student behavior and academic outcomes. Active students reduced their academic performance by 1.7 credit points per semester (12 % relative to baseline), despite maintaining similar exam registration patterns... Additionally, the reform increased the prevalence of ghost students by 10 percentage points...

So, removing fees in this context substantially increased the proportion of 'ghost students' by 10 percentage points, from a baseline that was already over 10 percent (Berens et al. present the data by study semester, and the 'ghost student' proportion varies between 10 percent and 20-25 percent, depending on year and study semester).

What explains the high impact of removing fees in Germany? Berens et al. highlight the role of incentives, and in particular the generous nature of public assistance available to students. Specifically:

...student status confers substantial benefits, generally independent of academic performance... These benefits include subsidized health insurance (until age 25), state-wide public transport access (worth EUR 2900 annually), and parental child allowance (EUR 2450 annually). About 16 % of students also receive need-based grants averaging EUR 6800 annually...

So, being classified as a student can be quite lucrative in Germany, even if the student is a 'ghost'. That might also explain the lack of effect of first-year fees free in New Zealand. While the fees are higher in New Zealand than in Germany, being a student in New Zealand is hardly a pathway to great riches (at least, not during the time spent as a student - see this post, and the links at the end of it). The student allowance is not very generous, and while there are some other perks to being a student, cheap movie tickets and public transport are not exactly worth a lot of money. So, it shouldn't be much surprise that the impact in Germany was much larger than for a similar policy change in New Zealand.

Another reason that the impact was not apparent in New Zealand could be that many students do not pay their tuition fees immediately. Instead, many (perhaps most) students' tuition fees are paid by student loans. 'Student Greg' is probably quite content to say that the student loan is 'Graduated Greg's' problem, and not worry about it today. So, from the perspective of 'Student Greg', first-year fees free doesn't really impact the decision to become a student or not. It doesn't change the costs of being a student for 'Student Greg', because they don't consider paying back the student loan as part of the costs of studying today. [**] And that might explain why there was no incentive effect of first-year fees free in New Zealand (also, fees-free papers are not free if students fail them, as I noted in this 2023 post).

The incentives in Germany and New Zealand, when the tuition fees were changes, resulted in quite different impacts. In Germany, where the benefits of being a student were higher, lower costs of being a 'ghost student' induced many people to enrol, whereas in New Zealand, where the benefits of being a student are lower, and the costs of tuition are typically deferred to the future, lower costs of being a 'ghost student' appear to have made no difference.

The nature of incentives, and the costs and benefits around the decision, definitely matter. The policy takeaway from this is that tinkering with fees alone may induce more (or less) 'ghost students', so the other immediate benefits and costs associated with student status also need to be considered. 

*****

[*] The data are for only one paper, but ECON100 and ECONS101 have been, for the most part, compulsory papers for business students. In a couple of years, some students could avoid the paper by taking all of the other first-year business papers. However, unless 'ghost student' status was more likely for students who did not take first-year economics, these results should be broadly representative.

[**] Essentially, 'Student Greg' is heavily discounting the future. In my ECONS102 class, we say that 'Student Greg' exhibits present bias, and is therefore only quasi-rational, not purely rational. Of course, not all students will have acted like 'Student Greg', but if enough of them did, that would explain the lack of incentive effects of the changes in first-year fees.

Thursday, 8 January 2026

'First in family' as a measure of disadvantage in higher education

In higher education policy circles, it is an article of faith that students who are the first in their family (usually in the sense that neither their parents, nor any older siblings, has already studied at university) appear to be at higher risk of being unsuccessful in university education. The rationale links to Bourdieu's concept of social capital - students being able to tap into who they know (their family) and importantly what their family knows (about university study) matters. Family members with past university experience can help with university-specific knowledge - things like how to choose majors, manage workload, seek extensions, interpret feedback, and navigate various university systems and processes. This all makes the challenges of studying at university a little easier.

So, I was surprised to learn from this 2020 article by Anna Adamecz-Völgyi, Morag Henderson, and Nikki Shure (all University College London), published in the journal Economics of Education Review (ungated earlier version here), that there is actually limited empirical evidence supporting first-in-family as an indicator of disadvantage. It is that empirical gap that Adamecz-Völgyi et al. attempt to fill, but the interesting thing about this paper is not so much that they find support for first-in-family as a measure of disadvantage, but the mechanism through which it works.

Adamecz-Völgyi et al.  use data from 7707 students from the Next Steps (formerly the Longitudinal Study of Young People in England, LSYPE), which followed a cohort of young people born in 1989/1990. The 'age 25' wave of that study captures most of the cohort after they have completed university education. Adamecz-Völgyi et al. look at various measures of disadvantage, and how well they predict students participating in, or graduating from, higher education. Aside from first-in-family, their battery of disadvantage measures (which they refer to as Widening Participation (WP) measures) includes whether the student had special education needs at high school, whether they were eligible for free school meals, whether their parents were of low social class (based on occupation), whether their family was part of the 20 percent most deprived families (based on a measure of deprivation), whether they had care responsibilities while at high school, whether they were non-white ethnicity, whether they have a disability, whether they lived in a single-parent household, whether they had ever been in care, and whether they lived in an area of high socioeconomic deprivation.

Adamecz-Völgyi et al. use a few different methods to establish whether first-in-family (which they refer to as 'potential FiF', because they only have data on parental education, and not the education of older siblings) is a good predictor of disadvantage (in terms of participating in, or completing) in higher education, including: (1) comparing the predictive power of each variable in separate models (compared using the 'Area under the Receiver Operating Characteristic' curve (AUC)); (2) looking at whether adding first-in-family to a model that already includes a parsimonious set of other measures of disadvantage improves predictions; and (3) using a 'random forest' model to rank the predictors in terms of importance. The AUC is a measure of how often the model correctly predicts a binary variable (in this case, whether a student enrols/does not enrol in university, or whether they do/do not complete university). The random forest model identifies which variables are the most important by running many regressions with different selections of variables. In their analyses, Adamecz-Völgyi et al. find that:

When we compare potential FiF to other WP indicators, it emerges as the most important measure until we condition on prior attainment and all measures end up similarly predictive. We provide evidence that the effects of family background manifest in educational attainment at an early age and pre-university educational attainment is the most important channel of the effect of parental education on HE participation and graduation.

So, this research supports the common belief that first-in-family is a good measure of disadvantage. Moreover, it shows that first-in-family picks up some dimension of disadvantage that other common measures do not. However, the mechanism through which first-in-family affects higher education participation and success is almost entirely through the students' success in pre-university education. Students who are first-in-family at university tend to have worse performance in high school, and that largely accounts for their lower performance in university. Adamecz-Völgyi et al. conclude that:

...being potential FiF (and having social and economic disadvantages in general) matters all along the production function of a child's human capital from early childhood to university. Thus, the educational achievement measures that a university can use are contaminated by this pre-existing disadvantage carried along since early childhood (or probably, since birth). They do not reflect the child's true capacity, but rather the interaction of their innate abilities and family circumstances. Thus, WP measures that simply favour the disadvantaged student out of two students having the same level of pre-university attainment are not enough to widen participation: on average, those from disadvantaged backgrounds are not going have the same pre-university educational attainment levels than those from advantaged backgrounds. The attainment gap must be addressed explicitly by CA [Contextual Admissions] measures.

Instead, I see two ways that universities may respond to these results. Conditional on pre-university educational attainment, first-in-family students do not have worse higher education outcomes than other similarly-prepared students. The reason first-in-family students do less well on average is that they tend to enter university with lower prior educational attainment, and it’s that pre-university gap that accounts for most of the observed difference. On one hand, as Adamecz-Völgyi et al. argue, first-in-family is an indicator of disadvantage, and from a social justice perspective universities should try to mitigate sources of disadvantage whenever they are apparent. On the other hand, these results could be read as suggesting that universities shouldn't worry about first-in-family students, because they perform as well as otherwise similarly-prepared students. The problem is the lack of pre-university educational attainment, and that needs to be addressed in pre-university education, not at university. University-level support may help at the margin, but it risks being an ambulance at the bottom of the cliff. Moreover, two otherwise similar students in terms of pre-university educational attainment could be treated very differently under targeted support policies (such as Contextual Admissions) when one is first-in-family and the other is not, raising issues of fairness.

I'm not going to take a stand on which of those two perspectives (social justice or fairness) is more important. They both have merit. If you accept first-in-family as a measure of disadvantage, the actionable question is whether universities can cost-effectively close preparedness gaps after entry, or whether they should rely on advocating for changes in pre-university education. At least, this research can provide us with confidence that first-in-family is indeed a suitable measure of disadvantage in higher education.

Monday, 15 December 2025

Grade inflation at New Zealand universities, and what can be done about it

Grade inflation at New Zealand universities has been in the news recently. This is a delayed reaction to this report from the New Zealand Initiative released back in August, authored by James Kierstead. He collected data on grade distributions from all eight New Zealand universities (via Official Information Act requests), and looks at how those distributions have changed over time. The results are a clear demonstration of grade inflation, and most clearly demonstrated in Figure 2.1 from the report:

Over the period from the mid-2000s to 2024, the proportion of New Zealand university students receiving a grade in the A range has increased at every New Zealand university, and by more than ten percentage points overall. Kierstead notes that:

Overall, the median proportion of A-grades grew by 13 percentage points, from 22% to 35%... The largest increases occurred at Lincoln, where the proportion of As grew by 24 percentage points between 2010 and 2024 (from 15% to 39%), more than doubling, and Massey, where they grew by 17 percentage points (from 19% to 36%) from 2006 to 2023.

A similar pattern of increases, although not as striking, is seen for pass rates, which in 2024 were above 90 percent at every university except Auckland. The results are also apparent across different disciplines, as shown in Figure 2.4 from the report:

Of course, this sort of grade inflation is common across other countries as well, and Kierstead provides a comparison that shows that New Zealand grade inflation is not dissimilar from grade inflation in the US, UK, Australia, and Canada.

Kierstead then turns his attention to why there has been grade inflation. He first dismisses some possible explanations such as better incoming students (NCEA results have not improved, although even if they had that might be due to grade inflation as well), more female students (the proportion of female students has been flat over the past ten years, while grades have continued to increase), better funding (bwahahahaha - in fact, funding per student has declined in real terms since 2019, while grades have continued to increase), and student-staff ratios (which have declined over time, but the student-academic ratio, which is the one that should matter most, has barely changed).

So, what has caused grade inflation? Kierstead describes it as a collective action problem, akin to the tragedy of the commons first described by Garret Hardin in 1968:

It is our contention that grade inflation is the product of a dynamic that is not dissimilar to the tragedy of the commons. Just like Hardin’s villagers, academics pursue a good (in this case high student numbers) in a rational way (in this case by awarding more high grades). And just as with Hardin’s villagers, negative consequences ensue, with a common resource (sound grading) being depleted, to the cost of every individual academic as well as others...

In the grade inflation game, the good that academics want to maximize is student numbers. Individual academics, on the whole, want to have as many students in their courses as possible. This suggests that they are popular teachers and can help get them promoted (and hence gain more money and prestige). It can also help make sure the courses they want to teach stay on the menu.

I like this general framing of the problem, where 'sound grading' is a common resource - a good that is rival and non-excludable. However, I would change it slightly, by thinking about the common resource as being A grades generally, which are depleted when the credibility of those grades reduces. In my slightly different framing, awarding A grades is rival in the sense that one person awarding more A grades reduces the credibility of A grades awarded by others. Awarding A grades is non-excludable in the sense that if anyone can award A grades, everyone can award A grades (while it is possible to prevent academics from awarding A grades, universities would probably prefer not to do so because that would reduce student satisfaction). So, while the social incentive for all academics collectively is to reduce the award of A grades to keep the credibility of those grades high, the private incentive for each academic individually is to increase the proportion of A grades awarded, leading to fame and fortune (or, more likely, leading to fewer awkward conversations with their Head of School as to why their grade distribution is too low, as well as better student evaluations - see here and here, for example). Essentially then, the incentives are for academics to inflate grades. The universities have few incentives to act to reduce grade inflation, since higher grades increase student satisfaction and lead to greater enrolments.

However, there is a problem. As Kierstead notes, grade inflation is well-termed because its effects are similar to the inflation that economists are more familiar with:

If universities hand out more and more As in a way that isn’t justified by student performance, the value of an A will go down. The same job opportunities will ‘cost’ more As as As flood the market. Students who worked hard will see the value of their As decrease over time, just as workers in the economy see their savings decrease in value due to monetary inflation.

So, what to do? Kierstead offers a few solutions in the report, including moderation of grades, reporting grades differently on transcripts, calculating grades differently, making post-hoc adjustments to grade point averages, having national standardised exams by discipline, changing the way that universities are funded to reduce the incentive to inflate grades, changing the culture of academics, and giving out prizes for 'sound grading'. I'm not going to dig into those different solutions, because sometimes the simplest one is the best one. With that in mind, I pick this:

Perhaps the simplest addition that could be made to student transcripts alongside letter grades is the rank that students achieved out of the total number of students on the course. So a student’s transcript might read, for example, ‘Classics 106: Ancient Civilizations: A- (27th of 252).’...

Adding ranking information restores some of the signalling value of grades without needing to reverse grade inflation itself. To see why, consider an example. If an employer has the transcripts of two students, one of whom got an A- grade in econometrics and ranked 17th out of 22 students, while the other student got a B grade and ranked 3rd out of 29 students, it's pretty clear that the grade might not be capturing the full picture of the students' relative merit. Kierstead worries about this simple solution because:

A limitation of rank-ordering is that it might suggest that students who achieved only a lowly ranking had performed badly, whereas they might well have performed very well in an especially difficult course.

Possibly, but the key point is not how well students did in the course, but how well they did relative to the other students in the class, which is exactly what the ranking provides. The benefit of this approach is that providing a ranking alongside the grade would reduce the incentives for students to cherry pick easy papers that award high grades, because a high grade on its own would not necessarily lead to a good ranking within the class.

Of course, there are potential problems with the simple solution. One such problem is that comparisons across different cohorts of students might not be fair. Taking the example of the two students I gave earlier, perhaps the student who got an A- grade and ranked 17/22 completed the paper in a cohort that was particularly smart, while the student who got a B grade and ranked 3/29 completed the paper in a cohort that was less smart. In that case, the grade without the ranking might be a better measure.

Kierstead's more complex solutions don't really deal well with the problem of between-cohort comparisons, and suffer from being more complicated for non-specialists to understand. A simple ranking, or a percentile ranking, is relatively easy for HR managers to interpret. Having said that, the between-cohort comparisons issue might not be too much of a problem in any case. My experience though, is for classes of a sufficiently large size (30 or more), the grade distributions do not differ materially (and if they do, it is usually because of the teaching or the assessment, not the students).

I can see some incentive issues though. Would students start to choose papers that they suspect that many weak students complete? Good students might anticipate that this would lead to a higher grade and a better ranking, which will look better on their transcript. On the other hand, is that really any worse than what students are doing now, if they choose papers that give out easy grades?

There are also potential issues with stigmatising students who end up near the bottom of a large class (how dispiriting would it be to have your transcript say you got a grade of E, and ranked 317th out of 319 students?). Of course, that could be solved to some extent by only providing ranking information for students with passing grades. And consideration would also be needed for how to deal with very small classes (is a ranking of 4th out of 5 students meaningful?).

Grade inflation is clearly a problem. It's not just nostalgia to say that an A grade is not what it used to be. Grade inflation has real consequences for employers, because the signalling value of high grades is reduced (see here for more on signalling in education). This means that there are also real consequences for high-quality students, who find it more difficult to differentiate themselves from average students. Solving this problem shouldn't involve government intervention to change university funding formulas, or trying to change academic culture. It shouldn't involve complicated statistical manipulations of grades. It really could be as simple as reporting students' within-class ranking on their academic transcripts.

The question now is whether any university would take it on themselves to do so. The credibility of university grades depends on it.

[HT: Josh McNamara, earlier in the year]

Read more:

Sunday, 14 December 2025

Online and blended learning lead to similar outcomes on average, at lower cost but lower student satisfaction

It's been a while since I've written about online or blended learning, which may seem surprising given the ample opportunities for us to learn about online learning during the pandemic. Perhaps I'm still dealing with the trauma of that, or perhaps I have just pivoted more to understanding the emerging role of AI in education. Nevertheless, I recently dipped my toes back into the research on online and blended learning, reading this 2020 article by Igor Chirikov (University of California, Berkeley) and co-authors, published in the journal Science Advances (open access).

Chirikov et al. evaluate a large multisite randomised controlled trial of online and blended learning in engineering, across three universities in Russia. As they explain:

In the 2017–2018 academic year, we selected two required semester-long STEM courses [Engineering Mechanics (EM) and Construction Materials Technology (CMT)] at three participating, resource-constrained higher education institutions in Russia. These courses were available in-person at the student’s home institution and alternatively online through OpenEdu. We randomly assigned students to one of three conditions: (i) taking the course in-person with lectures and discussion groups with the instructor who usually teaches the course at the university, (ii) taking the same course in the blended format with online lectures and in-person discussion groups with the same instructor as in the in-person modality, and (iii) taking the course fully online.

The course content (learning outcomes, course topics, required literature, and assignments) was identical for all students.

Their sample is made up of 325 second-year university students, with 101 randomly assigned to in-person, 100 to blended, and 124 to online. All students then completed the same final examination. Looking at student performance, Chirikov et al. find:

...minimal evidence that final exam scores differ by condition (F = 0.26, P = 0.77)... The average assessment score varied significantly by condition (F = 3.24, P = 0.039): Students under the in-person and blended conditions have similar average assessment scores (t = 0.26, P = 0.80), but those under the online condition scored 7.2 percentage points higher (t = 2.52, P = 0.012). This effect is likely an artifact of the more lenient assessment submission policy for online students, who were permitted three attempts on the weekly assignments.

The lack of a difference in student performance on average across different learning modes is a common feature of the literature (see the links at the end of this post). It would have been interesting if Chirikov et al. had undertaken a heterogeneity analysis to see whether online and blended modes advantage the more able and engaged students, while disadvantaging the less able and engaged students (also a feature of the literature on online and blended learning). The general result that online and blended learning provides benefits for top students but harms weaker ones is a point I’ve discussed many times before (see the links below for more).

Chirikov et al. then look at student satisfaction, and despite claiming that "we find minimal evidence that student satisfaction differs by condition", Table 3 in the paper does show that students in the online mode report a statistically significant five percentage points lower satisfaction than in-person students, while students in the blended mode report lower satisfaction (by about 2-2.5 percentage points) than in-person students, although the latter difference was not statistically significant.

Finally, Chirikov et al. evaluate the effect on the cost of education, finding that:

Compared to the instructor compensation cost of in-person instruction, blended instruction lowers the per-student cost by 19.2% for EM and 15.4% for CMT; online instruction lowers it by 80.9% for EM and 79.1% for CMT...

These cost savings can fund increases in STEM enrollment with the same state funding. Conservatively assuming that all other costs per student besides instructor compensation at each university remain constant, resource-constrained universities could teach 3.4% more students in EM and 2.5% more students in CMT if they adopted blended instruction. If universities relied on online instruction, then they could teach 18.2% more students in EM and 15.0% more students in CMT.

I don't think it will come as a surprise to anyone that online and blended learning are more cost-effective. There is little doubt that it has factored into some of the push towards online and blended learning across higher education over time.

Given that, in this study, both online and blended learning lead to similar outcomes on average, one might be tempted to suggest that they are good value for money from the university’s or funder's perspective. For cash-strapped institutions (or governments), the temptation to expand online provision on the back of such numbers is obvious. However, we should be cautious about drawing that conclusion. The lower student satisfaction in the blended and (especially) online modes should be a worry (at least to those who care about student satisfaction). And, as alluded to earlier, the average student performance can hide important heterogeneity between more engaged and less engaged students.

The real question here isn’t whether online and blended learning can be as effective on average, but whether we are comfortable trading lower satisfaction and potential for harms to less engaged students for lower cost of delivery and higher enrolments.

Read more: