Sunday, 25 May 2025

Universities' (and teachers') cheap talk on generative AI and assessment

Generative AI should be changing the way that universities assess students. I say "should be", rather than "is", because it seems to me that a lot of teaching staff really have their head in the sand on this, continuing to assess in a very similar way, and simply attaching a warning label ("thou shalt not use generative AI") to each assessment, as if that will make a difference. The futility of that approach is the topic of this new article by Thomas Corbin, Phillip Dawson (both Deakin University), and Danny Liu (University of Sydney), published in the journal Assessment and Evaluation in Higher Education (open access).

Corbin et al. focus attention on the university level frameworks, but many of the things that they say apply equally to each paper. When considering how to approach the impact of generative AI on assessment, and how assessment needs to change as a result of generative AI Corbin et al. distinguish two approaches: (1) discursive changes, which involve telling students what is and what is not permitted; and (2) structural changes, which involve changing the assessment itself so that the way that students may use AI (or not) is specifically factored into the assessment.

Corbin et al. make the important point that:

...existing frameworks predominantly rely on merely discursive methods which introduces significant vulnerabilities related to compliance and enforceability, ultimately undermining assessment validity and institutional reputation. Although these systems may have value in other areas, for example by assisting teachers to conceptualise the different ways AI may be used in a task, from a validity standpoint any change which is merely discursive and not structural is likely to cause more harm than good.

Discursive approaches include 'traffic light' systems, or various assessment scales, where teachers communicate to students what generative AI use is or is not allowed. They also include requirements for students to disclose the use of generative AI in their assessments. The problem with discursive changes to assessment is obvious:

Without reliable detection mechanisms, prohibitions against AI use remain merely discursive. This technological limitation exposes a more fundamental issue with discursive approaches. That is, they rely entirely on student compliance with rules that cannot be enforced.

There is not reliable way of detecting generative AI use in student assessment. The best that teachers can do is to rely on vibes. Or when a student writes in their essay that they are 'delving' into a 'rich tapestry' or a 'multifaceted realm' and trying to find the 'intricate balance' or a 'symbiotic relationship'.

Corbin et al. instead advocate for structural changes, which they define as:

Modifications that directly alter the nature, format, or mechanics of how a task must be completed, such that the success of these changes is not reliant on the student’s understanding, interpretation, or compliance with instructions. Instead, these changes reshape the underlying framework of the task, constraining or opening the student’s approach in ways that are built into the assessment itself.

They illustrate with some examples, starting with:

A traditional take-home essay (asynchronous) provides students with ample opportunity to use AI without detection, regardless of what instructions are provided. In contrast, a supervised in-class writing exercise (synchronous) inherently limits AI assistance by its very structure.

Justin Wolfers would approve. However, Corbin et al. rightly note that:

This doesn’t mean that all assessment should become synchronous and supervised; certainly, asynchronous assessment has valuable benefits for developing certain skills. The key is aligning the assessment structure with what we genuinely want to measure. If we want to develop a student’s ability to think deeply and develop complex arguments over time, an asynchronous format may be appropriate, but we would need to build in structural assessment elements that capture the development process rather than just the final product.

Corbin et al. don't leave us hanging. Even though they can't solve all of our AI-related assessment issues, they do offer some suggestions:

First, structural changes frequently involve reorienting assessment from output to process. Rather than evaluating only the final product, which could potentially be AI-generated, assessment may be designed to capture the student’s development and attainment of understanding and skill over time. This might mean building in authenticated checkpoints where students must demonstrate their evolving thinking. For instance, rather than simply submitting a final essay, students might need to participate in live discussions about their developing ideas or demonstrate how their thinking evolved through structured peer feedback sessions...

Second, structural changes often involve viewing assessment validity at the unit or module level rather than the task level. Instead of trying to ensure each individual assignment is AI-proof (an increasingly futile endeavour), educators can design interconnected assessments where later tasks explicitly build on a student’s earlier work.

This relates back to two earlier posts of mine. This post talks about assessment specifically, while this post talks about changing the way that students interact with generative AI in learning and assessment tasks, so that they skills are scaffolded through their degree. We do need to make changes to assessment practices. It is possible for assessments to change in ways that take account of students' access to generative AI. It is not necessary to make generative AI use forbidden in all situations. It is probably equally unhelpful to make it 'open season' on generative AI use either. As with all things, there is a balance to be had, and university teachers need to find that balance.

Corbin et al. conclude that discursive changes to assessment:

...remain powerless to prevent AI use when they rely solely on student compliance. They say much but change little. They direct behaviour they cannot monitor. They prohibit actions they cannot detect. In other words, when it comes to appropriate assessment change for a time of AI, talk is cheap.

Simply using a set of written rules on when and how students can use generative AI is, at best, ineffective, and at worst, may actively harm student learning. Those rules are cheap talk. We can do much better.

[HT: Maria Neal]

Read more:

Saturday, 24 May 2025

ChatGPT and economics homework questions

Back in December last year, I briefly discussed why we eliminated Moodle quizzes from my ECONS101 assessment (we still have quizzes, but they are not for credit, and they happen every day - more on that in a future post). The problem with Moodle quizzes is that there are browser extensions that will automatically answer Moodle quizzes using generative AI. That makes Moodle quizzes largely a waste of time as an assessment tool (although I believe that they still have value as a learning tool).

Of course, Moodle quizzes are not the only assessment that have been rendered obsolete by generative AI. As Justin Wolfers notes, any high-stakes at-home assessment is now essentially worthless. But it's not just high stakes assessment. Problem sets or homework assignments are also affected. This new article by Rachel Faerber-Ovaska (Youngstown State University) and co-authors, published in the journal Bulletin of Economic Research (open access) asks the question, "Has ChatGPT made economics homework questions obsolete?"

Faerber-Ovaska et al. test the ability of ChatGPT to answer 1112 multiple choice and 186 long answer questions (which they call essay questions) from the question banks of the 2nd edition of Principles of Economics by Greenlaw and Shapiro. They find that:

The bot answered 67.63% of the 1112 multiple-choice questions correctly.

Faerber-Ovaska et al. then looked at the characteristics of the questions that ChatGPT got wrong, and report that:

The inclusion of tables or figures, as well as higher levels of difficulty, were found to significantly decrease the odds of ChatGPT answering correctly. For example, the model estimated odds of ChatGPT answering a question with a table correctly were only 0.45, corresponding to an 80% lower probability compared to a question with no table. Overall, the bot struggled with material from chapters requiring visual interpretation, such as supply and demand, elasticity, theory of the firm, and financial economics.

As for the long answer (essay) questions:

...we found that the bot scored higher for clarity than for content. The bot score for content was an A on 72.0% of questions, whereas for clarity, the bot scored an A on 93.5% of the questions. Overall, for essay questions, the bot’s responses earned an A in 72.0% of the questions and a B in 18.3% of the questions.

It is worth noting that Faerber-Ovaska et al. were testing ChatGPT 3.5, and more recent versions of ChatGPT (and other large language models) are likely to perform even better. Nevertheless, in answering the question they post, Faerber-Ovaska et al. conclude:

Have the economics homework and test questions we currently rely on been rendered obsolete by ChatGPT? The answer is: yes, as they are used now.

The "as they are used now" is important in that sentence. Economics teachers (and teachers in other disciplines) need to change the way that we do things. Homework may still be effective as a learning tool, but it will not be effective if the approach is simply to have students turn in (or complete online) homework problems that are copy-pasted from a large language model. For the moment, these models are not great at drawing accurate diagrams in economics, but that just means that economics has mere moments more time to adapt than some other disciplines. Homework may still have a place in student learning, but it needs to be structured in a way that ensures that students engage, even if they are using generative AI. In my ECONS101 and ECONS102 classes, our low-tech solution is to have students complete homework in handwritten form only. The homework is not graded, but is built on in-person in tutorials (and completion of the tutorials is worth a small amount of marks). This ensures that even if students are using generative AI, they need to engage in class as well. This is reinforced by an approach to assessment that is heavily weighted towards in-person invigilated assessment, where the use of generative AI is unlikely (for now!).

Homework isn't dead (in economics, or in general). But its role in student learning and assessment needs to change.

Read more:

Friday, 23 May 2025

This week in research #76

Here's what caught my eye in research over the past week:

  • Huseynov (with ungated earlier version here) finds that students reduce their confidence regarding their future earning prospects after exposure to AI debates, and this effect is more pronounced after reading discussion excerpts with a pessimistic tone (but will their actual experience match their expectations, pessimistic or otherwise?)
  • Pipke (open access) studies 7,000 soccer penalty shootouts and 74,000 kicks and finds no evidence of a first-mover or second-mover advantage in winning probability
  • Böheim, Freudenthaler, and Lackner (open access) find that women’s NCAA basketball teams with a male head coach are 6 percentage points more likely to take risk than women’s teams with a female head coach
  • Clemens and Strain (with ungated earlier version here) study the interplay between minimum wages and union membership, and find that each dollar in minimum wage increase predicts a 5 percent increase (0.3 pp) in the likelihood of union membership among individuals ages 16–40, which may explain why unions are in favour of minimum wage increases
  • Yang and Zhou (open access) find that, after controlling for quality as well as author-, paper-, and journal-specific attributes, publications in economics with a Chinese first author receive 14% fewer citations
  • Abdulla and Mourelatos (open access) find that Russian migrants are significantly less likely than local Kazakhs, local Russians, or Kyrgyz migrants to receive job interview invitations in Kazakhstan, based on data collected after the start of the Russia-Ukraine War

In other news, I had two articles published in The Conversation this week, on the New Zealand Budget:

  • This article, co-authored with Michael Ryan, discusses the difficulty of economic forecasting and why this 2025 Budget was particularly tight
  • This article gave my quick take on the Budget (since it was published only a couple of hours after the Budget was announced), as well as summarising the key Budget announcements in each area

Thursday, 22 May 2025

Costco buttering up New Zealand consumers

Overseas, Costco uses a variety of products as loss leaders, including rotisserie chicken (see here). Right now in New Zealand, it appears to be butter. As the New Zealand Herald reported yesterday:

It was organised chaos this week at Costco when another delivery of butter arrived.

This butter is not just any butter – while the other supermarkets are selling a 500g slab for up to $10, Costco’s butter is $9.99 for a kg...

Chris Schulz, a senior investigative journalist at Consumer NZ, said it looked likely that the Costco butter was a loss leader.

“The retailer’s Facebook page is flooded with people speculating when the butter might be back on shelves, debating when to visit, and showing off when they do get it.

“With butter costing at least $17 per kilo elsewhere else, Costco’s pricing makes them look like the ‘good guys’ in contrast to our supermarket duopoly. Once they’re in store, I’m sure many people are picking up roast chickens, cheese, and giant tubs of biscuits too.”

As I note in my ECONS101 class, the ideal loss leading product is one that has a high price elasticity of demand, and lots of complementary goods. Elastic demand means that a decrease in price will increase the number of consumers by a lot. So, loss leading will get a lot more consumers in store. And complementary goods are goods where lowering the price of one good causes the consumer to buy more of the other good. Most supermarket staples that are regularly purchased will be complementary goods, because consumers tend to buy them together on the same shopping trip. So, lowering the price of one causes the consumer to buy more of the other goods on their shopping list.

In this case, by selling butter at a loss (and it must be at a loss, because there's no way that selling butter at half the price of other retailers is profitable), Costco is able to attract many more consumers, who then buy other things that Costco can profit from. The Herald article offers some examples, including this one:

Kaleb Halverson decided to start making the trip from New Plymouth to Auckland to deliver Costco’s 1kg blocks of butter at $9.99 to customers across the Taranaki region.

He only had a few orders at first, but they kept rolling in...

He brings back everything the store has to offer, but said butter is definitely top of the list. “It’s our hot item; at the moment, every order has butter.”

Other popular products are cleaning products and snacks.

Costco sells the butter at a loss, and makes up for it with greater sales of (and profits from) cleaning products and snacks.