Thursday, 31 October 2024

Book review: Wonderland (Steven Johnson)

When I think about the dramatic changes in society that have occurred since the end of the Industrial Revolution, one of the trends that stands out (to me) is the massive increase in leisure time. In the 19th Century, most people worked far more hours than they do today. The recent decades of that trend were well-described in Daniel Hamermesh's book Spending Time (which I reviewed here). What was left unexplored in that book was the way that leisure pursuits have affected the economy and society.

That is the purpose of Steven Johnson's book Wonderland, which is subtitled "How play made the modern world". Johnson describes the book as:

...a history of play, a history of the pastimes that human beings have concocted to amuse themselves as an escape from the daily grind of subsistence. This is a history of what we do for fun.

The book is comprised of chapters devoted to fashion and shopping, music, food, entertainment, games, and our use of public space. Each chapter is well written and well resourced, and a pleasure to read. Johnson is a great storyteller and the stories he presents are interesting and engaging.

However, from the first chapter, I struggled with the overall thesis of the book, which is that changes in leisure pursuits drove broader societal changes and economic changes. This is most glaringly demonstrated in the first chapter, where Johnson contends that it was the desire for fashion that drove the Industrial Revolution:

When historians have gone back to wrestle with the question of why the industrial revolution happened, when they have tried to define the forces that made it possible, their eyes have been drawn to more familiar culprits on the supply side: technological innovations that increased industrial productivity, the expansion of credit networks and financing structures; insurance markets that took significant risk out of global shipping channels. But the frivolities of shopping have long been considered a secondary effect of the industrial revolution itself, and effect, not a cause... But the Calico Madams suggest that the standard theory is, at the very least, more complicated than that: the "agreeable amusements" of shopping most likely came first, and set the thunderous chain of industrialization into motion with their seemingly trivial pursuits.

In spite of the excellent prose, I'm not persuaded by the demand-side argument for the Industrial Revolution, which flies in the face of lots of scholarship in economic history (as well as in history). Now, it may be that the first chapter just made me grumpy. But Johnson draws several conclusions which are, at best, a selective interpretation of the evidence. And at times, he makes comparisons that are somewhat odd, such as a comparison between the tools and technologies available to artists and scientists and those available to musicians in the 17th Century, concluding that there were fewer and less advanced tools available to artists and scientists than for musicians. There doesn't seem to be any firm basis to make such a comparison (how does one measure how advanced technologies in different disciplines are, in order to compare them?).

The final chapter, though, was a highlight to me. There was a really good discussion of the role of taverns in the American Revolution. And in that discussion, Johnson acknowledges that it is difficult to establish a causal relationship (which made me again wonder why he was unconcerned about the challenges of causality between shopping and the industrial revolution earlier in the book). I really appreciated the discussion of the work of Jürgen Habermas, Ray Oldenburg, and the "third places" (places of gathering that are neither work, nor home). It reminded me of my wife's excellent PhD thesis on cafés.

Overall, I did enjoy the book in spite of my griping about the overall thesis and the way that Johnson sometimes draws conclusions from slim evidence. If you are interested in the history of leisure pursuits, I recommend it to you.

Wednesday, 30 October 2024

Some notes on generative AI and assessment (in higher education)

Last week, I posted some notes on generative AI in higher education, focusing on positive uses for AI for academic staff. Today, I want to follow up with a few notes on generative AI and assessment, based on some notes I made for a discussion at the School of Psychological and Social Sciences this afternoon. That discussion quickly evolved into more of a discussion on intentional design of assessment more generally, rather than focusing on the risks of generative AI to assessment more specifically. That's probably a good thing. Any time academic staff are thinking more intentionally about the assessment design, the outcomes are likely to be better for students (and for the staff as well).

Anyway, here are a few notes that I made. Most importantly, the impact of generative AI on assessment, and the robustness of any particular item of assessment to generative AI, depends on context. As I see it, there are three main elements of the context of assessment that matter most.

First, assessment can be formative, or summative (see here, for example). The purpose of formative assessment is to promote student learning, and provide actionable feedback that students can use to improve. Formative assessment is typically low stakes, and the size and scope of any assessment item is usually quite small. Generative AI diminishes the potential for learning from formative assessment. If students are outsourcing (part of) their assessment to generative AI, then they aren't benefiting from the feedback or the opportunity for learning that this type of assessment provides.

Summative assessment, in contrast, is designed to evaluate learning, distinguish good students from not-so-good students from failing students, and award grades. Summative assessment is typically high stakes, with a larger size and scope of assessment than formative assessment. Generative AI is a problem in summative assessment because it may diminish the validity of the assessment, in terms of its ability to distinguish between good students and not-so-good students, or between not-so-good students and failing students, or (worst of all) between good students and failing students.

Second, the level of skills that are assessed is important. In this context, I am a fan of Bloom's taxonomy (which has many critics, but in my view still captures the key idea that there is a hierarchy of skills that students develop over the course of their studies). In Bloom's taxonomy, the 'cognitive domain' of learning objectives is separated into six levels (from lowest to highest): (1) Knowledge; (2) Comprehension; (3) Application; (4) Analysis; (5) Synthesis; and (6) Evaluation.

Typically, first-year papers (like ECONS101 or ECONS102 that I teach) predominantly assess skills and learning objectives in the first four levels. Senior undergraduate papers mostly assess skills and learning objectives in the last three levels. Teachers might hope that generative AI is better at the lower levels - things like definitions, classification, understanding and application of simple theories, models, and techniques. And indeed, it is. Teachers might also hope that generative AI is less good at the higher levels - things like synthesising papers, evaluating arguments, and presenting its own arguments. Unfortunately, it also appears that generative AI is also good at those skills. However, context does matter. In my experience, and this is subject to change because generative AI models are improving rapidly, generative AI can mimic the ability of even good students at tasks at low levels of Bloom's taxonomy, which means that tasks at that end lack any robustness to generative AI. However, at tasks higher on Bloom's taxonomy, generative AI can mimic the ability of failing and not-so-good students, but is still outperformed by good students. So, many assessments like essays or assignments that require higher-level skills may still be a robust way of identifying the top students, but will be much less useful for distinguishing between students who are failing and students who are not-so-good.

Third, authenticity of assessment matters. Authentic assessment (see here, for example) is assessment that requires students to apply their knowledge in a real-world contextualised task. Writing a report or a policy brief is a more authentic assessment than answering a series of workbook problems, for example. Teachers might hope that authentic assessment would engage students more, and reduce the use of generative AI. I am quite sure that many students are more engaged when assessment is authentic. I am less sure that generative AI is used less when assessment is authentic. And, despite any hopes that teachers have, generative AI is just as good in an authentic assessment as it is in other assessments. It might be better in fact. Consider the example of a report or a policy brief. The training datasets of generative AI no doubt contain lots of reports and policy briefs, so it has lots of experience with exactly the types of tasks we might ask students to complete in an authentic assessment.

So, given these contextual factors, what types of assessment are robust to generative AI. I hate to say it, and I'm sure many people will disagree, but in-person assessment cannot be beaten in terms of robustness to generative AI. In-person tests and examinations, in-person presentations, in-class exercises, class participation or contributions, and so on, are assessment types where it is not impossible for generative AI to influence, but where it is certainly very difficult for it to do so. Oral examinations are probably the most robust of all. It is impossible to hide your lack of knowledge in a conversation with your teacher. This is why universities often use oral examinations at the end of a PhD.

In-person assessment is valid for formative and summative assessment (although the specific assessments used will vary). It is valid at all levels of learning objectives that students are expected to meet. It is valid regardless of whether assessment is authentic or not. Yes, in case it's not clear, I am advocating for more in-person assessment.

After in-person assessment, I think the next best option is video assessment. But not for long. Using generative AI to create a video avatar to attend Zoom tutorials, or to make a presentation, is already possible (HeyGen is one example of this). In the meantime though, video reflections (as I use in ECONS101), interactive online tutorials or workshops, online presentations, or question-and-answer sessions, are all valid assessments that are somewhat robust to AI.

Next are group assessments, like group projects or group assignments, or group video presentations. The reason that I believe group assessments are somewhat robust is that it requires a certain amount of group cohesion to make a sustained effort at 'cheating'. I don't believe that most groups that are formed within a single class are cohesive enough to maintain this (although I am probably too hopeful here!). Of course, there will be cases when just one group member's contribution to a larger project was created with generative AI, but generally it would take the entire group to do so. When generative AI for video becomes more widespread, group assessments will become a more valid assessment alternative than video assessment.

Next are long-form written assessments, like essays. I'm not a fan of essays, as I don't think they are authentic as assessment, and I don't think they assess skills that most students are likely to use in the real world (unless they are going onto graduate study). However, they might still be a valid way of distinguishing between good students and not-so-good students. To see why, read this New Yorker article by Cal Newport. Among other issues, the short context window of most generative AI models means that it is not great at long-form writing, at least compared with shorter pieces. However, generative AI's shortcomings here will not last, and that's why I've ranked long-form writing so low.

Finally, online tests, quizzes, and the likes should no longer be used for assessment. The development of browser plug-ins that can be used to answer multiple-choice, true/false, fill-in-the-blanks, and short-answer-style questions automatically, with minimal student input (other than perhaps to hit the 'submit' button), makes these types of assessments invalid. Any attempts to thwart generative AI in this space (and I've seen things like using hidden text, using pictures rather than text, and other similar workarounds) are at best an arms race. Best to get out of that now, rather than wasting lots of time trying (but generally failing) to stay one step ahead of the generative AI tools.

Finally, I know that many of my colleagues have become attracted to getting students to use generative AI in assessment. This is the "if you can't beat them, join them" solution to generative AI's impact on assessment. I am not convinced that this is a solution, for two reasons.

First, as is well recognised, generative AI has a tendency to hallucinate. Users know this, and can recognise when a generative AI has hallucinated in a domain in which they (the user) have specific knowledge. If students, who are supposed to be developing their own knowledge, are being asked to use or work with generative AI in their assessment, at what point will those students develop their own knowledge that they can use to recognise when the generative AI tool that they are working with is hallucinating? Critical thinking is an important skill for students to develop, but criticality in relation to generative AI use often requires the application of domain-specific knowledge. So, at the least, I wouldn't like to see students encouraged to work with generative AI until they have a lot of the basics (skills that are low on Bloom's taxonomy) nailed first. Let generative AI help them with analysis, synthesis, or evaluation, while the student's own skills in knowledge, comprehension, and application allow them to identify generative AI hallucinations.

Second, the specific implementations of assessments that involve students working with generative AI are not often well thought through. One common example I have seen is to give students a passage of text that was written by AI in response to some prompt, and ask students to critique the AI response. I wonder, in that case, what stops the students from simply asking a different generative AI model to critique the first model's passage of text?

There are good examples of getting students to work with generative AI though. One involves asking students to write a prompt, retrieve the generative AI output, and then engage in a conversation with the generative AI model to improve the output, finally constructing an answer that combines both the generative AI output and the student's own ideas. The student then submits this final answer, along with the entire transcript of their conversation with the generative AI model. This type of assessment has the advantage of being very authentic, because it is likely that this is how most working people engage with generative AI for competing work tasks (I know that it's one of the ways that I engage with generative AI). Of course, it is then more work for the marker to look at both the answer and the transcript that led to that answer. But then again, as I noted in last week's post, generative AI may be able to help with the marking!

You can see that I'm trying to finish this post on a positive note. Generative AI is not all bad for assessment. It does create challenges. Those challenges are not insurmountable (unless you are offering purely online education, in which case good luck to you!). And it may be that generative AI can be used in sensible ways to assist in students' learning (as I noted last week), as well as in students completing assessment. However, we first need to ensure that students are given adequate opportunity to develop a grounding on which they can apply critical thinking skills to the output of generative AI models.

[HT: Devon Polaschek for the New Yorker article]

Read more:

Monday, 28 October 2024

Generative AI may increase global inequality

As I noted in a post earlier this month, the general public appears to be worried about the impact of generative artificial intelligence on jobs and inequality. Some economists are clearly worried as well. Consider this post on the Center for Global Development blog, by Philip Schellekens and David Skilling. They note three reasons why generative AI might increase global inequality, because: (1) richer countries are better equipped to harness AI’s benefits; (2) poorer countries may be less prepared to handle AI’s disruptions; and (3) AI is intensifying pressure on traditional development models.

I have a lot of sympathy for these arguments, but it is worth exploring them in a bit more detail. Here's part of what Schellekens and Skilling said on the first reason:

High-income countries, along with wealthier developing nations, hold a distinct advantage in capturing economic value from AI thanks to superior digital infrastructure, abundant AI development resources, and advanced data systems...

When many people may think about economic growth, we think about catch-up growth. Developing countries often have growth rates that exceed those in developed countries. There are vivid examples of catch-up growth, like the way many developing countries were able to bypass copper telephone lines and move straight to mobile telecommunications. Could AI be like that? It's a hopeful vision. However, the problem with that argument is that AI isn't quite the same as the telecommunications example. There is no outdated technology that is being replaced by AI (unless humans count?). So, developing countries can't leapfrog technology and catch up. If a country doesn't have the technology infrastructure and capital necessary to develop their own AI models, they will be forced to use models developed in other countries. That creates problems for developing countries, and Schellekens and Skilling note two particular concerns:

First, AI could reinforce the dominance of wealthier nations in high-value sectors like finance, pharmaceuticals, advance manufacturing, and defense. As richer countries use AI to enhance productivity and innovation, it becomes harder for poorer countries to penetrate these markets.

Second, while AI is poised to primarily disrupt skill-intensive jobs more prevalent in advanced economies, it can also undermine lower-cost labor in developing countries. Automation in manufacturing, logistics, and quality control would enable wealthier nations to produce goods more efficiently, reducing the need for low-wage foreign workers. This shift, supported by AI-driven predictive analytics and customization capabilities, may allow richer countries to outcompete on cost, speed, and product desirability.

Note that second argument says that in spite of any increase in inequality within developed countries (which is what the general public was most concerned about in my previous post), there would be increases in global inequality because of the differential impact on different labour markets. This is a consequence of past labour market polarisation, where different countries have become reliant on employment in different sectors.

On their second point, Schellekens and Skilling note that, while the social safety net in developed countries may insulate their populations from the negative impacts of AI (a point that I'm not sure that many would agree with), the situation in developing countries is quite different:

Limited resources and underdeveloped social protection systems mean they are less equipped to absorb the economic and social shocks caused by AI-driven disruptions. Many lower-income countries already struggle with high rates of informal employment and fragile labor markets, leaving workers highly vulnerable to sudden economic shifts.

The lack of fiscal space also restricts these countries from investing in crucial areas like reskilling programs, infrastructure upgrades, or targeted welfare schemes to support affected communities. Without such mechanisms, the impact of AI-related job losses could exacerbate unemployment and deepen poverty.

It would be interesting to see some research on the expected impact of generative AI on informal sector employment, but I except that Schellekens and Skilling are largely correct about the impacts on formal sector employment in developing countries.

Finally, on their third point, Schellekens and Skilling note that the model of development that many countries have followed in recent decades, moving first from an agrarian economy, into low-technology manufacturing (like garments), and then into higher-technology manufacturing over time, has become less viable for developing countries over time, and that generative AI may impact the obvious alternative, which is export-oriented service industries:

Countries like the Philippines and India have seen success in business process outsourcing, thanks to booming call center industries and IT services. But AI poses a threat to this model as well. AI has the potential to reduce the labor intensity of these activities, eroding the competitive edge in the international marketplace of lower-cost service providers.

If AI were to undermine labor-intensive service industries, developing countries may find it harder to identify viable pathways for growth, posing a significant challenge to long-term development and dampening the prospects of convergence.

The conclusion here is that generative AI may not only increase within-country inequality, but because of the differential impact on developed and developing countries, it may increase between-country inequality as well. This would potentially reverse decades of declining global inequality (see here and here).

Sunday, 27 October 2024

Airlines have to pay more compensation for death or injury, but it probably still isn't enough

The value of a preventable fatality (a more palatable term than the value of a statistical life) for New Zealand was increased last year to $12.5m (see here). That is the value that Waka Kotahi New Zealand Transport Agency uses in evaluating the benefits of road safety improvements, for example. The new value was a substantial increase from the previous value of $4.88 million.

So, I was interested to read this week that the International Civil Aviation Organisation (ICAO) has revised the amount that airlines must pay in compensation in the event of a death or injury, to just $335,000. As the New Zealand Herald reported:

Travellers will be eligible for higher compensation for international flights, with the International Civil Aviation Organisation (ICAO) setting new liability limits for death, injury, delays, baggage and cargo issues.

This means airlines must pay out at least $335,000 for death or “bodily injury” on flights as a result of the review of payment levels that come into force late this year.

While liability limits are set by the international Montreal Convention agreement, there are no financial limits to the liability for passenger injury or death if a court rules against an airline.

Why is the ICAO value so much lower? After some fruitless searching, I haven't been able to find anything to say how the ICAO sets its value. It dates back to 1999, where the value was set as 100,000 SDRs (Special Drawing Rights - an international reserve asset created by the International Monetary Fund, based on a basket of five currencies).

One reason that might account for this difference is the way that the two estimates are measured. The value of statistical life for New Zealand noted above is measured using the willingness-to-pay approach. Essentially, that method involves working out how much people are willing to pay for a small reduction in the risk of death, then scaling that value up to work out how much they would be willing to pay for a 100 percent reduction in the risk of death, which becomes the estimated value of a statistical life.

An alternative is to use the human capital approach, which involves estimating the value of life as the total amount of economic production remaining in the average person's life. The value of that production is estimated as their wages. Essentially then, this approach involves working out the total amount of wages that the average person will earn in their remaining lifetime. Typically, the human capital approach will lead to a much smaller estimate than the willingness-to-pay (WTP) approach (and for an unsurprising reason - people are worth more than just the value they generate in the labour market!).

So, this difference in approach might account for the different estimates. Why might the ICAO use the human capital approach? One reason may be that the human capital approach leads to lower liability for compensation (in cases where the airline is not found to be at fault - if the airline is found by courts to be at fault, then the compensation is uncapped). Given that many airlines that belong to ICAO are national carriers, each country has an incentive to try and limit the liability of their own airline to paying compensation. A second reason is explained in Kip Viscusi's book Pricing Lives (which I reviewed here). In the book, Viscusi argues that the WTP approach is more appropriate when considering what society is willing to pay to prevent deaths (e.g. in road safety improvements), and that the human capital approach is more appropriate approach when considering a particular life (e.g. in calculating a legal penalty for wrongful death). If we believe Viscusi's argument, then the human capital approach should be used by ICAO.

However, even if we believe that the human capital approach is the right approach (and I'm not convinced that it is), it probably still underestimates the compensation that should be paid, at least for New Zealanders. Consider the following details. The median age in New Zealand is 38.1 years (at the 2023 Census). Life expectancy (at birth) is 80 years for males, and 83.5 years for females. The median weekly earnings (from wages and salaries) was $1343 in June 2024, or $69,836 per year. Using those numbers, and assuming that the median-aged person works only until age 65, and using a social discount rate of 3 percent per year, the discounted value of future wages for the average New Zealander is $1.35 million. That is more than four times higher than ICAO's figure, and is estimated using the human capital approach. Even if we used a discount rate of 10 percent, rather than 3 percent, the value is still about $715,000, more than double the ICAO value.

The ICAO is seriously understating the value of compensation that should be paid in the case of a death on a flight (and where the airline is not at fault). It's just as well that these are rare events!