Saturday, 27 November 2021

How and when we develop our strategic reasoning

Some years ago, my son introduced me to the "Game of 21". The first player chooses a number (1 or 2), and then players take turns incrementing the count by 1 or 2. So, for example, if the first player chooses 1, then the second player could choose either 2 or 3, but if the first player chooses 2, then the second player could choose either 3 or 4. The winner is the player that chooses 21. My son beat me handily, but he knew the winning strategy, which is to always choose a multiple of 3, if one is available. To see why that's a winning strategy, we can work backwards from 21. If you choose 18, then your opponent must choose either 19 or 20, in which case you can choose 21. If you choose 15, then your opponent must choose either 16 or 17, in which case you can choose 18, and then 21 after their next choice. And so on (15, 12, 9, 6, and 3). It turns out that there is a clear second-mover advantage in the Game of 21, since the second player can always choose 3, regardless of what the first player chooses, and the first player can never choose a multiple of 3 if the second player does so.

The winning strategy seems obvious when it is explained to you, but it is far from obvious to most people before the game begins. How long does it take for people to figure it out, and would a shorter game (say, a "Game of 6") help? Those are the research questions that this 2010 article by Martin Dufwenberg (University of Arizona), Ramya Sundaram (George Washington University), David Butler (University of Western Australia), published in the Journal of Economic Behavior and Organization (ungated version here), tackle. Using a sample of 72 research participants, they had 42 of them pair up (in a round-robin format) for five rounds of the Game of 21 (G21) followed by five rounds of the Game of 6 (G6), and 30 of them did the reverse (five rounds of G6, then five rounds of G21). Essentially, they test whether playing the simpler G6, where recognising the multiple-of-3 winning strategy is much easier, helps players to recognise the winning strategy for G21. Indeed, that's what they find. Looking only at players who play second (the 'Green' player in their wording), since only those players have a dominant strategy, they look at the proportion of players playing a 'perfect' game.

First, Dufwenberg et al. note that:

...most subjects playing five rounds of G6 realize that G6 may be solvable by rational calculation...

...most subjects playing the Green position in G21 for the first time do not immediately figure out that choosing multiples-of-three is the best they can do... Across treatments, in G21, only 49 of 179 games (27%) are played perfectly... The rates of perfect play are especially low in the early rounds of the G21-then-G6 treatment (e.g. 2 out of 20, or 10%, in round 1).

Then, turning to their main research question, Dufwenberg et al. find that:

Green players play G21 perfectly in the G6-then-G21 treatment 37% of the time, compared to 21 percent in the G21-then-G6 treatment. This difference is significant at the 5% level (Z statistic = 2.20).

They also show that, among players who appear to have figured out the winning strategy in G21, players who played G6 before G21 choose the winning strategy earlier, on average, than those who played G21 before G6. That suggests that we can learn how to optimise in difficult strategic situations if we are first presented with similar but simpler situations.

An interesting side-point of the Dufwenberg et al. paper was the reference to level-k reasoning. Level-reasoning refers to the number of steps of reasoning a decision-maker is capable of undertaking. As they note:

...level-0 players may choose randomly across all strategies. Level-1 players assume everyone else is level-0, and best respond; level-2 players assume everyone else is a level-1 player, and best respond; etc...

G6 and G21 don't really require much in the way of steps of reasoning, because once you realise what the winning strategy is, it doesn't matter much what the other player does (unless they don't know the winning strategy).

However, one game that does test level-k reasoning is the 'beauty contest'. Each player must choose a number between 0 and 100, and the winner is the player who chooses the number that is closest to two-thirds (or some other fraction) of the average of all guesses. I played this game many times with students when I was teaching a third-year Managerial Economics and Strategy paper. Level-0 reasoning would lead to a player choosing randomly. If all players did that, then the rational choice for a Level-1 reasoning player would be to choose 33 (two-thirds of the average of 50). However, if you believed that everyone else was a Level-1 reasoning player, making you a Level-2 reasoning player, then you should choose 22 (two-thirds of 33). And, if you believed that everyone else was a Level-2 reasoning player, making you a Level-3 reasoning player, they you should choose 14 (two-thirds of 22). And so on. The Nash equilibrium here is for everyone to choose 1 (or 0, depending on how the game is scored). However, the outcome is never that the winning score is 0. From memory, the winning score in my class was always around 10-20.

How many steps of reasoning do people engage in? That question has drawn a lot of research attention. One interesting aspect is how we early in life we develop level-k reasoning. That's the topic of this new article by Isabelle Brocas and Juan Carrillo (both University of Southern California), published in the Journal of Political Economy (ungated earlier version here). Brocas and Carillo created a very simple three-player game that could be easily solved by backward induction (for those who understand some game theory). As they describe it:

...subjects were matched in groups of three and assigned a role as player 1, player 2, or player 3, from now on referred to as role 1, role 2, and role 3. Each player in the group had three objects, and each object had three attributes: a shape (square, triangle, or circle), a color (red, blue, or yellow), and a letter (A, B, or C). Players had to simultaneously select one object. Role 1 would obtain points if the object he chose matched a given attribute of the object chosen by role 2. Similarly, role 2 would obtain points if the object he chose matched a given attribute of the object chosen by role 3. Finally, role 3 would obtain points if the object he chose matched a given attribute of an extra object.

The accompanying Figure 1 in the paper helps to understand the game (although the figure is in black-and-white, and the description refers to colours, which doesn't help as much as it could!):

Player 3 is asked to match the shape, so they should choose the dark square C. Player 2 is asked to match the colour that Player 3 will choose, so they should choose the dark triangle B. Notice that Player 2 needs to undertake two steps of reasoning, working out what Player 3 is doing in order to work out what they should do. Player 1 is asked to match the letter that Player 2 will choose, so they should choose the light circle B. Notice that Player 1 needs to undertake three steps of reasoning, because they must work out what Player 3 will do and then what player 2 will d, in order to work out what they should do.

Brocas and Carillo run their experiment with a number of samples of children and young adults, and each research participant played the game 18 times (six times in each of the three positions). They expect to find:

...four types of individuals: R (subjects who always play randomly), D0 (subjects who play at equilibrium only if they have a dominant strategy), D1 (subjects who play at equilibrium when they have a dominant strategy and can best respond to a D0 type), and D2 (subjects who can play as D0 and D1, as well as best respond to D1).

And in terms of behaviour:

The predicted behavior is simple. R plays the equilibrium strategy one-third of the time in all roles, D0 always plays the equilibrium strategy in role 3 and one-third of the time in roles 1 and 2, D1 always plays the equilibrium strategy in roles 2 and 3 and one-third of the time in role 1, and D2 always plays the equilibrium strategy.

Their first study involves students from third to eleventh grade from a private school in Los Angeles, along with undergraduate students from USC. Classifying the research participants into the four types outlined above, Brocas and Carillo find that:

Subjects either recognize only a dominant strategy or always play at equilibrium. Also, some very young players display an innate ability to play always at equilibrium while some young adults are unable to perform two steps of dominance.

In other words, there are no D1 players, as every player who can reason beyond one step can reason all the way through the steps. Then, looking only at the 234 grade school students in their sample, Brocas and Carillo find that:

Performance in roles 1 and 2 increases significantly up to a certain age (around 12 years old), and then stabilizes...

So, older students perform better, but only up to the age of 12 years. They then go on to replicate similar findings for a Los Angeles public school (where overall performance was lower) - there is no difference in performance from sixth to eighth grade (12-14 years old).

Finally, Brocas and Carillo study a sample of students from kindergarten to second grade. They simplify the game so that there are only two players (rather than three), and only two attributes (rather than three). With their sample of 117 children, they find that:

Equilibrium behavior is not significantly different between K and grade 1 in roles 2 and 3, and they are both lower than in grade 2 (p < .02, FDR adjusted).

The evolution of strategic behaviour as people age is interesting. That isn't quite what Brocas and Carillo are studying, since they don't follow the same children over time, but instead look across cohorts of different ages. However, it's hard to see how or why there would be a cohort effect here, so possibly they are observing an age effect. Interestingly, most of the improvement in strategic reasoning happens between the ages of 8 and 12 (second to sixth grade), and there is little improvement after that. That doesn't quite accord with the USC students performing better than the 11th-graders, so perhaps we need to know a little bit more about the evolution of strategic reasoning among older adolescents. However, in relation to younger children, Brocas and Carillo note that:

Existing research shows that by 7 years of age children may think ahead and form correct anticipations... Children have also been shown to develop inductive logic between the ages of 8 and 12...

Those are the sorts of skills that are used in developing level-k reasoning, so the mechanisms underlying the increase in strategic reasoning between ages 8 and 12 seem plausible. However, this clearly needs to be unpacked a bit more, and that would be a fruitful avenue for future research.

Overall, these two studies help us to understand a little bit more about how (and when) our reasoning in strategic games develops.

[HT: My colleague Steven Tucker for the Dufwenberg et al. study; Marginal Revolution for the Brocas and Carillo study]

Friday, 26 November 2021

Simulation evidence that alcohol minimum pricing is better than increasing excise tax

If alcohol is too cheap (see this post), then the two main policy options that the government has is to increase alcohol excise tax (which would increase the price of all alcoholic drinks), or to introduce a minimum unit price (which would increase the price of cheap alcoholic drinks, but probably leave more expensive options unchanged in price). Which is better?

On that topic, I just read this 2010 article (open access) by Robin Purshouse (University of Sheffield) and colleagues, published in the prestigious median journal Lancet. They constructed a complex simulation model from cross-sectional consumption survey and alcohol purchase data (differentiating between on-premise and off-premise purchases, and type of beverage), as well as health data, for England. Importantly, they disaggregate the effects of changes in price on groups based on the level of drinking: moderate (including non-drinkers); hazardous; and harmful. This seems to me to be one of the most thorough exercises of this type that I have seen. The most obvious flaw is the use of cross-sectional data, where longitudinal data would provide better estimates of the own-price and cross-price elasticities of the various beverage types.

They investigate a wide range of pricing policies, with different levels of change in price. Their model allows them to estimate the effects on alcohol consumption (based on own-price and cross-price elasticities), and the effects on health care costs (based on health economic models) and health gains measured in Quality-Adjusted Life Years (QALYs; based on econometric models linking consumption to alcohol-attributable medical conditions). Their findings are most easily summarised in Figure 1 from the article:

Unsurprisingly, within any type of policy, larger increases in price have more positive effects. However, the more interesting result is comparing across different policies. Purshouse et al. find that:

...notable between-policy differences exist. For example, a £0·45 minimum price would be more effective overall than a 10% general price increase, but is achieved with a much lessened effect on moderate drinkers’ spend and larger increases in spend for harmful drinkers. This differential effect arose because minimum price policies target cheap alcohol products, which make up a higher proportion of the average selection of alcohol purchases for heavier drinkers than for moderate drinkers.

So, policies that have the same overall effect on alcohol consumption can have very different effects in terms of reducing alcohol-related harm. My takeaway from the results overall is that it appears that minimum unit prices work better than increasing prices across-the-board through excise tax increases. This would accord with other research, although it is not a reason to discard excise taxes entirely.

Understanding the effects of potential policy options is important. In Purshouse et al.'s discussion of their results, they make what seems to me to be a really important point:

For policy makers, a balance between reduction in health harms and increased consumer spending might be important for proportionality, and one implication of our study is that minimum pricing strategies might help achieve this balance. For example, a general 10% price rise is estimated to reduce consumption by 4·4% and alcohol-related harm by £3·5 billion over 10 years, but a minimum price of £0·45 could produce a similar overall consumption effect, while achieving greater reductions in harm and a rebalancing of spending effect away from moderate drinkers towards heavier drinkers.

So often, public health researchers ignore the trade-offs inherent in their policy recommendations, or lack any sense of the proportionality of those recommendations. The sort of simulation exercise that Purshouse et al. conducted allows for quite a deep exploration of various pricing policy options. They make the point that their modelling approach can be used as a template for other countries. It would be great to pull together something like this for New Zealand, which might provide the evidence to support minimum unit pricing here.

Read more:

Wednesday, 24 November 2021

Coronavirus lockdowns and educational inequality in German high schools

The rapid shift to online learning affected schools (and teachers, and students) at all levels. Some schools (and teachers) were better prepared than others, having resources that were more easily adapted to online teaching modes. Some students were better prepared than others, having access to devices and stable internet connections, in order to more fully participate in online learning. The unfortunate thing is that the students who had the lowest access to online learning are likely to be those who were already under-achieving. At least, that is the headline result from this new article by Elisabeth Grewenig (Leibniz Center for European Economic Research) and co-authors, published in the journal European Economic Review (ungated earlier version here).

Grewenig et al. use data collected from 1099 parents of school-aged (i.e. not university) students in Germany, collected as part of the ifo Education Survey. The survey collected data on students' time use, both during June 2020 (when lockdowns were in effect and there was basically no in-person teaching), and retrospectively for the period before the coronavirus pandemic. Time use was separated into several categories: (1) school-related activities (school attendance; or learning for school); (2) activities 'deemed conducive to child development' (reading or being read to; playing music and creative work; or physical exercise); and (3) 'activities deemed generally detrimental' (watching television; gaming; social media; or online media); and (4) relaxing.

Comparing students' time use during the lockdowns with their time use before the pandemic, Grewenig et al. find that:

...the school closures had a large negative impact on learning time, particularly for low-achieving students. Overall, students’ learning time more than halved from 7.4 h per day before the closures to 3.6 h during the closures. While learning time did not differ between low- and high-achieving students before the closures, high-achievers spent a significant 0.5 h per day more on school-related activities during the school closures than low-achievers. Most of the gap cannot be accounted for by observables such as socioeconomic background or family situation, suggesting that it is genuinely linked to the achievement dimension. Time spent on conducive activities increased only mildly from 2.9 h before to 3.2 h during the school closures. Instead, detrimental activities increased from 4.0 to 5.2 h. This increase is more pronounced among low-achievers (+1.7 h) than high-achievers (+1.0 h). Taken together, our results imply that the COVID-19 pandemic fostered educational inequality along the achievement dimension.

So, low-achieving students (defined as those in the bottom half of the grade distribution for this sample for German and mathematics combined) reduced their study time by more than high-achieving students. To the extent that study time leads to greater academic achievement, this can only lead to an increase in the disparity in academic performance between students at the top and those at the bottom. The really disheartening finding though was that:

...only 29% of students on average had online lessons for the whole class (e.g., by video call) more than once a week. Only 17% of students had individual contact with their teacher more than once a week... The main teaching mode during the school closures was to provide students with exercise sheets for independent processing (87%)... although only 37% received feedback on the completed exercises more than once a week...

The distance-teaching measures over-proportionally reached high-achieving students. Low-achievers were 13 percentage points less likely than high-achievers to be taught in online lessons and 10 percentage points less likely to have individual contact with their teachers... Low-achievers were also less likely to be provided with educational videos or software and to receive feedback on their completed tasks.

If you thought that teachers, having scarce online teaching time available, would prioritise the low-achieving students, perhaps because the high-achieving students are more self-motivated and/or have better learning support through their parents, you would be sorely mistaken. That strongly suggests to me a failure in the way that German teachers were supported in their rapid shift to online teaching activities, since it was entirely foreseeable that low-achieving students would be more greatly affected by the changes. Alternatively, supporting the low-achieving students in low socioeconomic families to have better access to online resources would no doubt have helped as well (although New Zealand's experience suggests that something more proactive than simply having support or resources available for those who ask for it is required).

One issue with this research is the use of retrospective recall about students' time use from the period before coronavirus. Grewenig et al. argue that the degree of social desirability bias is low, and that the results are similar to those from the German Socioeconomic Panel (GSEP), where students report their own time use. However, comparing those two sources, it is clear that the reported number of hours of school-related activities before coronavirus is much higher in this sample than in the GSEP. That needn't be a problem, unless the disparity differs between parents of high-achieving students and parents of low-achieving students. Presumably, all parents are roughly equally able to observe their children's time use during lockdown. That probably is less likely of the period before coronavirus. If parents of low-achieving students are more likely to overestimate the number of hours of school-related activities than parents of high-achieving students, then that would bias the results towards showing a bigger decline in school-related activities for low-achieving students. Since those students are low-achieving, it is entirely plausible that they usually spend less time on school-related activities than their parents think they do. Unfortunately, there is no way to easily identify whether that is a problem in this sample.

With that caveat in mind, this study does point to an issue that we should be concerned about, which is how the pandemic has affected student learning, and in particular whether it has increased educational inequality. Hopefully, this is not a general result that extends beyond the German schooling system, but unfortunately it seems likely that it is.

Tuesday, 23 November 2021

The lockdown 'baby boom' in proper context is anything but a boom, and possibly not even related to the lockdowns

I was interested to read this New Zealand Herald article this week:

We've all heard the jokes about how lockdown leads to a "baby boom" - but it turns out being stuck at home does lead to a rise in birth rates.

New information from Stats NZ for the year ending in September 2021 confirms an increase in live births compared to the same time last year.

The data reveals there were 59,382 live births registered in Aotearoa, an increase from 57,753 last year.

And the fertility rate has risen slightly as well, sitting at 1.66 births per woman, up from 1.63 at the same time in 2020...

Significantly, the number of live births as at September 2021 is the highest since 2015 - long before the pandemic changed all of our lives and lockdown was the last thing on anyone's mind.

This was a little bit of a surprise, as the recent births data has shortly historically low birth rates in New Zealand. So, a 'baby boom' would come as a surprise. However, when we actually look at the data, we find that calling it a 'boom' is a mischaracterisation. Here's the data on the raw number of births by quarter in New Zealand, from 1991 to 2021 [*]:

The number of births per quarter fluctuates between about 13,500 and 16,500. There was a bit of a downward trend from 1991 to 2003, then an uptick, before the downward trend resumed from about 2009. You can see the recent rise in births at the end of the series. Indeed, the number of births is at its highest level since 2015. You might even convince yourself that this constitutes a 'baby boom'. However, then you'd also need to believe there was a boom from 2007 to 2011, where the number of births per quarter was mostly at or above the number in Q3 of 2021.

There is a problem with looking at the raw number of births though, and that is that it doesn't account for the size of the population. Population has grown a lot over the 30-year timespan shown in the graph above. To account for that, I calculated the number of births per 100 women aged 15-49 years (you can call this the period fertility rate; I use the rate per 100 women, because that makes the numbers a bit easier to interpret). [**] Here's the result for New Zealand as a whole, since 1996:

The trends are somewhat similar to the previous graph, although the overall downward trend is much more obvious. The recent increase in the birth rate is still apparent, but by itself there isn't much to suggest a 'baby boom', maybe just a slight reversal of the recent trend. As you can see, the rate is lower than it was in 2017, and for basically the entire period prior to 2013. It remains to be seen whether the increase in birth rate in Q3 of 2021 is a brief spasm in the data (similar to Q2 of 2015), or the start of a change in fertility trends. My intuition is that it is the former.

So, was this increase in births caused by lockdown? It is easy to speculate that it is, given the timing. However, we can do a little better than that. Auckland has suffered from longer periods of lockdown than the rest of the country. So, if there is a baby boom driven by lockdowns, it's likely that it would be more apparent for Auckland than for the rest of the country. That isn't what we see though. Here's the birth rates for Auckland over the period since 1996:

That doesn't look much different to New Zealand as a whole (which isn't a surprise - more than a third of the New Zealand total is contributed by Auckland). Also, if we compare the change in the birth rate across regions between Q3 of 2019 and Q3 of 2021 (I chose 2019 for the comparison, since it is the most recent year with no effect of coronavirus or lockdowns), we see this:

The biggest increase in the birth rate between 2019 and 2021 has been in the Tasman and Gisborne regions. Auckland barely features at all. The birth rate in Q3 of 2021 is actually lower in Wellington and the West Coast than it was in Q3 of 2019. It's hard to make a case that lockdowns are a cause for the increase in births, unless you can somehow make the case that the lockdowns had a bigger effect in Tasman and Gisborne, and smaller in Wellington and the West Coast (or, you can show that there is some other socio-demographic or economic effects that are able to explain the cross-region differences that seem to more than offset any impact of lockdowns). Now, you could argue that the biggest difference in the effects of lockdown between Auckland and the rest of the country is actually happening now, and so the regional differences in the effect of lockdowns should become apparent in the births data for Q1 of 2022. I guess we will wait and see for that.

So, while there has certainly been an increase in births, it is hardly a 'baby boom' (unless you have an extraordinarily liberal interpretation of what constitutes a boom). And, it is hard to make a case that it was caused by lockdown (unless lockdown and other socio-demographic or economic changes affected birth rates in different regions in some idiosyncratic way, such that Auckland ended up having a very low increase in the birth rate).

*****

[*] The data come from Statistics New Zealand, Infoshare.

[**] The calculations here are based on the births data from Infoshare, plus subnational population estimates for each region, from NZ.Stat. As the data are only provided for 30 June of each year (and only since 1996), I take the 30 June population as the denominator for the rates for Q3 of each year, and use a linear interpolation between each Q3 value to obtain population estimates for the other quarters.