Showing posts with label Survivorship bias. Show all posts
Showing posts with label Survivorship bias. Show all posts

Sunday, 30 March 2025

Another study of MasterChef that doesn't tell us much because of survivorship bias

Data from sports and games can tell us a lot about decision-making and behaviour. That's because the framework within which decisions are made, and behaviour takes place, is well defined by the rules of the sport or game. That's why I really like to read studies in sports economics, and often post about them here. I also like to read studies that use data from game shows, where the framework is clearly defined.

While those sorts of studies can tell us a lot, they still need to be executed well, and unfortunately, that isn't always the case. Consider this post, where I outlined a clear problem of survivorship bias in the analysis of a paper using data from MasterChef. Sadly, that paper is not alone as one of the authors, Alberto Chong (Georgia State University) has made a similar mistake in a follow-up paper, again using the same dataset from MasterChef.

This new paper intends to look at the relationship between exposure to anger and performance. As Chong explains:

Being exposed to anger in others may provide a burst of energy and increase focus and determination, which maybe translated into increased performance. However, the opposite may also be true. Exposure to anger in others may cloud judgment, impair decision-making, and may end up decreasing performance. In short, understanding whether the link between these two variables is positive or negative is an empirical question.

And if you've ever watched MasterChef (the US version), you will know that anger is a key feature of the series. For example:

So, Chong looks at whether exposure to the angry reactions of the judges affects contestants' performance overall, including their final placement, as well as the number of challenges they placed in the top three, their probability of placing in the top three, and their probability of winning. The dataset covers all seasons of MasterChef from 2010 to 2020. Exposure to anger is measured as "the number of times that any of the contestants have been exposed to anger by any of the judges". Chong finds that:

...people who are exposed to anger appear to react positively to anger by improving their final placement in the competition likely as a result of increased focus and determination. In particular, we find that it is associated with contestants improving around 1.5 placement positions or higher in the final standings. We also find that the probability of winning the competition increases by around 2.2 percent.

However, there is a problem, and that problem is survivorship bias. Contestants who remain in the show for longer have more opportunity to be exposed to anger from the judges. So, even if angry reactions are completely randomly assigned to contestants, those who survive for more episodes will both attract more angry reactions and have a higher placing overall. There is a mechanistic relationship that drives a negative correlation between placing in the show and exposure to anger. The analysis needs to condition the exposure to anger on the number of opportunities for judges to be angry. So, rather than the number of times exposed to anger, the key explanatory variable should be the proportion of times the contestant is exposed to anger.

Now, in my previous post on analysis of this dataset I demonstrated using some randomly generated data why survivorship bias was a problem. I'm not going to do that this time, because the issue is substantively the same (even if the specific numbers will be different). However, as I noted then, this study is crying out for a replication along with the other one, and together they would make a great project for a motivated Honours or Masters student. Then these studies might live up to the ideal of telling us something about decision-making and behaviour.

[HT: Marginal Revolution]

Monday, 22 May 2023

New economic ideas, and the rift between academic economists and government economists

On NZAE's Asymmetric Information blog yesterday, Dennis Wesselbaum reported on the latest (sixth) NZAE member survey. The theme of the survey was views on Kate Raworth's Doughnut Economics (which I reviewed here), Mariana Mazzucato's Mission Economy (which I reviewed here), and the circular economy. The results are eye-opening:

One takeaway from this survey is that academic economists disagreed that these concepts are improving economic policy analyses, and that these concepts should be taught as part of the economics curriculum at all. In the words of one respondent: “As an economics lecturer, I am loathe to introduce non-science based concepts into the curriculum. (This is the same reason I/we are reluctant to teach Modern Monetary Theory - another popular set of non-scientific econ "meme" theories)”.

However, the exact opposite is true for government economists. While the survey is non-representative for either group, this suggests an important and potentially dangerous divergence of views about the usefulness of these fringe concepts.

I'm not sure that I would go so far as to label the difference as 'dangerous', but certainly noteworthy. As one of the academic economist respondents to the survey, I thought it worth adding a bit of additional context.

Consider Question 2 from the survey: "To the extent that [each of Doughnut Economics, Mission-based economics and the Circular Economy] has been incorporated into economic policy analysis by Ministries and Agencies, it has improved the quality of that analysis." Here's the resulting graphs, summarising the survey results for that question:


On the right-hand graphs, the purple bars are academic economists' views, and it is clear that there is a strong belief that each of doughnut economics, mission economy, and the circular economy have reduced the quality of economic policy analysis. Mission economy is viewed a less negatively than the other two, and that is pretty much how I ranked them in my responses. For the government economists (the green bars in the graph), uncertain was the mode response. Aside from being uncertain, government economists tend to have positive views about the impact of these ideas on the quality of economic policy analysis.

This divergence of views makes me wonder about the underlying cause. There is probably little difference in the undergraduate economics training between government economists and academic economists, but academic economists are more likely to have a PhD. However, it seems unlikely that the PhD makes that much of a difference. In what other ways are academic economists and government economists different? Perhaps academic economists have a generally more sceptical or critical viewpoint of new ideas that do not yet have robust empirical support (which is arguably the case for all three ideas)? Or, perhaps academic economists, stuck in their ivory towers, are simply out of touch with the latest ideas? Either of those characterisations of academic economists could be accurate. On the other hand, perhaps government economists are more likely to 'toe the party line', supporting the latest fads or fashions that the government is in favour of? Or perhaps there is a survivorship bias, in that government economists who disagree with the ideas that are currently in favour depart government jobs for the private sector (or academia)? Either of those characterisations of government economists could be accurate. Or, perhaps, a mixture of all of those characterisations explains the difference in views between academic economists and government economists - that would be my guess.

We don't know the cause of the differences, but there is at least a little bit of support for the argument that academic economists are generally more sceptical about new ideas:

The final question asked whether these concepts should be taught as part of the core syllabus in economics... We again find substantial differences between academia and government respondents: academics tend to disagree, while government economists tend to agree that these concepts should be taught.

Even if we don't believe in these ideas, there is something to be said for at least exposing students to them. Students should know that these ideas exist, and are currently influential in policy circles. If we want our students to get good government economic policy jobs, and these ideas are important for economists in those jobs to understand, we should at least be teaching our students to understand them. And we should be teaching students how to critically examine these ideas, so that students (and future government economists) have realistic expectations about what should be contributing to sensible economic policy analysis.

[Update: Eric Crampton makes many of the same points here]

Thursday, 23 March 2023

MasterChef, fear of failure, and survivorship bias

In our most recent Waikato Economics Discussion Group session, we discussed this recent article by Alberto Chong (Georgia State University) and Marco Chong (Bethesda-Chevy Chase High School), published in the journal Kyklos (ungated earlier version here). They used hand-collected data from ten seasons of the US version of MasterChef (from 2010 to 2020), to investigated whether fear of failure leads to better, or worse, performance. This is an important question, as they note that:

...fear of failure is rather common and widespread in societies. In the United States, for instance, it has been estimated that around 30 % of the population is terrified of failure, and it ranks among the worst fears that the population endure in this country...

Fear of failure has previously been studied in a number of contexts, including sports, business, and education. Most of the literature finds that fear of failure reduces performance. That may be because fear of failure leads to emotional paralysis, inhibiting people from their 'usual' level of performance. On the other hand, fear of failure could lead to increased performance, by providing additional focus that leads to greater creativity and the ability to take calculated risks. MasterChef is a particularly interesting context in which to study this, because:

As the home cooks are judged by world class chefs and restaurateurs and watched by television audiences that range in the millions the potential shame and embarrassment of failing under these circumstances is very significant... The fact that the judges tend to be rather harsh with the contestants further compounds to this, more so given that these home cooks come with high self-esteem and egos, as they are typically considered as cooking luminaries in their immediate circles of friends, families, coworkers, neighbors or clubs and associations.

In other words, failure in MasterChef is very public, and directed at someone who is probably not used to failure. Chong and Chong collected data on the final ranking of 197 contestants across the ten seasons of MasterChef in their sample. They measured fear of failure as the sum of two variables:

The first is the number of times that a contestant ends up among the bottom three entries in any particular cooking challenge. The second is the number of times that a contestant ends up surviving a Pressure Test...

Chong and Chong also created a measure of 'extreme fear of failure', which was simply the number of times that a contestant survived a Pressure Test. They then look at the relationship between their measures of fear of failure and the contestants' final ranking in the season, controlling for individual characteristics, the number of rounds the contestant participated in, and a measure of positive reinforcement (made up of the number of times the contestant won a Mystery Box or elimination or team challenge, plus the number of times that they placed in the top three in a challenge). Chong and Chong find that:

...on average, individuals that are on the verge of being eliminated, but are able to survive and stay in the competition, end up doing better in the final rankings, all else being equal. In particular, we find that the higher the number of times a contestant is put in this situation, the higher his or her final placement will be among all the contestants. Overall, we find that depending on the measure used, an increase in one unit in our fear of failure index is linked to an increase of between almost one position to four positions in the final competition placement, the latter in situations of extreme fear of failure.

In other words, contestants who experience a greater fear of failure over the course of the season, perform better. Or do they?

Think carefully about how Chong and Chong measured fear of failure - the number of times that a contestant won a Mystery Box or elimination or team challenge, plus the number of times that they placed in the top three in a challenge. Contestants that last longer in the season will obviously find themselves in those situations more often than contestants who are eliminated early. So, contestants who last longer in the season will rank higher overall, and have a higher measure of fear of failure, simply because they lasted longer in the season. In other words, there is a clear survivorship bias in their analysis.

But wait! Didn't Chong and Chong control for the number of rounds that the contestant participated in? They did, and in theory that should mean that their results compare participants in terms of fear of failure, holding the number of rounds they participated in constant. However, that's not quite the case, because of the way that the number of rounds variable is included. The number of rounds variable is a measure of the exposure of each contestant to the potential for fear of failure. But that assumes that the exposure variable is linear, when it really isn't. If being at risk of elimination is randomly allocated to contestants, then there is a 15 percent (3/20) chance of being in the bottom three when there are 20 contestants, but a 50 percent (3/6) chance of being in the bottom three when there are just six contestants left. The effect of the number of rounds is clearly non-linear, but they control for a linear relationship.

To illustrate the problem here, I constructed some simulated data and ran some analyses (in Excel). I assumed that there were initially 20 contestants, and that each round, one of them was eliminated. I randomly determined which contestant was eliminated, and which two of the other contestants were in the bottom three. I ran this through until a 'final' 17th round, where there were four contestants remaining. I then calculated a measure of fear of failure (the number of times the contestant was in the bottom three), and the ranking of each contestant. Then, I ran a multiple regression model, with ranking as the dependent variable, and fear of failure as the explanatory variable, controlling for the number of rounds that the participant survived for. The outcome was that fear of failure had a coefficient of -0.350 with a p-value of 0.009 (highly statistically significant). In other words, in the data that was simulated totally at random, the result was a statistically significant relationship. It's picking up survivorship bias.

Then, instead of controlling for the number of rounds, I calculated a measure of expected exposure to fear of failure. This was the probability that a contestant would find themselves in the bottom three in a round, then summed up for each round that they participated in. When I run the same multiple regression, but controlling for expected exposure instead of the number of rounds, the coefficient on fear of failure was -0.21 and statistically insignificant (p-value of 0.38).

And, just in case anyone thinks all of this was purely coincidental, I created a whole new random dataset, and repeated the exercise. The second time, controlling for rounds the coefficient on fear of failure was -0.22, and statistically significant with a p-value of 0.035. When controlling for expected exposure, the coefficient on fear of failure was -0.27 and statistically insignificant (p-value of 0.13).

I'm sure I could run the analysis many times more and get similar results. While my data setup is not identical to theirs, it does enough to illustrate that their measure of fear of failure is susceptible to survivorship bias, because a randomly simulated dataset leads to similar results (albeit with a smaller coefficient).

Chong and Chong say that their data are available on reasonable request. This study is crying out for a replication with a better measure of fear of failure. That could either be a measure of expected exposure to fear of failure (as in my simulated dataset), or dummy variables to each number of rounds. Approaching the analysis either way (or both) would make a great Honours project for a suitably motivated student.

Friday, 12 March 2021

How not to measure the long term consequences of the Hiroshima atomic bomb

In a new article published in the Journal of the Japanese and International Economies (sorry, I don't see an ungated version online), Satoshi Shimizutani (Nakasone Yasuhiro Peace Institute) and Hiroyuki Yamada (Keio University) look at the long-term impact of the Hiroshima atomic bomb blast in 1945. They use data from the 2011-12 and subsequent waves of the Japanese Study on Aging and Retirement (JSTAR), along with a supplementary survey conducted in 2017. Their sample includes 653 people living in Hiroshima in the JSTAR sample, 297 of whom are defined as 'affected' by the Hiroshima bomb (primarily either because they were survivors of the Hiroshima bomb, or because they parents were). Comparing those two groups across a wide range of socio-economic variables, they find that:

...more than 60 years after the tragic event, survivors and their children are not seriously disadvantaged in marriage status or educational attainment but some significant distinctions between the affected and the non-affected group is observed in such aspects as combination of married couples, work status, mental health, and expectations.

Specifically, they find that affected people are more likely to have inherited their house, affected women (but not men) are more likely to be self-employed, work in small firms, and hold managerial positions, while affected men (but not women) report a lower subjective probability of living to age 85, and more depressive symptoms. There are a whole range of variables where the differences are not statistically significant.

There are a couple of problems with this study. The first is that they studied a wide range of socio-economic variables, but made no adjustment for the multiple comparisons that they made. Simply by chance, five percent of all comparisons are likely to be 'statistically significantly different' at the five percent level of significance. The fact that they don't really observe any effects that are consistent for both men and women should also be a big red flag here. Why would there be employment effects for affected women, but not men? Shimizutani and Yamada don't provide a strong theoretical reason to support their results, and fail to acknowledge the big limitation on them that the multiple comparisons creates.

The second, and probably more serious, problem is survivorship bias. The majority of the worst-affected people won't have survived to 2011, in order for their data to be included. Only the long-term survivors are included in the sample. So, at best this study can say that it looks at the long-term consequences of the Hiroshima atomic bomb, for those that survived to at least 2011. The long-term effects on the whole population affected in 1945 or thereafter were likely much greater, and much of those effects are unobserved. Again, this limitation wasn't acknowledged by Shimizutani and Yamada.

The short-term and long-term health consequences of the atomic bombings in 1945 have been extensively studied. The socio-economic consequences have received much less research attention. Unfortunately, this study doesn't do an adequate job of filling that gap.

Sunday, 29 December 2019

E-bike commuting, happiness, and survivorship bias

If I told you that e-bike commuters are happier than drivers are, would you conclude that travelling to work by e-bike made people happier? Perhaps you would, but if you did, you would be confusing correlation with causation. Perhaps happier people are more likely to commute using e-bikes? Or perhaps only higher-income people can afford an e-bike, and higher-income people are happier? It is difficult to say. However, jumping straight from the observed correlation into studying why e-bike commuters are happier shouldn't be your next step, especially not if you are going to base your study on talking to only 24 e-bike commuters.

However, that's exactly what this study published in the Journal of Transport & Health (sorry, I don't see an ungated version online), by Kirsty Wild and Alistair Woodward (both University of Auckland), did. It was covered in the New Zealand Herald earlier this year, but I held off on writing about it until I had a chance to read the research myself.

The problem isn't so much the research itself - interviewing e-bike commuters about what they like about e-bike commuting is fine. However, extrapolating that to answer the question about what should be put in place to encourage e-bike commuting, as this study does, is fraught. The reason is survivorship bias.

Almost by definition, if you interview current e-bike commuters, then you're interviewing people who tried e-bike commuting, and liked it. However, there are a bunch of people who tried e-bike commuting and hated it - they don't commute by e-bike any more, and they didn't get interviewed. In other words, the current e-bike commuters are the survivors from a larger group of people that have tried e-bike commuting at some time.

The problem in this case is that those two groups (survivors and non-survivors) are different. At the very least, the survivors like e-bike commuting, and the non-survivors don't (or, at least, they don't like it enough to continue commuting by e-bike). Interviewing the survivors tells you nothing about what the non-survivors liked or didn't like about e-bike commuting. It could be that the things that the survivors like about commuting by e-bike are exactly the things that the non-survivors hated about it. And you can't tell from this research, because former e-bike commuters (the non-survivors) were not interviewed.

So, if you decide to base decisions about cycling infrastructure on what current e-bike commuters like about it, there is no guarantee that e-bike commuting would increase as a result. The people who want to commute by e-bike with the current infrastructure are already commuting by e-bike. What you really want to know is, what do the people who don't currently commute by e-bike want?

I'd find research like this a lot more plausible if they had interviewed people who gave up on commuting by e-bike, or if they interviewed both groups. Or, even better, if they ran an experiment where they distributed e-bikes randomly to some people and asked them to use them for commuting, and then interviewed that experimental group about what they liked and did not like.

Otherwise, you simply get more of what the survivors already like, and don't necessarily create the right environment to increase e-bike commuting at all.