Showing posts with label Risk aversion. Show all posts
Showing posts with label Risk aversion. Show all posts

Tuesday, 1 April 2025

The emerging debate on Oprea's paper on complexity and Prospect Theory

Late last year, an article in the American Economic Review by Ryan Oprea caught my attention (and I blogged about it here). It purported to show that the key experimental results underlying Prospect Theory may in part be driven by the complexity of the experiments that are used to test them. These were extraordinary results. And when you publish a paper with extraordinary results, that could potentially overturn a large literature on a particular theory, then those results are going to attract substantial scrutiny. And indeed, that is what has happened with Oprea's paper.

The team at DataColada, most well-known for exposing the data fakery of Dan Ariely and Francesca Gino (and the resulting lawsuit, which was dismissed), have a new working paper, authored by Daniel Banki (ESADE Business School) and co-authors, looking at Oprea's results (see also the blog post on DataColada by Uri Simonsohn, one of the co-authors). To be clear before I discuss Banki et al.'s critique, they don't accuse Oprea of any misconduct. They mostly present an alternative view of the data and results that appears to contradict key conclusions that Oprea finds in his paper. Oprea has also provided a response to some of their critique.

I'm not going to summarise Oprea's original paper in detail, as you can read my comments on it here. However, the key result in the paper is that when presented with risky choices, research participants' behaviour was consistent with Prospect Theory, and when presented with choices that involved no risk at all but were complex in a similar way to the risky choices ('deterministic mirrors'), research participants' behaviour was also consistent with Prospect Theory. This suggests that a large part of the observed results that underlie Prospect Theory may arise because of the complexity of the choice tasks that research participants are presented with.

Banki et al. look at a number of 'comprehension questions' that Oprea presented research participants with, and note that:

...75% of participants made an error on at least one of the comprehension questions, such as erroneously indicating that the riskless mirror had risk.

Once the data from those research participants is excluded, Banki et al. show that research participant behaviour differs between lotteries and mirrors for the research participants who 'passed' the comprehension checks (by getting all four of the comprehension questions correct on their first try). This is captured in Figure 2 from Banki et al.'s paper:

The two panels on the left of Figure 2 show the results for the full sample, and notice that both lotteries (top panel) and mirrors (bottom panel) look similar in terms of results. In contrast, when the sample is restricted to those that 'passed' the comprehension checks, the results for lotteries and mirrors look very different. Which is what we would expect, if research participants are not 'fooled' by the complexity of the task.

Banki et al. provide a compelling reason why the results for the research participants who failed the comprehension checks looks the same for lotteries and mirrors: regression to the mean. As Simonsohn explains in the DataColada blog post, this arises because of the way that a multiple-price list works:

When the dependent variable is how much people value prospects, regression to the mean creates spurious evidence in line with prospect theory. When people answer randomly for 10% chance of $25, they overvalue it, because the “right” valuation is $2.50, and the scale mostly contains values that are higher than that. When people answer randomly for 90% chance of $25, they undervalue it, because the “right” valuation is $22.50 and the scale mostly contains values that are lower than that. Thus, random or careless responding will produce the same pattern predicted by prospect theory.

Oprea responds to both of these points, noting that:

...a range of imperfectly rational behaviors including noisy valuations, anchoring-and-adjustment heuristics, compromise heuristics and pull-to-the-center heuristics will all tend to produce prospect-theoretic patterns of behavior simply because of the nature of valuation. BSWW offer this possibility as an alternative to the Oprea (2024)’s account of his data, but in fact these are examples of exactly the types of cognitive shortcuts Oprea (2024) was designed to study.

In other words, Banki et al.'s results don't refute Oprea's results, but are very much in line with Oprea's. One thing that Oprea does take issue with is Banki et al.'s use of medians as the preferred measure of central tendency. Oprea uses the mean, and when reanalysing the data with the same exclusions as Banki et al., Oprea shows that the mean results look similar to the original paper. So, Banki et al.'s results are not simply driven by excluding the research participants who failed the comprehension checks, but also by switching from using the mean to using the median.

On that point, I'm inclined to agree with Banki et al. The median is often used in experimental economics, because it is less influenced by outliers. And if you look at Oprea's data, there are a lot of large outliers, which become quite influential observations when the mean is used as the summary statistic. However, the outliers are likely to be the observations you want to have the smallest effect on your results, not the largest effect.

Oprea also critiques Banki et al.'s interpretation of the comprehension questions. Oprea rightly notes that:

...it is important to emphasize that these training questions weren’t designed to measure beliefs (e.g., payoff confusion), and because of this they are poorly suited to the task BSWW repurpose it for, ex post. Indeed, evidence from the patterns of mistakes made in these questions suggests that overall training errors largely serve as a measure of the cognitive effort (an important ingredient in Oprea (2024)’s account) subjects apply to answering these questions, and that BSWW therefore substantially overestimate the level of payoff confusion with which subjects entered the experiment.

In other words, the 'comprehension questions' are not comprehension questions at all, but they are really 'training questions' that were used to train the research participants to understand the choice tasks that they would be presented with. And so, using those training questions overall as a measure of understanding misses the point, and seriously underestimates the amount of understanding of the task that research participants had by the time they had completed the training questions.

Oprea's response is good on this point. However, if the training questions had really done a good job of training the research participants, then all participants should have had a similar level of understanding by the end of the training questions, and there should be no detectable differences in behaviour between those with more, and those with fewer, 'failed' training questions. That wasn't the case - the behaviour of the research participants who made errors in training was much more likely to be the same for lotteries and mirrors than was the behaviour of research participants who made no errors. To clear this up, it would have been interesting to have research participants also complete 'comprehension questions' at the end of the experimental session, to see if they still understood the tasks they were being asked to complete. At that point, those failing the comprehension questions could be dropped from the dataset.

One point of Banki et al.'s critique that Oprea hasn't engaged with (yet, although he promises to do so in a future, more complete response), is their finding that a larger than 'usual' proportion of the research participants fail 'first order stochastic dominance' (FOSD). A failure of FOSD in this context means that a research participant valued a lottery (or mirror) lower than a similar lottery that was strictly better. For example, valuing a 90% chance of receiving $25 less than a 10% chance of receiving $25 is a failure of FOSD. Banki et al. show that:

We begin by examining G10 and G90. Violating FOSD here involves valuing the 10% prospect strictly more than the 90% one. Across all participants (N = 583), 14.8% violated FOSD for mirrors, and 13.9% for lotteries. These rates are quite high given that the prospects differ in expected value by a factor of nine.

Those failure rates are much higher than for other similar research studies. Banki et al. note an overall rate of 20.8 percent in the Oprea results, compared with an average of 3.4 percent across eight other highly cited studies. It will be interesting to see how Oprea responds to that point in the future.

This is an interesting debate so far. Oprea does a good job of summing up where this debate should probably go next:

Ultimately, however, these questions and ambiguities can only be fully resolved by further research. While BSWW’s critique has not convinced me that the interpretation offered in Oprea (2024) is mistaken, I am eager to see new experiments that deepen, alter, or even overturn this interpretation. First, concerns that the Oprea (2024)’s results are a consequence of the design being too confusing to yield insight can only really be resolved one way or another by followup experiments that vary his procedures, instructions and other design choices in such a way as to satisfy us that the Oprea (2024) results are (or are not) overfit to that design.

Indeed, more follow-up research is needed. Prospect Theory hasn't been overturned, yet (and as I noted in my earlier post, it is consistent with a lot of real-world behaviour). However, now we know that it may be vulnerable, and Oprea's paper provides a starting point for testing more thoroughly how much of the experimental results arise from complexity.

[HT: Riccardo Scarpa]

Read more:

Wednesday, 15 January 2025

Lab experimental vs. real-world measures of risky choice

Before Ryan Oprea caused us to question all lab experimental measures of risky choice behaviour (as noted in yesterday's post), one main concern about experiments in the lab was whether they accurately reflected real-world decisions. The news is somewhat mixed, as I've written about before (see here and here). And that is quite aside from concerns that experimental subject pools made up of (usually undergraduate) students are not representative of real-world populations (see here).

Nevertheless, the last somewhat hopeful point in yesterday's post suggested that real-world behaviour might still be consistent with the experimental results (even if we cannot believe the experiments). On that note, I recently read this 2016 article by Arjan Verschoor, Ben D’Exelle, and Borja Perez-Viana (all University of East Anglia), published in the Journal of Economic Behavior and Organization (open access). They compare a measure of risk preferences estimated from a 'lab-in-the-field' experiment among over 800 farmers in rural Uganda, with measures of risk taking based on their agricultural choices.

Verschoor et al. compare two types of decisions. The first they term 'narrow-bracketed', where:

...the decision-maker does not consider their consequences together with the consequences of other decisions)...

So, narrow-bracketed decisions are those that are quite independent of other decisions, and therefore simpler. In contrast, a decision that is interdependent with other decisions, which Verschoor et al. liken to "portfolio management", is more complex.

Comparing the results of the lab experiment with real-world decisions that are narrow-bracketed (a fertiliser purchase decision) and not (the decision on whether or not to grow and sell crops for the market), Verschoor et al. find that:

Controlling for other determinants of risk-taking in agriculture, we find that risk-taking in the experiment is associated with the relatively straightforward investment decision of fertiliser purchase. However, for more involved livelihoods strategies that call not only on willingness to take risks but also on other attributes of entrepreneurship, viz. moving away from subsistence farming to growing crops for the market (measured in two alternative ways), we find no evidence of an association with risk-taking in the experiment. By contrast, a hypothetical willingness to take large-scale risks, elicited through a questionnaire, is associated with both fertiliser purchase and growing crops for the market (however measured), suggesting that this is a better proxy for entrepreneurship broadly defined.

In other words, when the real-world decision was reasonably straightforward, and reasonably independent of other decisions ('narrowly bracketed') the farmers behaved in line with their risk preferences measured in the experiment. However, when the real-world decision was more complex and inter-related with other decisions, there is little association with the risk preferences measured in the experiment. Verschoor et al. note that:

The decision to buy fertiliser is a straightforward investment decision that raises both the expected profit and the spread of possible profits within an existing livelihoods strategy... Decisions to grow cash crops or to grow for the market more broadly, on the other hand, are complex, multi-dimensional decisions that invoke not only risk preferences but also the nebulous notion of entrepreneurship.

It is interesting to think about these results alongside those in the Ryan Oprea research I discussed yesterday. Oprea found that risk preferences consistent with behavioural economics (specifically Prospect Theory) only arose because of the complexity of the experimental task used to measure them. Verschoor et al. note in the discussion of their underlying theoretical model that:

...prospect theory, which only considers changes to wealth relative to a reference level, correspondence between the two domains [risky choices in lab experiments and risky choices in real life] is assured provided both are narrowly bracketed...

Combining the two sets of results (from Verschoor et al. and Oprea), I infer that perhaps narrowly-bracketed decisions, which nevertheless involve complex choices in the lab, may show preferences consistent with Prospect Theory. Real-world decisions that are narrowly bracketed demonstrate similar preferences. An open question is whether the real-world decisions are consistent with Prospect Theory because of the complexity of those decisions. Verschoor et al. don't give us any steer on that (and neither should they, given that their article was published eight years before Oprea's).

When you move beyond narrow bracketing, adding interdependence with other decisions and therefore even more complexity, then decisions are not consistent with the lab-estimated risk preferences. We don't know whether that makes them inconsistent with Prospect Theory, but I expect it does. Does that mean that adding even more complexity makes decision-makers more rational? Or perhaps, more salience of the decision (or higher stakes of the decision) makes them more rational? It's not simply the real-world context, or the fertiliser decisions would also be affected.

Now of course, I am trying to reconcile just two sets of results from a much larger literature here, and probably going too far in doing so. However, it will be interesting to see where the lab experiments vs. real-world research literature goes next.

Read more:

Tuesday, 14 January 2025

Are experimental measures of loss aversion and behaviour under risk just an artefact of complexity?

Loss aversion has been under fire in the economics literature recently (see here and here). As one of the foundations of behavioural economics, this is a big deal. So, I was interested to read this recent paper by Ryan Oprea (University of California, Santa Barbara), published in the journal American Economic Review (ungated earlier version here). Oprea essentially tests the key tenets of Prospect Theory, that when faced with a risky choice such as a lottery, people are risk averse when it comes to gains, but risk seeking when it comes to losses. Oprea's argument is that we observe that behaviour in lottery experiments, not because it is real, but because it is an artefact of the complexity of the lotteries that the research participants are faced with.

Here's what Oprea did:

In each task in our experiment, we elicit subjects’ dollar valuations for a set of 100 “boxes,” each of which contains some dollar amount. For example, in one of our tasks (called G90), we ask subjects to value a set consisting of 90 boxes that each contain $25 and 10 boxes that each contain $0. Acquiring a set of boxes influences the subject’s earnings in the experiment according to a payoff rule, and we compare how subjects value these sets under two contrasting payoff rules.

By opening one of the boxes from the set at random and paying the subject the amount inside, we turn the set into a lottery (i.e., G90 becomes a risky prospect of earning $25 with probability 0.9), and the dollar value the subject attaches to it becomes a certainty equivalent: the certain dollar amount the subject judges to be equivalently valuable to the risky lottery.

Using those results, Oprea replicates the key results from Prospect Theory, which he refers to as the 'fourfold pattern' of risk (a term that actually comes from Kahneman and Tversky), as well as loss aversion. Then:

Our contribution is to compare these valuations to the valuations of what we call “deterministic mirrors” of the same lotteries. A deterministic mirror of a lottery consists of the same set of 100 boxes used to describe the lottery but is characterized by a different payoff rule: instead of paying the dollar amount in one of the 100 boxes selected at random as a lottery does, a mirror pays the sum of the rewards in all of the boxes, weighted by the total number of boxes. Thus, instead of paying $25 with probability 0.9 (as a lottery does), the mirror of G90 pays 0.9 × $25 = $22.50 with certainty.

In other words, the 'deterministic mirror' of a lottery retains all of the complexity associated with the choice, but eliminates all of the risk (because the amount received is certain, rather than risky). So, if the 'fourfold pattern' is real and arises from the riskiness of the lottery, it should disappear in these experiments. Instead, using data from 673 research participants (and with similar results in a second sample of 489 research participants):

...we find that

(i) The fourfold pattern arises in the valuations of deterministic mirrors just as it does in lotteries, and with roughly the same strength. Importantly, this means that we find strong evidence of what is usually called “probability weighting” in settings without probabilities.

(ii) Loss aversion arises in deterministic mirrors even though at the relevant margins they cannot actually produce losses. Thus, we find strong evidence of what is usually called “loss aversion” in settings without risk of loss.

(iii) Across subjects, the severity of each of these anomalies in lotteries is strongly predicted by their severity in deterministic mirrors, suggesting that the behaviors in the two settings are strongly linked, deriving from a common behavioral mechanism (which, clearly, cannot be grounded in risk or risk preferences).

In other words, Oprea finds strong evidence that it is complexity that drives the 'fourfold pattern' of risk in lottery experiments, because when risk is removed (but complexity remains), the 'fourfold pattern' is still there. On top of that, loss aversion remains even when there is no risk of loss. So, loss aversion may also be an artefact of complexity of lottery experiments. Oprea concludes that:

First, theories of risk preferences designed to explain these anomalies (e.g., prospect theory) are unlikely to contain much normative content and therefore should not be accommodated in the inference of welfare or the design of policy. Second, our finding of systematic departures from neoclassical benchmarks in perfectly deterministic settings suggests that many of our descriptive theories of preferences for risk are really descriptive theories of the way people evaluate complex things.

That's a really nice way of saying that behavioural economists may need to reconsider some of their key theories, because the lab experiments they have been using to verify them do not stand up to this scrutiny. And Oprea's results may also help to explain some of the recent anomalies in the loss aversion literature (see here and here).

Oprea's results are important, and even though the working paper version of this article has already been cited over 50 times, I still don't think this research has received the attention that it deserves (and see Eric Crampton's take here). However, it may not be time to throw away behavioural economics or loss aversion entirely. Oprea notes that:

We do not claim, for instance, on the basis of these data that risk preferences or even loss preferences do not exist but only that they are unlikely to be reliably revealed in lottery valuations.

That is an important caveat. Behavioural economists may simply need to find a new way of demonstrating the 'fourfold pattern' of risk, and loss aversion, without resorting to complex lotteries. These effects may still be real. After all, there is a lot of real-world behaviour that is very consistent with loss aversion (see my various posts on that topic here).

Read more:

Tuesday, 1 March 2022

The non-effect of images of half-naked women on economic behaviour

You can get away with a lot of crazy things in a lab experiment (tempered somewhat by the fact that whatever you propose has to be approved by an ethics committee). For example, this 2018 paper by Evelina Bonnier (Stockholm School of Economics) and co-authors describes a lab experiment where research participants were shown advertisements, of three kinds:

...a treatment where advertisements contain half-naked women dressed in bikini or underwear, a treatment where advertisements contain fully dressed women, or a control condition where there are no women present.

Participants were then tested on their risk taking preferences, their willingness to compete, and the mathematics performance. Specifically:

Risk taking is measured from having participants complete two multiple-price lists in which they face a series of decisions between two lotteries, where one is more risky than the other. Participants make 20 such choices and our measure of risk taking is defined as the fraction of more risky lottery choices. The competitiveness task... participants in a first stage perform a math task and get paid according to a piece-rate scheme. In a second stage, participants perform a similar math task and are paid according to a competitive winner-takes-all tournament scheme. In a third stage, participants get to choose between the two payment schemes before performing the task for a third time. Willingness to compete is measured from this binary choice. As a measure of math performance, we use the number of correctly solved math problems in the first non-competitive stage of the competitiveness task. 

The research participants were 648 students (331 men and 317 women) at the University of Copenhagen or the University of Valencia. Bonnier et al. find that there are:

...no treatment effects on any of the three main outcome measures for female participants. For men, we also find no effect on math performance or willingness to compete, but suggestive evidence that men take more risk after having been exposed to images of half-naked women... Moreover, we do not find a significant difference between the half-naked treatment and the fully dressed treatment on risk taking for men.

Risk taking was 4.8 percentage points higher among men who saw the images of half-naked women compared with the control group who saw no women (47.1% vs. 42.8%), but this difference was only statistically significant in some (but not all) analyses. That probably reflects that the real effect (if any) is quite small, and so you would need a much larger sample in order to robustly find any effect of half-naked pictures on men's risk taking behaviour. That there was no statistically significant effect for women probably accords with expectations, although risk taking was about 4 percentage points higher for women who saw images of fully dressed women than the control group women. Interestingly, although this difference was statistically insignificant, it was larger than the difference for men.

So, more research is needed on this topic, if we really wanted to know how exposure to images of half-naked women affects economic behaviour. And, of course, as Bonnier et al. note in their conclusion, in the interests of gender equity:

...it would be interesting to explore the effects of exposure to images of half-naked men.

[HT: Marginal Revolution, back in 2018]

Monday, 3 June 2019

What Jeopardy and Junior Jeopardy can tell us about gender differences in risk taking

Game shows are fun, and funny. As a bonus, they can provide a window into the contestants' decision making in a setting where the rules are known (and if you haven't already seen it, you should check out the show Golden Balls that I blogged about here). And they can provide data that economists can exploit to understand that decision-making.

That is exactly what this 2017 article (ungated) by Jenny Säve-Söderbergh and Gabriella Sjögren Lindquist (both of Stockholm University), published in The Economic Journal, does. Säve-Söderbergh and Sjögren Lindquist (hereafter SSSL) use data from the Swedish edition of the game show Jeopardy and Junior Jeopardy to investigate gender differences in risk taking, and the influence of the gender composition of the other contestants. Specifically, they are looking at whether women (and girls) make different decisions when competing against men (and boys) than when competing against other women (and girls). They also look at whether the differences are the same for adults (in Jeopardy), as for 10-11 year old children (in Junior Jeopardy).

Specifically, they look at what happens when the contestants receive a Daily Double, where they have the option to wager some of their current score on getting the answer (or, since this is Jeopardy, the question) right. Using data from 2000 and 2001 (206 shows of Jeopardy, with 449 contestants) and from 1993-2003 (85 shows of Junior Jeopardy, with 222 contestants), they find that:
...there is no gender gap in wagering among children, in contrast to the results for adults. This result is robust to controls for absolute performance, the difficulty level of the questions, experience, relative performance and performance feedback, in addition to whether children shared the game earnings with their classes. Our second finding is that male and female risk taking differ with age in different ways: whereas girls wager more than women, boys wager less than men.
That in itself is interesting. Girls aged 10-11 years are more risk-takers than boys of the same age, but this reverses among adults. It has been established in many studies that men are less risk averse than women (although those findings are contested), but girls being less risk averse than boys is a surprise.

SSSL then go on to find that:
...female behaviour is sensitive to social context. In particular, despite the high-stakes setting and the lack of strategic advantage created by providing incorrect answers, girls perform worse (answering the Daily Double incorrectly more often and winning less often) and employ less gainful wagering strategies when they are randomly assigned a group of boy opponents compared with when they are randomly assigned a same-gender group of opponents or a mixed-gender group of opponents... Conversely, women wager less if they are randomly assigned a group of male opponents... The performances of boys and men do not change with the social context...
So, essentially boys and men don't seem to care who they are playing against, but girls and women do. And it seems that the social context affects girls more than adult women, since it impacted girls' success in getting the Daily Double question right. SSSL interpret this as potentially showing stereotype threat (where girls perform worse than boys when there is a belief that on average, girls will perform worse than boys). To support this, the authors note that their results may:
...reflect feelings of intimidation in the presence of boys that therefore causes girls to be prone to making mistakes.
It is definitely concerning if girls as young as 10 are being affected by stereotype threat. You could put this down to being just one study in one particular (and fairly unique) context. However, it seems that there is a long history of studies that have identified stereotype threat among children (see here for an overview). If we're concerned about gender gaps among adults, and among university students, then it appears that solutions need to start from a much younger age.

Friday, 19 January 2018

Professional tennis players are optimisers

With plenty of action in Melbourne at the Australian Open this week, it seems timely for me to write a post about tennis. I've already noted in an earlier post that tennis players appear to be loss averse. But are they optimising nonetheless? Do they make decisions that maximise their chances of winning (which would also be consistent with loss aversion)?

A recent paper by Jeffrey Ely (Northwestern University), Romain Gauriot (University of Sydney), and Lionel Page (Queensland University of Technology), published in the Journal of Economic Psychology (sorry I don't see an ungated version) provides us with some answer. The authors look specifically at the risk behaviour of servers on first and second serve:
When serving, players can opt for risky serves which are more likely to fail but are harder to return if successful or more conservative serves which are less likely to fail but are also easier to return.
The key is whether players behave differently on first and second serves (more on that in a moment). However, simply comparing first and second serves is not so straightforward. The authors correctly note that there is:
...a potential caveat with raw data on tennis serve: it can be characterised by a selection problem. First serves are always observed while second serves are only observed when the first serve failed. This means that second serves may be more likely to be observed when serving is harder than usual either for natural reasons (e.g wind conditions), fitness (e.g. tiredness late in the match) or strategic reasons (e.g. opponent having learned how to return the player’s serve).
Their solution is quite ingenious:
To cleanly compare first and second serves one ideally wants to observe some random events which determines in a given situation whether a serve is going to be a first or a second serve. We argue that such a situation occurs when the ball hits the tape (top of the net) on the first serve. The impact with the net gives the ball an unpredictable trajectory leading the ball to be either in or out. It introduces the required randomness as a first serve follows a ball let which lands in the court and a second serve follows a ball which lands outside the court.
The serve immediately following a 'let serve' is randomly either a first serve (if the 'let serve' landed in) or a second serve (if the 'let serve' landed out). Ely et al. use a dataset from 3,188 matches, involving over 690,000 serves, of which 7,605 follow a 'let serve' and are the core sample of interest. They test four conditions which would imply that players are correctly maximising their chance of winning:

  1. That first serves are more risky than second serves (the probability that a serve lands in is lower for first serves);
  2. That first serves are harder to return than second serves (players are more likely to win the point on their first serve);
  3. Using two first serves is a suboptimal strategy (it leads to a lower probability of winning the point); and
  4. Using two second serves is also a suboptimal strategy.
They find that:
...the serves from professional tennis players meet four conditions which make them consistent with the optimal strategy of risk taking between first and second serves. This result is observed both overall and when splitting the sample by gender and ranking.
So, it appears that professional tennis players are optimisers. Which we should expect - they are trained professionals who have developed skills in strategic play over many years.

Read more:



Tuesday, 12 December 2017

How not to measure sexual risk aversion

Risk aversion seems like such a simple concept - it is how much people want to avoid risk. Conventionally, economists measure the degree of risk aversion of a person by how much of an expected payoff they are willing to give up for a payoff that is more certain (or entirely certain). If you're willing to give up a lot, you are very risk averse, and if you are not willing to give up a lot, you are not very risk averse. But notice that the measure of risk aversion is all about behaviour, either as a stated preference (what you say you would do when faced with a choice between a more certain outcome, and a less certain outcome that has a higher payoff on average) or a revealed preference (what you actually do when faced with that same choice).

So, I was interested to read this recent paper by Stephen Whyte, Esther Lau, Lisa Nissen, and Benno Torgler (all from Queensland University of Technology), published in the journal Applied Economics Letters (sorry, I don't see an ungated version). In the paper, the authors claim to be comparing "risk attitudes towards unplanned pregnancy and sexually transmitted diseases (STDs)" between health students and other students. It is an interesting research question, since you might expect health students to be better informed about the actual risks of sexual behaviour.

However, when you look at the measure they used for risk attitudes, it becomes immediately clear that there is a problem:
To assess participant perceptions of the safety of different forms of contraception and sexual contact in relation to unplanned pregnancy and STDs, they were asked to rate, on a seven-point scale from 0% safe to 100% safe, the level of safety of each of six options. The six responses were then summed and divided by the number of responses to create a measure of average individual attitudes towards the specific risk.
The six options for risk of unplanned pregnancy were condoms; contraceptive pill; sex during menstruation; intrauterine devices; withdrawal method; and contraceptive implant; and the six options for risk of STDs were oral sex; physical contact; kissing; digital penetration; anal penetration; and vaginal penetration. At least, I think that's the case, as it was a little unclear from the paper.

However, their measure is clearly not a measure of risk attitudes (or risk aversion) at all. It is a measure of 'perceptions of safety'. Notice that the measure doesn't ask about students' sexual behaviour at all, and doesn't ask about a trade-off decision. So, it won't tell you much at all about risk aversion. In order to turn it into a measure of (sexual) risk aversion, you would at the very least need to ask the students to choose between two (or more) of the options, with different levels of risk and different levels of either 'beneficial payoff' or (more likely) cost.

Perceptions of safety of the different options is one component of the decision of which option to engage in (or to engage in none of them), but alone it does not tell you about risk aversion. A student might report that they believe the options convey a low degree of safety, but that doesn't mean that the student is risk averse. It just means that they believe that the options presented to them are high risk (low safety). Similarly, a student who reports that the options convey a high degree of safety is not necessarily less risk averse than a student who reports that the options convey a low degree of safety.

How would we expect health students to be different from other students? You might expect health students to be better informed about the actual safety associated with the different options (at least, you'd hope that they would learn this in their health studies!). In other words, you might expect other (non-health) students to over- or under-estimate the degree of safety of the different options to a greater extent than health students. Let's say that non-health students are more likely to over-estimate safety. They are more likely to take risks with their sexual health and in terms of unplanned pregnancy than are health students, because the health students are better informed about the real levels of safety of each option. This would manifest in higher measures of 'perception of safety' among non-health students than among health students. And these authors would interpret this as greater risk aversion among health students, when in fact it is entirely driven by the non-health students being misinformed relative to the health students.

Notice also that the measure of 'perceptions of safety' increases if students believe that oral sex is safer (in terms of avoiding risk of STDs), or if kissing is safer, or if vaginal sex is safer, with no consideration of the actual level of risk associated with each option. It would have been better to evaluate some of the options separately, rather than all together, since evaluating them all together really turns their measure into a general measure of 'perceptions of safety of sexual activity'.

That latter problem aside, the results of the paper are still interesting, provided you interpret them (correctly) in terms of 'perceptions of safety of sexual activity'. Students who reported as virgins had lower perceptions of safety (which might explain in part why they are still virgins). Older students had lower perceptions of safety (I guess, you learn from your mistakes, or your friends' mistakes?). Male students had higher perceptions of safety in terms of STDs, but not in terms of unplanned pregnancy (this one was a bit or a surprise, as I would have expected the opposite). Non-religious students (which the authors label as atheists) had lower perceptions of safety in terms of STDs, but higher perceptions of safety in terms of unplanned pregnancy (I guess the religious students are more worried about pregnancy, which can't easily be hidden from their peers and family, than they are about STDs, which can?).

Anyway, even though the results are interesting, it doesn't change the fact that this is not the way to measure risk aversion.

Sunday, 12 March 2017

The irrationality of NFL play callers, part 2

A few weeks ago I wrote a post about the irrationality of NFL offensive play callers, specifically that they fail to adequately randomise their play choices, with the implication that defensive play callers should be able to (and do) exploit this for their own gain. Why would they do this? The Emara et al. paper (one of the two I used in the post) suggested:
Perhaps teams feel pressure not to repeat the play type on offense, in order to avoid criticism for being too “predictable” by fans, media, or executives who have difficulty detecting whether outcomes of a sequence are statistically independent. Further, perhaps this concern is sufficiently important so that teams accept the negative consequences that arise from the risk that the defense can detect a pattern in their mixing.
Which seems like a plausible suggestion. Last week I read this recent post by Jesse Galef on the same topic:
In football, it pays to be unpredictable (although the “wrong way touchdown” might be taking it a bit far.) If the other team picks up on an unintended pattern in your play calling, they can take advantage of it and adjust their strategy to counter yours. Coaches and their staff of coordinators are paid millions of dollars to call plays that maximize their team’s talent and exploit their opponent’s weaknesses.
That’s why it surprised Brian Burke, formerly of AdvancedNFLAnalytics.com (and now hired by ESPN) to see a peculiar trend: football teams seem to rush a remarkably high percent on 2nd and 10 compared to 2nd and 9 or 11.
What’s causing that?
Galef argues that there are two possibilities (note that the first one is similar to the suggestion by Emara et al.):
1. Coaches (like all humans) are bad at generating random sequences, and have a tendency to alternate too much when they’re trying to be genuinely random. Since 2nd and 10 is most likely the result of a 1st down pass, alternating would produce a high percent of 2nd down rushes.
2. Coaches are suffering from the ‘small sample fallacy’ and ‘recency bias’, overreacting to the result of the previous play. Since 2nd and 10 not only likely follows a pass, but a failed pass, coaches have an impulse to try the alternative without realizing they’re being predictable.
Galef then goes through some fairly pointy-headed methodological stuff, before arriving at his conclusion:
If their teams don’t get very far on 1st down, coaches are inclined to change their play call on 2nd down. But as a team gains more yards on 1st down, coaches are less and less inclined to switch. If the team got six yards, coaches rush about 57% of the time on 2nd down regardless of whether they ran or passed last play. And it actually reverses if you go beyond that – if the team gained more than six yards on 1st down, coaches have a tendency to repeat whatever just succeeded.
It sure looks like coaches are reacting to the previous play in a predictable Win-Stay Lose-Shift pattern...
All signs point to the recency bias being the primary culprit.
However, I'd still like to see some consideration of risk aversion here. Galef controlled for game situation and a bunch of other game- and team-level variables, but not individual-level variables related to the coaches (he did control for quarterback accuracy, but as far as I can see that might be a team-level variable if the team has changed quarterback mid-season).

This is yet more evidence that there is an exploitable trend in NFL offensive play calling, but the reason underlying this trend is still not fully established. Defensive play callers need not care about the reasons why though - they should be adjusting their strategies now.

Read more:


Monday, 20 February 2017

The irrationality of NFL play-callers

I recently read two papers that both essentially conclude (based on different aspects) that NFL play callers are not rational (or more specifically, not rational and risk neutral - an important point I'll return to at the end of the post). Recall that a rational decision-maker weighs up the costs and benefits of a decision, and when faced with mutually exclusive options (such as choosing which play to run in an NFL game), they should choose the option with the greatest net benefit (benefits minus costs).

The first paper (by Jonathan Hartley, an MBA student at the Wharton School at the University of Pennsylvania) looks at play-callers' choices between an extra point attempt and a two-point attempt following a touchdown. A rational and risk-neutral play-caller should choose whichever play provides the greatest expected benefit (expected number of points). In this case, Hartley found:
Between 2002 and 2014, the extra point conversion rate was 99.2% (out of 7738 attempts). As the average two point conversion rate remained 0.475, the expected points from a two-point conversion remains 0.95 below the automatic 0.992 points...
Over 2 seasons since the implementation of the new rules [increasing the distance the extra point try is attempted from], the extra point conversion rate has fallen from 0.992 to 0.95. Moreover, the total number of 2 point conversion attempts per season has nearly doubled...
In other words, when the NFL changed the extra point to being attempted from a greater distance (thereby making it more difficult), the expected value of an extra point try fell from 0.992 to 0.95 points. The expected value of a two-point conversion remained steady at 0.95 points. So, a rational and risk-neutral play-caller should now be indifferent between an extra point attempt and a two-point conversion. However, as Hartley shows in the paper, most teams still attempt very few two-point conversions, even those teams that have a history of success at them. The paper itself is pretty rough, but I wish the MBA students here could do this sort of work!

The second paper, by Noha Emara (Rutgers), David Owens (Haverford College), John Smith (Rutgers), and Lisa Wilmer (Florida State), is forthcoming in the Journal of Behavioral and Experimental Economics (ungated earlier version here), and looks at serial correlation in play-calling. Serial correlation occurs when you have a time series (like a series of plays) and where each observation in the time series is related (positively or negatively) to the observation or observations earlier in the time series. Obviously, an NFL offensive play-caller wants to call players in a random way - what we call a mixed strategy. Mixed strategy is particularly important in sports - think of the choice of where to serve in tennis, or where to shoot a penalty or which way to dive as a goalkeeper in soccer (see here or here for more on this). If an NFL offensive play-caller doesn't effectively randomise their play calling, then the defence can potentially exploit some prior knowledge of the play about to be called.

Humans are rubbish at trying to create random series, and indeed that's what Emara et al. found, based on their dataset of more than 200,000 plays from the 2000-2012 NFL seasons:
...the previous pass variable is negative and significant in each specification. This provides evidence that, even after controlling for down, distance, field position, and other observables, play calling exhibits significant negative serial correlation. The Previous pass-Previous failure interaction estimate is negative significant in both of the specifications where it appears, suggesting that play calling becomes even more negatively serially correlated following a failed play.
To translate, play-callers are significantly more likely to call a running play after a previous passing play, and to call a passing play after a previous running play, than would be expected if they were selecting plays randomly. And on top of that, if the previous play was a failure (e.g. if it lost yards), then they are even more likely to change the play type on the following play.

To make things worse, Emara et al. find evidence that teams would be better off if they ran more plays that were the same as the previous play:
We find that a rush following a rush gains 0.14 more yards than a rush following a pass. We also find that a pass following a pass gains 0.21 more yards than a pass following a rush. Estimates are more pronounced when we also control for whether the previous play was a failure. We find that a rush gains 0.24 more yards more following a failed rush than following a failed pass. Also, a pass gains 0.34 more yards following a non-failed pass than following a non-failed rush.
In summary, we find evidence that the efficacy of a play, as measured by yards gained, increases if it follows a play of the same type.
The results is even stronger on second down plays, but not so much for third down plays. However, the take-away message, like that of the first paper, is that play-callers are not being purely rational.

However, there is a caveat here. If we think that, based on this evidence, that play-callers should be calling more two-point conversions and switching up play types less often, then we may be forgetting that there is also a wider game at play here. If play-callers are risk averse, then this affects their decision-making. The two-point conversion may have the same expected value as an extra point attempt, but it is riskier (see also this post on NBA three-pointers from last week), so a risk averse play caller may avoid the two-point conversion more than the simple comparison of expected values would suggest.

But what about the play-callers in the second paper? Emara et al. have thought about this, and this is what they offer:
Perhaps teams feel pressure not to repeat the play type on offense, in order to avoid criticism for being too “predictable” by fans, media, or executives who have difficulty detecting whether outcomes of a sequence are statistically independent. Further, perhaps this concern is sufficiently important so that teams accept the negative consequences that arise from the risk that the defense can detect a pattern in their mixing.
Making play calls that the fans think are predictable (but which are actually more random) may make the play-caller themselves at risk of losing their job (or at least, of looking like they are doing a poor job). So, play-callers may attempt to make their play calls look more random by switching (from run to pass or vice versa) more often than they should, even though this is actually less random and costs the team in terms of yards gained per play.

The question is, now that these trends are known, will any team want to exploit them?

[HT: Marginal Revolution, here and here]