Sunday, 6 January 2019

Parachutes may be ineffective as safety devices

A lot of what you read about research in mainstream media is taken from press releases (by universities or research institutes), or at best, is taken from the abstracts of research papers. Randomised controlled trials are the gold standard in evaluation research (especially in health), so many readers might not take a critical eye when reading the studies. This would be a mistake.

A good example of where a lack of critical reading would go horribly wrong is this article from the Christmas issue of the British Medical Journal, by Robert Yeh (Harvard Medical School) and co-authors, entitled "Parachute use to prevent death and major trauma when jumping from aircraft: randomized controlled trial". Here's the first sentence of the conclusions from the abstract:
Parachute use did not reduce death or major traumatic injury when jumping from aircraft in the first randomized evaluation of this intervention.
One could conclude from that sentence that parachutes are ineffective as safety devices. However:
Compared with individuals screened but not enrolled, participants included in the study were on aircraft at significantly lower altitude (mean of 0.6 m for participants v mean of 9146 m for non-participants; P<0.001) and lower velocity (mean of 0 km/h v mean of 800 km/h; P<0.001).
The authors tried to enrol people into their randomised controlled trial by asking people:
...whether they would be willing to be randomized to jump from the aircraft at its current altitude and velocity.
They would be randomised into making the jump either with, or without, a parachute. The only people willing to be randomised in the study, unsurprisingly, were those on stationary aircraft on the ground. This limits the external validity of their sample a little, but does allow them to conclude that:
...although we can confidently recommend that individuals jumping from small stationary aircraft on the ground do not require parachutes, individual judgment should be exercised when applying these findings at higher altitudes.
This study was, of course, a response to an early BMJ article on parachutes from 2003, which concluded that:
...the effectiveness of parachutes has not been subjected to rigorous evaluation by using randomised controlled trials.
Well, now the effectiveness of parachutes has been evaluated, and found wanting. Read both papers (they're open access); they're hilarious (as has been the case with many past papers in the Christmas issue of the British Medical Journal, such as this one I blogged about in 2017).

[HT: Thomas Lumley at StatsChat, whose pet hate is journalists quoting uncritically from press releases or abstracts of non-peer-reviewed papers]

Friday, 4 January 2019

Gender differences in multiple choice answering

In a post in 2017, I discussed multiple choice and constructed response questions in exams. One of the key points of that post was to highlight the gender difference in performance in multiple choice questions (female students do worse, but they do better on constructed response questions). It turns out that result is fairly common in the academic literature, but the reason why female students don't perform as well as male students on multiple choice questions is less clear. Maybe it is that female students don't respond well to high pressure or competitive situations (and multiple choice questions are reasonably high pressure). Women are more risk averse than men, so maybe it is related to that? Female students might be more likely to skip questions in order to avoid the risk of losing marks [*]. Also, men are more overconfident than women, so maybe it is related to that? Again, male students might be less likely to skip questions, because they are less likely to be unsure they have the right answer.

So, I was quite interested to read this 2017 article by Gerhard Riener and Valentin Wagner (both Düsseldorf Institute for Competition Economics), published in the journal Economics of Education Review (sorry, I don't see an ungated version, but it looks like it might be open access anyway). Riener and Wagner conduct an experiment with 2060 German secondary school students across 89 classes in 25 schools, who each sat a maths test (based on the Math Kangaroo questions).

There were three levels of questions (easy, worth 3 marks; medium, worth 4 marks; and difficult, worth 5 marks). Students lost one mark for every question they got wrong. So, there is an incentive not to guess the easy questions if you don't know (because there is a 1/5 chance of getting three marks, and a 4/5 chance of losing a mark, so the expected value of guessing is -0.2 [0.2 * 3 + 0.8 * -1]). There is no incentive either way for the medium questions (the expected value is 0 [0.2 * 4 + 0.8 * -1]), and a positive incentive to guess on the difficult questions (the expected value is 0.2 [0.2 * 5 + 0.8 * -1]). So, Riener and Wagner expect students to skip more of the easy questions, and fewer of the difficult questions. It turns out that wasn't the case, since they:
...find that the number of skipped questions is increasing in difficulty.
I'm not surprised by this at all. Students don't understand expected value, so it wouldn't surprise me that they didn't realise that there was a positive expected value for guessing, for the difficult questions. And the positive expected value is a relevant consideration for a risk neutral decision-maker. Since there were only 14 questions in the test (and only 4 difficult questions), then it is relatively risky to guess on one of them, in terms of the impact on overall score.

That wasn't the main results from the paper though, which concerned the gender differences, and the experiment. The experiment was that students in some classes were rewarded, if their test score was better than their earlier mid-term result (they didn't know about the experiment until after the mid-term, so there is no risk of the students engaging in strategic behaviour). The reward (according to Riener and Wagner, a source of extrinsic incentive to do well in the test) was one of: (1) a medal, awarded in front of the rest of the class; (2) a letter to their parents, praising their good performance; (3) a "no homework" voucher, which entitled them to take a day off homework; or (4) a "surprise gift" (which was actually a combination of (1) and (2)).

They find that:
Females always tend to skip more questions than males regardless of whether they are incentivized or not. However, incentivized pupils tend to skip fewer questions than non-incentivized pupils...
...girls in our low stakes baseline treatment skip significantly more questions than boys... However, the gender gap depends on item difficulty. While girls skip as many questions as boys when items are easy... they skip significantly more questions for medium... and difficult questions... Interestingly, providing extrinsic incentives for performance, and hence increasing the stakes, closes the gender gap in skipping test items. 
So, providing an incentive closed the gender gap. They also find that the gender gap is only present in academic high schools (Gymnasium) and not in vocational high schools (Gesamtschule, Realschule, and Hauptschule). However, I'm not convinced by those results, as the vocational students sat an easier test, with fewer questions used in the analysis, which renders the results not comparable with the academic high school students.

Riener and Wagner argue that their results are:
...suggestive evidence that the gender gap could be explained by a stereotype threat. Girls in high school only skip significantly more questions if the questions are difficult, although the attractiveness of answering is higher for difficult questions than for easy questions. Further support for a stereotype threat explanation is the fact that the gender gap vanishes if the difficulty of the task is made less salient (shifting the focus of pupils to winning an extrinsic reward).
It's possible that this is stereotype threat. However, it is also possible that the types of rewards that they were offering were the types of rewards that girls value more than boys (especially given that the students got to choose their preferred reward). Perhaps the reward increased the stakes of the test for girls, but not for boys? In any case, it's hard for me to see how the offering of the reward reduces stereotype threat. So, although they were able to eliminate the gender difference in performance, I don't think this paper really helps us get to the bottom of why female students usually perform worse than male students in multiple choice questions.

*****

[*] This applies if there are negative marks for a wrong answer. It is a bit harder to argue this when there is no penalty for a wrong answer (or at least, no difference in the penalty between skipping and getting the answer wrong).

Read more:


Wednesday, 2 January 2019

Book review: Foolproof

One good thing about the Christmas break is that I have a lot of time for catching up on reading. So, it only took a couple of days for me to read Greg Ip's Foolproof: Why Safety Can Be Dangerous and How Danger Makes Us Safe.

I quite enjoyed the book, as at its heart it is about moral hazard and especially about unintended consequences, which regular readers of this blog will recognise is one of my favourite topics. However, Ip continually returns to financial crises, which nicely knits the book together into a coherent narrative. If you've already read books about the Global Financial Crisis, then I doubt you would gain much from reading this one, but Ip writes it in a very accessible way, and the links to moral hazard and unintended consequences mean that almost any student of economics will learn a lot, along with the general reader. The overall message is that because our economic institutions (like central banks) work hard to reduce risks to the economy, we end up taking more risks, so that when the system fails, it does so catastrophically.

Although the book has financial crises as an underlying theme, I generally enjoyed some of the other examples more (although I must admit, a lot of the stuff about the Volcker years in the U.S. was new to me). For example, this bit on floods:
Gilbert White, an obscure government geographer who had been pursuing graduate studies part-time at the University of Chicago, noticed that the frenzy of levee and dam building in the 1930s had not solved flooding; instead it had created a new problem: more homes, factories, and farms had sprung up on the floodplain, so more destruction ensued when floods overtopped the levees.
And this bit on helmets in the National Hockey League (NHL):
Helmets became mandatory for new National Hockey League players in 1979. Thereafter the number of head fractures went down, while the number of spinal injuries went up. The conclusion of several specialists was that a more aggressive style of play, perhaps encouraged by the wearing of helmets and full face masks, was causing players to hit one another harder in ways that made spinal injuries more likely.
The lessons from these other examples though, are targeted towards finance, such as this on the 1987 stockmarket crash:
Nonetheless, the crash taught an important lesson about insurance against financial catastrophes. It works when only a few people buy it; when everyone does, it not only makes the catastrophe more likely, it threatens the survival of the system...
Just as flood and earthquake insurance enable more people to live in flood- or earthquake-prone regions, insurance against market disruptions enables more investors to pile into those markets and perversely make the event more likely and more severe. Portfolio insurance had enabled this with stocks in 1987, and now CDSs [Credit Default Swaps] did the same with mortgages.
Ip doesn't get everything right though, at least from my perspective. In one section, he does a poor job of explaining Prospect Theory, and consequently the book stumbles over the distinction between loss aversion and risk aversion. Similarly, most economists would disagree with Ip that risk aversion is characteristic of behavioural economics, and not traditional economics. However, the general reader will not notice these errors.

As you might expect of a book about financial crises, Ip does come up with solutions. He concludes that:
Our goal should be to eliminate big disasters, not small ones, to accept a bit more risk and instability today in return for more reward and stability in the long run.
An analogy here (surprisingly not used in the book) is allowing children to take risks and hurt themselves a little, so that they can learn about their own limits and avoid catastrophe in the future. A few small boo-boos in the financial system should help us to reduce the chance of serious life-threatening events in the future. Overall, this is a good book for those interested in financial crises, but who don't want to do a deep dive into a book heavy on theory.

Tuesday, 1 January 2019

Economics students are extraverted, disagreeable, but emotionally stable

I've written a couple of times about personality differences of students by academic discipline (see here and here). In one of the studies I discussed, students high in extraversion were more likely to choose to study law, or business and economics, and less likely to choose science, technology, engineering and mathematics (STEM). However, both posts were based on a single study.

A 2016 systematic review by Anna Vedel (Aarhus University), published in the journal Personality and Individual Differences (ungated version here), summarises the results of twelve studies, across seven countries (in Europe, Israel and the U.S.), and including 13,389 students. All studies compared students' Big Five personality traits by discipline, although only four studies separated out economics students (and in one of those, "economics" included marketing, management, accounting, and public administration). Vedel finds that:
Economics and Business scored consistently lower than other groups [for neuroticism]...
Economics, Law, Political Sc., and Medicine scored higher than Arts, Humanities, and Sciences [for extraversion], and the differences often represented medium effect sizes...
Humanities, Arts, Psychology, and Political Sc. scored higher than other academic majors [for openness], and effect sizes were often moderate or even large in comparisons with Economics, Engineering, Law, and Sciences.
Law, Business, and Economics scored consistently lower than other groups [for agreeableness], and a few medium effect sizes were found in comparisons with Medicine, Psychology, Sciences, Arts, and Humanities...
Arts and Humanities scored consistently lower than other academic majors [for conscientiousness], and medium effect sizes were found in comparisons with Sciences, Law, Economics, Engineering, Medicine, and Psychology.
If we take those results as representative (which might not be as big a stretch as it sounds, as the effect sizes were reasonably consistent across studies), then economics students are more extraverted, but less neurotic (alternatively, more emotionally stable) and less agreeable than other students. The big question now is, how do we leverage those traits to improve student outcomes, or to provide better advice to students?

[HT: Marginal Revolution, back in August]

Read more: