Thursday, 9 November 2017

Female student performance in high-stakes biology exams

Phys.org reported on a new study a couple of weeks ago:
A new study of students in introductory biology courses finds that women overall performed worse than men on high-stakes exams but better on other types of assessments, such as lab work and written assignments. The study also shows that the anxiety of taking an exam has a more significant impact on women's grades than it does for men.
 "It was striking," said Shima Salehi, a doctoral student at Stanford Graduate School of Education and one of the study's two lead authors. "We found that these types of exams disadvantage women because of the stronger effect that test anxiety has on women's performance."
The original study is available here (open access), published in the online journal PLoS ONE. The authors were Cissy Ballen (University of Minnesota), Shima Salehi (Stanford), and Sehoya Cotner (University of Minnesota). The results are based partly on data from 1205 first-year biology students over ten sections, with the results on test anxiety (and 'interest in course content') based on survey data from 372 students over three sections. In the paper, the key research questions were:
1) What is the extent of the gender gap in incoming academic preparation among students? 2) What is the extent of the gender gap in exam grades and non-exam grades? 3) Do women and men report different levels of test anxiety and interest in science? 4) Do these two affective factors influence performance outcomes in undergraduate biology courses?
The authors found that there was a significant gender gap in academic preparation among students, with ACT (American College Test) scores on average about 0.28 standard deviations lower for female than for male students. There was also a difference in exam grades between female and male students, of 0.15 standard deviations. However, to me the key result is:
When we included incoming ACT score in the model as a fixed effect, the gender gap in exam performance disappeared...
In other words, the performance gap in exams between female and male students was almost entirely explained by differences in student quality (as measured by the ACT score). There was no need for the authors to dig into text anxiety or interest in course content, especially given that the results they present based on their mediation analysis actually don't show anything because the combined paths are not statistically significant. Female students did worse because they were worse students, not because of some gender bias or because of test anxiety.

Or maybe not. I noticed that at one point in the paper, the authors note that the exam grades were "multiple-choice exam grades", which implies (to me) that the exams were wholly multiple choice. And we know based on past research that female students have a disadvantage in multiple choice questions. In the Phys.org article, one of the authors is quoted as saying:
We want to figure out what kind of instructional methods will ensure that everyone can navigate successfully through these courses and have a wider range of career options.
Worry less about the instructional methods. Ditch the multiple choice in your exams, or replace them with a mixture of multiple choice and constructed response. Your female students will appreciate it.

Sunday, 5 November 2017

Book Review: The Knockoff Economy

Intellectual property rights face a significant trade off, identified by the economist William Nordhaus. The trade-off is between having weaker (or shorter) intellectual property rights, which would lead to under-investment in intellectual property development, or having stronger (or longer) intellectual property rights, which would lead to under-consumption of intellectual property. For instance, if the government strongly protects intellectual property (through longer periods of copyright or patent protection), then that increases the incentives for inventors or artists to invest the time and effort necessary to create new inventions, write new books, create new artworks and so on. This is because the strong intellectual property rights create limited natural monopolies for the rights holders, allowing them to raise the price and increase their profits. However, those higher prices reduce the consumption of the intellectual property relative to the case where intellectual property rights were weaker (or less long-lasting).

However, is it always the case that stronger intellectual property rights foster innovation, and weaker intellectual property rights deter inventors or creators from inventing or creating? This is the question that is addressed in a 2012 book by Kal Raustiala and Christopher Sprigman, entitled The Knockoff Economy: How Imitation Sparks Innovation. In the book, the authors look mainly at three industries (fashion, cuisine, and stand-up comedy) and show that substantial innovation occurs in each case in spite of a lack of strong enforcement of intellectual property rights. Indeed, in the examples discussed in the book, copyright and patents are either not used, or are not available. And yet, in all cases there is a great deal of ongoing innovation. This narrative provides a strong counter-argument to the seemingly-constant increases in the strength and length of intellectual property rights protection being granted in many western countries, especially the U.S.

As I was reading through the book, I made a large number of notes of things I wanted to discuss in my review, but there is really no way I could address them all and keep this post manageable. Because in each of the three cases that make up the first three chapters of the book (fashion, cuisine, and stand-up comedy), the reasons why innovation remains high are quite different. On fashion, the authors note that:
...the apparel industry is not just surviving - it is thriving. Extensive and legal copying accelerates the fashion cycle, banishing once-desired designs to the dustbin of apparel history (perhaps later to the dusted off and reintroduced) and sending the fashion-conscious off in search of the new, new thing.
In the fashion industry, the act of copying drives innovation because there are customers who really want something new, but they are only driven to something new once many others have started wearing the old, new fashions. Cuisine, though, is different. It is robust to copying because you aren't really buying just the meal but an experience, which is difficult to copy:
The dish you crave must be purchased as part of a larger, multifaceted transaction, replete with various courses, beverages, and side dishes. There are ambience, service, energy, and other intangibles in the mix. All of these factors work together. Copying one aspect - the main dish - may be easy. Copying the experience in full is virtually impossible. The experience is less one of buying a product and more that of enjoying a performance.
Chefs may copy each others' dishes, but they also care about their reputation, which is even more the case for stand-up comics. It is social norms that keeps copying of stand-up jokes in check. The authors write that:
...comedians' norm system includes informal but powerful punishments. These start with simple bad-mouthing and ostracism. If that doesn't work, punishments may escalate to a refusal to work with the offending comedian. Occasionally, comedians threaten joke thieves and even beat them up. None of these sanctions depend on legal rules - indeed, when comedians resort to threatening or beating up other comics, that's obviously against the law. Yet these tactics work.
That last point reminded me of Elinor Ostrom's work on informal agreements to deal with common resource problems, and it would have been good if the authors had also noted this parallel. The idea of the fashion cycle made me think about viral smartphone apps - could a case be made for removing copyright protection from smartphone apps, in order to drive more innovation (if indeed, we want more innovation in that space)?

The book looks also in less detail at a number of other areas of innovation including football (of the American variety), fonts, finance, and databases. In these cases, the authors draw the important distinction between 'pioneers' (those who first invent something) and 'tweakers' (those who take the original invention and improve it). This process of pioneer innovation followed by tweaking has been a driver of improved quality in many cases, and the authors essentially argue that this is likely to be true in many more domains. It is an attractive argument, particularly when you consider an area like pharmaceuticals, where monopoly pricing is problematic.

Overall, I really enjoyed this book. If you're interested in intellectual property or innovation, then this is definitely a good read.

Friday, 3 November 2017

Why Pharmac might be better not to fund next-generation drugs

As reported by the New Zealand Herald earlier this week, the government is to investigate a new fund to give New Zealanders access to costly new-generation medicines:
The Cancer Society has called for an early-access scheme, and Labour's previous health spokeswoman Annette King repeatedly called for one, saying that when in Government Labour would look at what funding was needed.
New Health Minister David Clark told the Herald the Government wanted to explore how such a scheme could operate.
The United States and Britain have versions of early-access schemes to let certain patients access ground-breaking drugs.
There is a real problem with funding of these schemes for very expensive treatments. While these treatments may be effective and have highly positive outcomes for the patients that receive them, focusing on the patients who will receive the treatment ignores the opportunity costs (this is a point I have made before about Pharmac funding, here and here). The appropriate way to decide on which treatments are funded is by considering their cost-effectiveness, not by considering which treatments generate the most negative media attention for the government.

A focus on cost-effectiveness ensures that scarce healthcare resources are being used where they will generate the greatest benefit for society. A treatment is cost-effective if it increases a person's health at a lower cost than alternative treatments. Since not all treatments provide the same health benefits (and many have negative side effects, etc.), we need some way of consistently measuring the health gains from a treatment, and measuring the cost per unit of health gain. To do this, we could use Quality-Adjusted Life Years (QALYs - a measure that combines length of life and quality of life) as our measure of health gain, [*] and cost-per-QALY-gained as a measure of which treatments are most cost-effective. A treatment that provides the same increase in QALYs for lower cost, or more QALYs for the same cost, should be preferred for funding.

That might sound unfair (especially to patients who miss out on funding, or their family or friends), but the alternative is even more unfair. If we ignore cost-effectiveness and simply fund any treatment that generates negative media attention (within the same fixed budget), then the healthcare budget will generate a lower total improvement in health. Funding expensive and less-cost-effective treatments has serious costs in terms of decreases in overall health and wellbeing of the population.

Even if the government increases funding for Pharmac, that increased funding should not necessarily go to these next-generation treatments, as there may be other currently-unfunded treatments that are most cost-effective and those should be funded first. Indeed, funds for next-generation treatments are not necessarily a good thing, as the Herald article notes:
The Cancer Drugs Fund in the UK has been overspending despite budget increases, resulting in a number of treatments being taken off its list.
An analysis in the leading cancer journal Annals of Oncology found the medicine funded through the British scheme was not worth the money, as only 18 of the 47 treatments prolonged the patient's life.
One of the paper's authors, Professor Richard Sullivan of King's College London, said the fund had been a "massive health error", and the populism that drives public policy has no place in health.
We need to be careful that our healthcare decision-making is made on the basis of what will generate the greatest gains in health for the budgeted amount, rather than making populist decisions that will make us worse off.

Read more:

[*] An alternative is to measure health using the number of Disability-Adjusted Life Years (DALYs) averted. DALYs are a measure of health lost due to illness or injury, which can be used in place of QALYs (you can read more about QALYs and DALYs here).

Wednesday, 1 November 2017

Students vs. representative samples in experimental economics

Experimental economics is an excellent tool for testing economic theories, and the effects of economic institutions, in an environment where other factors can be carefully controlled by the experimenter. In lab experiments, where the environment and the choices being made by participants are mostly artificial, most of the conditions related to the decision can be controlled (in contrast with field experiments), which eliminates most sources of bias that can affect people's decision-making. However, most of the samples used in lab experiments are convenience samples made up of university students, or even university economics students. One might rightly question whether the results obtained from lab experiments are sensitive to this selective sample, and whether the results are really generalisable to the population.

That is why I found this 2015 paper by Alexander Cappelen (Norwegian School of Economics), Knut Nygaard (Oslo and Akershus University College of Applied Sciences, Erik Sorensen, and Bertil Tungodden (both Norwegian School of Economics), and published in the Scandinavian Journal of Economics (ungated earlier version here), of great interest. In the paper, the authors compare the results of two commonly used experiments (the dictator game; and a generalised trust game), for three samples:
  1. 120 economics students;
  2. 119 humanities, science, or social science (excluding economics) students; and
  3. 136 participants recruited from a sampling frame representative of the whole population of Norway.
Because of the nature of the two experiments the authors run, it allows them to tease out some of the underlying moral motives for peoples decisions. That is, they could infer whether people made decisions based on efficiency (maximising the gains from the experiment for everyone), equity (everyone receiving the same gains from the experiment), or reciprocity (if the other person helps you in the experiment, you help them). They found that:
...students differ fundamentally from a representative sample, both in the relative importance assigned to different moral motives and in the level of pro-sociality. Moreover, we show that one needs to be careful when generalizing about gender effects on the basis of selected student samples. Both for the dictator game and the trust game, the role of gender in the student samples does not carry over to the representative sample. Finally, we show that economics students behave less pro-socially compared to non-economics students, but the two student groups are similar in the relative importance they assign to different moral motives.
We find that both equality and efficiency are important motivational forces among male representatives, whereas female representatives seem to move from a concern for equality in non-strategic environments to a focus on reciprocity in strategic environments.
So perhaps there is good reason to be a little skeptical of experimental results based on student-only samples. And in case you think that it isn't that big and issue, the authors point out in the introduction:
Among the papers published on social preferences in the top five economics journals from 2000 to 2010, only four out of 24 papers report from experiments on non-student samples, and only two of the papers report from experiments performed outside the lab...
Eight of the 24 papers published in the top five economics journals from 2000 to 2010 report from experiments on economics students, whereas nine papers rely on other student populations or do not report detailed background information on the students.
So it's clearly possible that this could be a big problem for generalising from these results. Which actually relates to a point I have made before about experimental results from student samples.