Monday, 7 March 2022

Research-based instructional strategies and flipped classrooms in economics

I'm still searching for the 'killer research' that convinces me that flipping the classroom will do good overall. Or even better, a convincing or growing consensus in the research literature. Unfortunately, I have so far found neither of those things. The latest bit of my search took me to two articles published in the 2018 issue of the AEA Papers and Proceedings, on using research-based instructional strategies (RBIS). In short, RBIS involves applying strategies from cognitive science that have been shown to improve learning gains for students. That sounds very promising.

The first article was this one (ungated earlier version here) by Austin Boyle and William Goffe (both Penn State University). They applied RBIS to two sections of macroeconomics principles in 2015. Their RBIS approach involved flipping the classroom, along with pre-class assignments (allowing for just-in-time teaching), clicker questions and worksheets, regular quizzes, quiz reflections, and homework (to achieve spaced repetition). Using the Test of Understanding of College Economics (TUCE), and comparing the learning gains of their class (of around 500 students) to those in the TUCE norming sample, Boyle and Goffe find that:

The mean pre-TUCE score was 11.90 questions correct (out of 30), and the mean post-TUCE score was 19.91. The mean learning gain, g = (post − pre)/(30 − pre) , was 0.44. In the TUCE norming population, the mean pre-TUCE score was 9.80 questions correct and the mean post-TUCE score was 14.19, for a gain of 0.22. Thus, students in this study had twice the gain of the norming population.

That all sounds very promising, but there are a couple of problems here. First, there is no true control group. The comparison with the norming sample is interesting, but not definitive. Second, there is no randomisation and so there could be selection into the class. Students who would do worse may disproportionately drop out of the class, and there doesn't appear to have been any adjustment made for that.

The second paper is this one (sorry I don't see an ungated version) by Sarah Cosgrove and Neal Olitsky (both University of Massachusetts, Dartmouth). They applied RBIS to an economic principles class, and importantly they have a control class (that uses traditional teaching methods). Unfortunately, other than flipping the classroom, the article doesn't give much specific detail about what the RBIS strategies that were employed actually were. Unlike Boyle and Goffe, we do find out that a little more than 10 percent of students (in both treatment and control groups) dropped the class, and there were similar proportions that dropped out of both groups (that doesn't mean that there is no selection bias, though, if for example better students dropped out of the control group, while worse students dropped out of the treatment group). Cosgrove and Olitsky find that:

...the average effect of treatment is an additional 3.2 questions or additional relative learning gains of approximately 15 percent. The potential outcome means results indicate that students in the treatment group answered 11.6 more questions correctly or 62.6 percent, on average, on the final exam than they did on the pretest; whereas the control group improved by 8.4 questions or 47.7 percent, on average.

However, this research also has problems. There was no randomisation of students between treatment and control groups, and there are large differences between the two groups. The treatment group had many more White students (64 percent, compared with 37 percent in the control group), more US citizens (98 percent, compared with 78 percent), and fewer low income students (11 percent, compared with 29 percent). Even if you control for those variables in the analysis (as Cosgrove and Olitsky did), it is hard to believe that they don't contribute significantly to the observed differences in performance.

I'd like to know more about these RBIS strategies, but unfortunately the format of AEA Papers and Proceedings precludes a long exposition of the teaching methods. However, given that we can't have strong confidence in the results (in both cases because the treatment and control groups were too different), I think I'll wait until I see some higher quality research on this.

Read more:

Sunday, 6 March 2022

Eurocentrism and the use of country names in the titles of research papers

If you read enough empirical research papers, you will eventually start to notice that the names of countries are sometimes included in paper titles. This practice is far from universal, but it provides an immediate signal of the context of the research. I'm not generally in the habit of putting country names in my research paper titles, but I have had journal editors request it. Now, if you have read many empirical research papers, you may note that there is one striking aspect to the practice of country names in paper titles: the US appears to be almost entirely immune to it. It is rare to see the US named in a paper title.

Now, a new research paper by Andrés Castro Torres (Max Planck Institute for Demographic Research) and Diego Alburez-Gutierrez (Universitat Autonoma de Barcelona) takes a closer look at country names in paper titles. Specifically, they collected data from over 1.2 million publication records on Scopus between 1995 and 2020, of which over 560,000 mentioned at least one country in the title or abstract. Castro Torres and Alburez-Gutierrez focus on this narrower set of publications, and define whether they are 'localised' or not, where a 'localised' paper includes the name of the country in its title. They then look at localisation rates (the proportion of publications that focus on a country that have that country's name in the title), and find that:

...although articles about Europe and North America dominate our sample, they have the lowest localization rates. The vast majority of the research articles we study focuses on countries in the global North - more than 60% of the total articles mention a European or North American country in their abstract... but the localization rate of these articles is the lowest, hovering around 42% for papers published between 1996 and 2020... This percentage contrasts with the localization rates in other regions, particularly in Eastern and South-Eastern Asia and Sub-Saharan Africa. The ratio of the localization rates between these regions and Europe and North America ranges from 1.5 to 1.8 throughout the period of analysis.

So, localisation is far more likely for research outside of the global North (which includes Europe and North America, as well as Australia and New Zealand). Moreover, there is no trend in localisation - it is stable over the entire time period that Castro Torres and Alburez-Gutierrez examined:

We find no clear temporal trend. Even though localization rates decreased over time, they did so very slowly (e.g., the coefficient for the last period indicates that, compared to articles published between 2005 and 2010, those published in the last five years are exp(-0.06) = 0.94 times as likely to be localized).

Looking at individual countries, Castro Torres and Alburez-Gutierrez confirm my suspicions, finding that:

...compared to the US, the other top 10 countries of study display higher localization rates. The lowest coefficient among these countries pertains to the UK (0.54), implying that papers about the UK are 1.72 times more likely to be localized than papers about the US. This is a very significant gap, and it suggests that, even at the top of numerical dominance, the lack of localization is not explained by the use of the English language alone. By contrast, the largest coefficient pertains to articles about China. This coefficient implies a relative risk of 3.67, meaning that, while two-thirds of US-focused papers are not localized, more than two-thirds of China-focused articles are.

In other words, papers about China are nearly four times as likely to mention China in the title, than papers about the U.S. are to mention the U.S. in the title. And the comparison is qualitatively the same for every other country. I was surprised though, that as many as one-third of papers on the U.S. mention the U.S. in the title.

Finally, Castro Torres and Alburez-Gutierrez show that the geographical results hold across 27 subfields of the social sciences, although interestingly:

The localization rate ranges from 0.33 among papers classified as “Applied Psychology” (n = 3,422) to 0.66 among papers classified as “Development” (n = 69,724). Articles on “Political Science and International Relations” (n = 52,961) and “Demography” (n = 21,328) display localization rates very similar to “Development,” namely 0.64, whereas articles on “Psychology (miscellaneous)” (n = 247), “Experimental and Cognitive Psychology” (n = 1,416), and “General Psychology” (n = 3,771) display localization rates below 0.4...

Does any of this matter though? Castro Torres and Alburez-Gutierrez raise the problem of Eurocentrism:

...Eurocentrism, understood as a worldview that considers Western thought as culturally and intellectually superior, continues to shape the global production of social sciences, including its questions, methods, and approaches...

One problematic aspect of this perspective is that it glosses over the historical contingencies and structural violence that produced and sustain Western hegemony, including the imposition of metrics that makes the West the “default case” and the search for universal, timeless, and context- and value-free knowledge in science. This might result in societal processes observed in countries of the global North, such as market-based economic growth, rising human development, liberal democracy, market/trade integration, and globalization, being considered the “default” cases towards which other nations and societies ought to converge in the mid- or long term...

As I mentioned earlier in the post, having the country name in the title of the paper provides context for the results. It is fair to question the external validity of research beyond its particular cultural context. However, those questions are sidelined to some extent when the context is hidden from view. Now, this wouldn't be a problem if researchers and policy-makers carefully read research and noted the cultural context in which it was conducted and whether it is suitable to assume that those results apply more widely. However, many do not, and there is a fair amount of research that is cited without having been read (e.g. see here). That leads to a situation where, as Castro Torres and Alburez-Gutierrez conclude, failing to note the country context in the paper title:

...can be misleading, if not outright harmful, in particular when evidence-based policy is involved.

That might be overstating the case a little, but it is certainly something about which we should be more aware.

[HT: David McKenzie, on the Development Impact Blog]

Saturday, 5 March 2022

More on the value of words in wine descriptions

I posted last month about consumers' willingness to pay for wine bullshit. That was based on research by Kevin Capehart, which was published in the Journal of Wine Economics. Now, Capehart has another new article published in the same journal on a related topic (sorry, I don't see an ungated version online). In this new article, Capehart follows up on earlier research by Coco Krumme (described here). Essentially, Capehart uses a dataset of 120,000 wine descriptions, and classifies them into 'high price' (over US$50) or 'low price' (under US$15). He then trains a Naive Bayesian Classifier to predict which wines belong to the low-price category and which belong to the high-price category, based on the words in their descriptions.

Capehart is mostly able to reproduce very similar results to the earlier Krumme work, but perhaps more interestingly:

...I find for the dataset studied here that there do seem to be two mostly distinct vocabularies for high- and low-priced wines. Out of the roughly 20,000 unique words used to describe the over-$50 and/or under-$15 wines, only 42% of those words overlap by being used to describe wines in both price categories. The remaining 58% of words are non-overlapping.

In other words, the descriptions of high-priced wines use very different words than the descriptions of low-priced wines. You might argue that is because high-priced wines have different characteristics than low-priced wines, their descriptions should include different vocabularies. However, the key question that isn't answered (and which Capehart alludes to in his conclusion) is, does the vocabulary relate to the quality of the wine, is it simply a signal of the price? In other words, do those who are writing the descriptions choose their vocabulary based on the price of the wine, or based on the quality of the wine? We'd need more research in order to answer that question, and more like Capehart's earlier work.

Read more:

Thursday, 3 March 2022

Unexpected alcohol prohibition and mortality in South Africa

The District Health Boards in New Zealand are busy cancelling elective surgeries and non-urgent care in order to deal with the ongoing surge in coronavirus cases. In 2020 and 2021, South Africa took things much further, on a number of occasions banning the sale of alcohol in order to reduce alcohol-related hospital presentations and injuries (see here). This new working paper by Kai Barron (Wissenschaftszentrum Berlin für Sozialforschung) and co-authors looks at the impact of the second ban (from 13 July to 18 August 2020) on mortality and other related outcomes. This particular ban makes a good natural experiment because:

...it was unexpected. The alcohol ban was announced in the evening of Sunday, July 12, 2020, and came into immediate effect from Monday morning on July 13, 2020. Second, it was implemented in the middle of the so-called “Level 3” COVID-19 policy response period during which time other policies and regulations were largely held constant.

In other words, no one could have anticipated the ban and adjusted their behaviour beforehand, so it provides a clean 'break' in the time series data, and no other policy could make it difficult to isolate the effect of the alcohol ban. [*]

Barron et al. focus on the effect on mortality due to unnatural causes (which includes road traffic injuries, interpersonal violence, and suicides) in the first instance. They use data on daily mortality from 2017 to 2020, and employ a difference-in-differences approach. I haven't seen it employed in quite this way before. They essentially treat the alcohol ban dates in each year from 2017 to 2019 as a control, and the same dates in 2020 as the treatment period. So, their difference-in-differences is the difference between the alcohol ban dates and the rest of the year in 2020 compared with the difference between the alcohol ban dates and the rest of the year in 2017 to 2019. In one of the footnotes, they argue that this approach:

...can be justified when there is strong year-on-year temporal regularity in the outcome of interest.

There does seem to be a persistent trend across the years 2017 to 2019, and Barron et al. control for day-of-the-week effects, week-of-the-month effects, and month fixed effects, which should capture most of the regularities over time. Their key results are best summarised in Figure 2 from the paper:

Notice that there's a drop in mortality from unnatural causes in 2020/21 every time there was an alcohol ban in 2020. Barron et al. are focused on the middle period, where the ban was unanticipated, and not correlated with other policy changes. The effect looks smaller there, but in their preferred econometric model they find that:

...the alcohol ban reduced unnatural mortality by 21.99 deaths per day... 

That doesn't sound like a lot, but it represents about a 14 percent reduction in deaths from unnatural causes. Barron et al. then go on to show that:

For men, the pattern is similar to that observed in the population as a whole, with the estimates indicating that the alcohol ban reduced mortality by approximately 21 deaths per day... For women, we find no significant impact of the alcohol ban on mortality.

That shouldn't be too surprising, given that men are at greater risk of deaths from unnatural causes. And then when they look at young adults (aged 15 to 34 years) they find that:

 ...the alcohol ban reduced mortality amongst men in this age-group by approximately 12 deaths per day... and may have had a small impact on the mortality of younger women.

So, overall, the alcohol ban reduced mortality from unnatural causes, almost entirely among men, and over half of the reduction was among young men. Barron et al. then go on to report results week-by-week using an event study, and it looks like most of the effect is concentrated in the first few weeks of the ban. That suggests that there might be some adaptation, or that there was less compliance with the ban as it extended over time. Those results should caution us from assuming that a complete prohibition on alcohol would lead to a large improvement in mortality.

Finally, using police data Barron et al. find that there were:

..at least 77 fewer homicides, 790 fewer assaults and 105 fewer rape cases reported per week during the alcohol ban period in comparison to the preceding five weeks. These constitute a drop in each outcome of 21%, 33% and 19% respectively.

Altogether, this supports a conclusion that the alcohol ban reduced interpersonal violence, leading to less alcohol-related harm, and lower mortality. We know that alcohol is associated with harm. This study has given us a particularly clean estimate of just how much harm it can cause.

[HT: Marginal Revolution]

*****

[*] Actually, that's not quite true, as there was also a curfew put in place, but the curfew lasted much longer than the alcohol ban, and Barron et al. show that removing the ban but leaving the curfew in place was associated with a bounce back in mortality and other outcomes.