Monday, 9 May 2022

A meta-analysis of meta-analyses of the value of statistical life

Individual research studies provide a single estimate (or a small number of related estimates) of whatever the researchers are trying to measure. There are many reasons that a single study might provide a biased estimate of what is being measured, including the data that are used and the methods that the researchers choose. For someone else reading the literature on a particular topic, it can be difficult to identify what the 'best' measure is, when there are many different estimates, all based on different data and methods. In those situations, meta-analysis can help, by combining the published estimates in a particular literature into a single overall estimate. Meta-analysis can even take account of publication bias, where statistically significant estimates are more likely to be published than statistically insignificant results.

But what should you do when there are multiple meta-analyses to choose from, each using different collections of estimates from other studies, and employing different methods? Is it time for a meta-analysis of meta-analyses? A meta-meta-analysis?

That is essentially what this 2021 NBER Working Paper by Spencer Banzhaf (Georgia State University) does for the literature on the value of statistical life in the US. As Banzhaf explains:

The Value of Statistical Life (VSL) is arguably the single most important number used in benefit-cost analyses of environmental, health, and transportation policies...

When choosing a VSL or range of VSLs, analysts must sift through a vast literature of hundreds of empirical studies and numerous commentaries and reviews to find estimates that are (i) up to date, (ii) based on samples representative of the relevant policy contexts, and (iii) scientifically valid... The US EPA is the only one of the three agencies that uses a formal meta-analysis. It uses a value of $9.4m with a 90 percent confidence interval of $1.3m to $22.9m (US EPA 1997, 2020). However, even today, these estimates are based on very old studies published between 1974 and 1991...

Perhaps one reason for this surprising gap is that we now have an embarrassment of riches when it comes to summarizing VSL studies. With so many to choose from, the process of selecting which meta-analysis to use, and defending that choice, might feel to some analysts almost like picking a single "best study."... Comparing these meta-analyses, many analysts may conclude that, as with the individual studies underlying them, each of them has a bit of something to offer, that no single one is best. Thus, the old problem of selecting a single best study has just been pushed back to the problem of selecting a single best meta-analysis.

Banzhaf collates the results from five recent meta-analyses of the VSL in the US, and applies a 'mixture distribution' approach:

Essentially, I place subjective mixture weights on eight models from five recent meta-analyses and reviews of VSL estimates applicable to the United States. I then derive a mixture distribution by, first, randomly drawing one of the eight meta-analyses (the mixture component) based on the mixture weights and, second, randomly drawing one value from the distribution describing that component's VSL (e.g., a normal distribution with given mean and standard deviation), and, finally, repeating these draws until the simulated mixture distribution approximates its asymptotic distribution.

Banzhaf finds that:

...the overall distribution has a mean VSL of $7.0m in 2019 dollars... The 90% confidence interval ranges from $2.4m to $11.2m.

For comparison, the VSL in New Zealand used by Waka Kotahi is $4.42 million (as of June 2020, which equates to about US$2.8 million). However, what was interesting about this paper wasn't so much the estimate, but the method for deriving the estimate (by combining the results of several meta-analyses). That's not something I had seen applied before, but with the growth of meta-analysis across many fields in science and social science, methods for combining meta-analysis estimates need active consideration.

Also, I thought New Zealand was an outlier in basing the VSL on seriously outdated estimates. New Zealand's VSL was estimated at $2 million in 1991, and has been updated since then by indexing to average hourly earnings. There are two sources of bias that are quite problematic when we continue to use essentially the same measure for over thirty years. First, VSL measures are based on either revealed preferences (how people react to the trade-off between money and risk of death), or stated preferences (how people say they would react to the same trade-off). However, preferences change over time, and our VSL estimate is still based on the trade-off as established in 1991. If New Zealanders have become more risk averse (in relation to risk of death) over the last thirty years, then the VSL will be an underestimate. Second, it assumes that hourly earnings is the best way of indexing the measure over time. This is adequate for a measure that was based on observed hourly wages at the time it was first estimated (such as if the VSL was based on trade-offs in the labour market). However, it assumes that the risk profile of jobs across the labour market is constant over time, and that certainly isn't the case. It is likely that risks are lower now, so the indexed VSL may be too high.

Given that these two biases work in opposite directions, it seems to me that it is well past due for an update of the underlying estimate. Unfortunately, we can't easily rely on a meta-analysis (or a meta-meta-analysis) as we lack a large number of underlying estimates. There is a clear opportunity for more New Zealand-based research on this.

[HT: Marginal Revolution, last year]

Sunday, 8 May 2022

Why are aid projects less effective in the Pacific?

In a new article published in the journal Development Policy Review (open access), Terence Wood, Sabit Otor, (both Development Policy Centre, Australia), and Matthew Dornan (World Bank) attempt to answer the question of why aid projects are less effective in the Pacific. In case you wonder whether the premise for their question is correct, here's their Figure 1, which shows how much less effective aid projects are in the Pacific, compared with the rest of the world:

Putting aside the fact that the y-axis for column graphs should start at zero (and so to the naked eye this figure very much overstates the difference between Pacific countries and other countries), the probability of an aid project under-delivering is significantly higher in the Pacific than in other developing countries. To answer the question of why, Wood et al. collate data on aid project effectiveness from a range of donors, including:

...the Australian Government Aid Program; the World Bank; the ADB; the UK’s Department for International Development (DFID) (now part of the Foreign, Commonwealth & Development Office); Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ), the German government’s development agency; KfW, the German government’s development bank; the International Fund for Agricultural Development (IFAD), a specialized agency of the United Nations; Japan International Cooperation Agency (JICA), the Japanese government aid program; and The Global Fund to Fight AIDS, Tuberculosis and Malaria (GFATM).

In each case, effectiveness is standardised to a score from one (worst) to six (best). Their dataset includes 4128 projects from 1996 onwards. They apply causal mediation analysis, which effectively means that in their primary analysis they add various plausible factors that might explain the Pacific's poor aid performance, one at a time, to a regression model that includes a dummy variable for the Pacific. They then look at the effect of adding each of these mediating variables on the coefficient of the Pacific dummy variable, and its statistical significance. Wood et al. find that:

First, when governance is added, the Pacific coefficient actually becomes larger (that is, its difference from zero becomes greater). This suggests governance is a moderating variable: because good governance boosts aid project effectiveness, and because governance is better in the Pacific, the finding indicates the negative effect of the Pacific on project effectiveness would actually be greater were it not for the positive influence of comparatively good governance. Adding the freedom variable reduces the magnitude of the Pacific effect considerably. Growth and GDP also reduce the magnitude but their impact is small. Remoteness, on the other hand, has a substantial impact, and for the first time the coefficient of the Pacific’s effect on project effectiveness ceases to be statistically significant. When population is included, the coefficient for the Pacific changes substantially again, actually becoming positive albeit not statistically significantly different from zero.

The fact the Pacific coefficient is effectively zero at the end of the analysis suggests the negative effect of the Pacific on project effectiveness is completely mediated by these variables.

In other words, aid projects in the Pacific are less effectiveness because of the remoteness of developing countries in the Pacific and their small population sizes. Wood et al. find similar results using two alternative methods of causal mediation analysis. Then, looking at how the characteristics of aid projects vary in effectiveness between the Pacific and other developing countries, they find that:

Although project duration and size have some impact on project effectiveness more generally, neither appears to have a differing impact on project effectiveness in the Pacific compared to the rest of the developing world. Indeed, the only variable for which any of the interaction terms is significant, is sector, and in particular humanitarian emergency work...

...no other sector’s performance differs between the Pacific and elsewhere in a manner that is statistically significant or in any way substantively meaningful. However, humanitarian projects do perform worse in a manner that is statistically significant.

One thing did concern me a little about the analysis. The Pacific countries are the most remote and smallest in population size, so it is possible that it isn't remoteness or small size that create the problems, but something else about the Pacific that is instead being captured by the remoteness and population size variables (that is, an omitted variable problem). That concern could have been allayed to some extent by showing that remoteness and population size were related to aid effectiveness when the Pacific countries were excluded.

Now, putting that aside and taking the results as given, the problem for aid agencies is that remoteness and small population size are not things that can be easily changed. It's not a question of changing the observable (and measurable) characteristics of aid projects, to make them more effective. However, Wood et al. are not so easily dissuaded, recommending that:

...as the main constraints to effective aid are constraints that cannot be shifted or which should not be changed, donors ought to focus foremost on adapting their practice. Successful adaptation is not likely to involve changes in sectoral focus or project size or duration, but rather working in a manner appropriate to giving aid in difficult circumstances...

More investment in building donors’ own expertise in the region will also likely help, as will more investment in gold standard evaluations that allow donors to learn from the specific challenges confronting their work in the Pacific.

That strikes me as a recommendation that could have been made without the necessity of going through the research exercise. It almost goes without saying that adapting practice to location conditions and building expertise in the region are important. In that case, aid agencies may simply have to accept that it will take more time and effort to conduct effective aid projects in the Pacific, and/or that those projects will not be as effective as they are in other developing countries.

Saturday, 7 May 2022

Blogging and the ethics of critique

Berk Özler at Development Impact has a very interesting and thought-provoking post about senior researchers blogging about, and critiquing, research work by junior researchers. Özler wrote:

The objection is that senior people should not be criticizing papers by juniors. The latter, whether grad students or assistant professors may have a fair amount riding on the work in question and the platform that the senior person has and the inequality in power makes such criticism unfair. But this is deeply unsatisfactory: can a senior researcher, however defined (by tenure, age, success in publications, otherwise fame, etc.), never discuss the work of junior people publicly? I don’t think of myself as senior but, very unfortunately, others do. I’d like to shed this persona and just be “one of the researchers,” who can excitedly discuss questions that I am geeked about with anyone of any age, gender, etc. But increasingly, the ideas are taking a back seat to who is voicing them, which makes me do a double take.

The related issue, terminology, that comes up is “punching down.” 

Regular readers of this blog will know that I regularly critique the papers and books that I read. Even with papers that I like, and where I think the authors have done a good job, there's often some aspect of it that I wished they explained better, or where I would have approached things differently. If the authors are junior, am I "punching down"? I (usually) treat the critiques on this blog with the same care and attention as I do as a journal reviewer (and I do a lot of reviewing), offering neither fear nor favour to the things I read, regardless of who are the authors. I have adopted the same approach in my new role as Associate Editor at the Journal of Economic Surveys. I'd like to think that I'm tough, but fair.

As for "punching down", that would assume that I am somehow elevated above the authors of the paper or book I am critiquing. I certainly don't have an outsized platform through this blog, given that the regular readership when I am not teaching could comfortably fit in a minivan. However, I am 15 years out from my PhD, and longevity in the profession brings with it a certain amount of experience. Özler makes that point that he doesn't "think of [himself] as senior but, very unfortunately, others do". Possibly I am in that position as well. And the question that Özler raises is particularly important in economics, which has been subject to severe (and warranted) criticism for its negative culture in recent times.

This presents a challenging dilemma for senior researchers. On the one hand, it is unfair when a senior researcher uses their platform (however modest) to attack the work of a junior researcher unfairly. On the other hand, the quality of research overall suffers if (even relatively good) research by junior researchers is immune to any form of criticism. Is there a way forward? Özler offers some further thoughts, from his perspective:

This does have an effect on me as a long-time blogger: how do I stay an effective public intellectual, which means, borrowing from Henry Farrell’s “In praise of negativity” in Crooked Timber, “no more or no less than someone who wants to think and argue in public.” That’s me! The other day philosopher Agnes Callard said “ARGUING IS COLLABORATING.” My best papers involved countless hours of vehement arguing with my co-authors: if you’re right, you can’t sleep because you want to convince your colleague. If you’re wrong, you can’t sleep because you can’t believe you missed that point. Either way, though, you get up the next morning and go talk to your colleague – either to argue more or to concede that you were wrong. So many times, a co-author and I slept on a discussion, only to meet the next day intending to take up the other’s position. It’s fun to argue about important development research topics in public – that’s why I blog and, more importantly, that’s how I mostly blog. If I have to worry about the potential blowback after a post because I blogged about a junior person’s paper, it is a significant disincentive for me to write.

I have some sympathy for that view. I am part of the community of researchers interested in particular topics. I advance my views on specific research because it interested me, usually based on topic, or sometimes based on the research methods employed. For the most part, I blog for myself as much as I do for others. Otherwise, I would likely choose different topics to blog about (note the difference in topics between when I am teaching, and my audience shifts to predominantly first-year students (as it will in July), and when I am not teaching). Blogging about research creates an aide memoire for later reference. I've lost track of the number of times I've been in conversation and thought, "I've read something on that", and a quick search of my blog has served as an effective reminder and a link to relevant research. Storing my critiques as part of that process creates an efficiency, as I don't have to carefully re-read a paper to recognise its weaknesses.

Is it "punching down" for me to critique papers I have read? I don't think so, and the comments to date at the bottom of Özler's post haven't changed my mind.

Friday, 6 May 2022

What World of Warcraft could have taught us about epidemics

I was interested to read this 2007 article by Eric Lofgren (Tufts University) and Nina Fefferman (Rutgers University), published in Lancet Infectious Diseases (ungated here). Lofgren and Fefferman look at the interesting case of an epidemic that suddenly erupted in the World of Warcraft online role-playing game in 2005. As they outline:

On Sept 13, 2005, an estimated 4 million players... of the popular online role-playing game World of Warcraft (Blizzard Entertainment, Irvine, CA, USA) encountered an unexpected challenge in the game, introduced in a software update released that day: a full-blown epidemic. Players exploring a newly accessible spatial area within the game encountered an extremely virulent, highly contagious disease. Soon, the disease had spread to the densely populated capital cities of the fantasy world, causing high rates of mortality and, much more importantly, the social chaos that comes from a large-scale outbreak of deadly disease...

Is this sounding somewhat familiar? You can read more about the outbreak here (or in the Lofgren and Fefferman article, which has much more detail). While the episode presents an interesting example of unintended consequences, Lofgren and Fefferman highlight the potential for online games to improve our understanding of how epidemics spread, and what might be effective in mitigating their impacts. They note that:

In nearly every case, it is physically impossible, financially prohibitive, or morally reprehensible to create a controlled, empirical study where the parameters of the disease are already known before the course of epidemic spread is followed. At the same time, computer models, which allow for large-scale experimentation on virtual populations without such limitations, lack the variability and unexpected outcomes that arise from within the system, not by the nature of the disease, but by the nature of the hosts it infects. These computer simulation experiments attempt to capture the complexity of a functional society to overcome this challenge.

Online gaming worlds may even have enough social elements to mimic real world responses to a disease outbreak:

In the case of the Corrupted Blood epidemic, some players - those with healing abilities - were seen to rush towards areas where the disease was rapidly spreading, acting as first responders in an attempt to help their fellow players. Their behaviour may have actually extended the course of the epidemic and altered its dynamics - for example, by keeping infected individuals alive long enough for them to continue spreading the disease, and by becoming infected themselves and being highly contagious when they rushed to another area.

Lofgren and Fefferman also highlight some of the practical issues with using online gaming worlds as tools for research:

Studies using gaming systems are without the heavy moral and privacy restrictions on patient data inherent to studies involving human patients. This is not to say that this experimental environment is free from concerns of informed consent, anonymity, privacy, and other ethical quandaries. Players may, for example, be asked to consent to the use of their game behaviour for scientific research before participating in the game as part of a licence agreement... Lastly, the ability to repeat such experiments on different portions of the player population within the game (or on different game servers) could act as a detailed, repeatable, accessible, and open standard for epidemiological studies, allowing for confirmation and the alternative analysis of results...

If only this accidental epidemic outbreak in World of Warcraft had been followed up with detailed research (either on this outbreak, or on other controlled outbreaks in that game or other games). What might we have been able to learn about disease dynamics, social distancing, lockdowns or stay-at-home orders, etc.?

Finally, and importantly, should researchers be looking to partner with game developers now, to engineer outbreaks in more current massive multiplayer games, like Rust, Sea of Thieves, or Skyrim?

[HT: Tim Harford]