Tuesday, 15 March 2022

Coin-operated fountains as a public good

I'm in Whanganui this week on a writing retreat. This afternoon, we went for a walk around Rotokawau Virginia Lake. When we arrived, we noticed a pretty sad looking fountain in the front of the lake, called the Higginbottom Fountain. However, after walking around the lake and on our way back to the car, I noticed a sign that made it clear that the fountain was coin operated ($1 for 10 minutes, and $2 for twenty minutes). One $2 coin later:

That got me thinking about paying for public goods, which are goods that are non-rival (one person consuming the good doesn't reduce the amount of the good or service available for everyone else) and non-excludable (if it is available for anyone, then it is available for everyone). Fountains like the Higginbottom Fountain are a good example of a public good. One person admiring the fountain or taking a picture of it doesn't stop anyone else from doing the same (non-rival), and if the fountain is viewable by anyone, it is viewable by everyone (non-excludable).

The problem with public goods is that private willingness-to-pay for the public good is often not enough to ensure that the public good is provided to the socially efficient quantity (on a related point, see here). So, if no one is willing to pay the full cost of operating the fountain by themselves, then no one will pay and no one will benefit from it, even if everyone in total would be willing to pay the cost. The problem is that many people will elect to be free riders - receiving benefit from the public good, even though they aren't paying anything towards it. In this case, until we arrived no one was willing to pay the $2 cost by themselves, but everyone ends up benefiting after we paid the cost ourselves (and if we hadn't, the fountain would not have operated).

In discussing in my ECONS102 class the problem of paying for public goods when there are free riders, I often use street lights as an example. I then make the joke that street lights you could make street lights excludable (and limit the free rider problem) if they were coin-operated - pedestrians put a coin in the base of each light, in order to make the next one light up. Well, it appears that the joke is on me. You can sometimes fund public goods that way.

Friday, 11 March 2022

The effects of training policy makers in econometrics

David Card, Joshua Angrist, and Guido Imbens shared the Nobel Prize in economics last year, for their contributions to the 'credibility revolution' in economics. The revolution involved the adoption of a range of new econometric methods designed to extra causal estimates, and set a much higher standard of what constitutes sound evidence for policy making. However, policy makers have not necessarily caught up. It seems that there could be substantial gains in improved policy to be achieved from training policy makers to use the insights from the credibility revolution.

That is essentially what this 2021 paper by Sultan Mehmood (New Economic School), Shaheen Naseer (Lahore School of Economics) and Daniel Chen (Toulouse School of Economics) sets out to investigate. Mehmood et al. conducted a thorough randomised controlled trial involving deputy ministers in Pakistan. As they explain:

We conducted a randomized evaluation implemented through close collaboration with an elite training academy. The Academy in Pakistan is one of the most prestigious training facilities that prepares top brass policymakers—deputy ministers—for their jobs. These high-ranking policy officials are selected through a highly competitive exam: about 200 are chosen among 15,000 test-takers annually.

There are a lot of moving parts to this research, so it is difficult to summarise (but I'm going to try!). First, Mehmood et al.:

...conducted a baseline survey and asked the participants to choose one of two books (1) Mastering ’Metrics: The Path from Cause to Effect by Joshua Angrist and Jörn-Steffen Pischke or (2) Mindsight: The New Science of Personal Transformation by Daniel J. Siegel...

Actually, they asked the deputy ministers to choose a high or low probability of receiving each of the two books, and then (importantly) they randomised which book each deputy minister actually received (the randomisation is important, and a point we will return to later). The book Mastering 'Metrics (which I reviewed here) is important, because it is the effect of assigning that book that Mehmood et al. set out to test. Mastering 'Metrics is essentially an exposition of the methods that constitute the credibility revolution, and it presents the randomised controlled trial (RCT) as the 'experimental ideal'. However, the treatment doesn't stop only with the assignment of a book to read:

The meat of our intervention is intensive training where we aim to maximize the comprehension, retention, and utilization of the educational materials. Namely, we augmented the book receipt with lectures from the books’ authors, namely, Joshua Angrist and Daniel Siegel, along with high-stakes writing assignments... As part of the training program, deputy ministers were assigned to write two essays. The first essay was to summarize every chapter of their assigned book, while the second essay involved discussing how the materials would apply to their career. The essays were graded and rated in a competitive manner. Writers of the top essays were given monetary vouchers and received peer recognition by their colleagues (via commemorative shields, a presentation and discussion of their essays in a workshop within the treatment arm). Deputy ministers in each treatment group also participated in a zoom session to present, discuss the lessons and applications of their assigned book in a structured discussion.

Performance in the training programme is highly incentivised, not only because of the rewards on offer, but because the grades matter for the future career progress of each deputy minister. So, they had a strong incentive to participate fully. 

Mehmood et al. then test the effect of being assigned the Mastering 'Metrics treatment on a range of outcomes measured four to six months after the workshop, finding that:

While attitudes on importance of qualitative evidence are unaffected, treated individuals' beliefs about the importance of quantitative evidence in making policy decisions increases from 35% after reading the book and completing the writing assignment and grows to 50% after attending the lecture, presenting, discussing and participating in the workshop. We also find that deputy ministers randomly assigned to causal training have higher perceived value of causal inference, quantitative data, and randomized control trials. Metrics training increases how policymakers rate the importance of quantitative evidence in policymaking by about 1 full standard deviation... When asked what actions to undertake before rolling out a new policy, they were more likely to choose to run a randomized trial, with an effect size of 0.33 sigma after completing the book and writing assignment (partial training) and 0.44 sigma after attending the lecture, presentation, discussion and workshop (full training). We also observe substantial performance improvements in scores on national research methods and public policy assessments.

Mehmood et al. also conducted a field experiment to evaluate how much the deputy ministers would be willing to pay for different types of evidence (RCTs, correlational data, and expert bureaucrat advice), and find that:

...treated deputy ministers were much more willing to spend out of pocket (50% more) and from public funds (300% more) for RCTs and less willing to pay for correlation data (50% less). Demand for senior bureaucrats’ advice is unaffected.

Mehmood et al. then ran a second field experiment, where:

First, we elicited initial beliefs about the efficacy of deworming on long-run labor market outcomes. Then, they were asked to choose between implementing a deworming policy versus a policy to build computer labs in schools... Next, we provided a signal - a summary of a recently published randomized evaluation on the long-run impacts of deworming... After this signal, we asked the same deputy ministers about their post-signal beliefs and to make the policy choice again.

Mehmood et al. find substantial effects:

From this experiment, we observe that only those assigned to receive training in causal thinking showed a shift in their beliefs about the efficacy of deworming: the treated ministers became more likely to choose deworming as a policy after receiving the RCT evidence signal. The magnitudes are substantial - trained deputy ministers doubled the likelihood to choose deworming, from 40% to 80%. Notably, this shift occurs only for those ministers whose previously believed the impacts of deworming were lower than the effects found in the RCT study, while those who previously believed the effects were larger than the estimate reported in the signal did not shift their choice of policy.

Importantly, experimenter demand effects are likely to be limited, because as part of the experiment they recommended that the deputy ministers should choose the policy to build computer labs. If anything, these effects are going to be underestimates.

Next, Mehmood et al. test the effects on prosocial behaviour, which might be a concern if you think that studying economics makes students less ethical (see here, or here, or here). On this point:

The administrative data also included a suite of behavioral data in the field, for example, a choice of field visits to orphanages and volunteering in low-income schools. This allowed us to assess potential crowdout of prosociality, an oft-raised concern about the teaching of neoclassical economics... We detected no evidence of econometrics training crowding out prosocial behavior - orphanage field visits, volunteering in low-income schools and language associated with compassion, kindness and social cohesion is not significantly impacted. Scores on teamwork assessments as a proxy of soft skills were also unaffected...

Finally, the randomisation of deputy ministers to books was important, as that provides an indication of the initial preferences for each book. Ministers who preferred the 'Mastering Metrics book could differ in meaningful ways from those who preferred Mindsight, and that might affect the estimated impact of the treatment. Instead, Mehmood et al. note that:

A typical concern in RCTs is that the compliers respond to treatment and we estimate Local Average Treatment Effect (LATE) since we do not observe defiers. It is a plausible concern that people who demand to learn causal thinking may be more responsive to the treatment assignment. Thus estimates of the treatment impacts would be uninformative on those who are potential non-compliers. In our unique experimental set-up, we developed a proxy for compliers through those who demanded the metrics book; we show that the effects are the same for both the high and low demanders... we observe no significant differences between the treatment effects for low and high demanders of metrics training.

There is a huge amount of additional detail in the paper. The overall takeaway is that understanding the importance of causal estimates makes a significant difference to the preferences and decision-making of policy makers, and therefore can contribute to better decision-making. We have these results for Pakistan, but it would be interesting to see if they hold in other contexts. And if they do, then the training of policy-makers should include training in basic econometrics.

[HT: Markus Goldstein at the Development Impact blog]

Wednesday, 9 March 2022

Decision-makers don't dislike uncertain advice, but they do dislike advisors who are uncertain

One aspect of my research and consulting involves producing population projections, which local councils use for planning purposes. Population projections involve a lot of demographic changes, each of which are not perfectly known beforehand, so projections inherently have a lot of uncertainty (for more on that point, see here). [*] Understandably, while planners are trying to make plans for an uncertain future, reducing the uncertainty of that future makes their jobs easier. In my experience, planners are usually looking for projections that give them one single number for the total population (for each year) to focus their planning on. [**] So, it seems natural to me that decision-makers would be averse to uncertainty, and prefer to receive predictions or projections that convey a greater degree of certainty.

It turns out that might not be the case. This 2018 article by Celia Gaertig and Joseph Simmons (both University of Pennsylvania), published in the journal Psychological Science (ungated version here) demonstrates that decision-makers don't have an aversion to uncertain advice at all. Gaertig and Simmons conducted a number of experimental studies where they presented research participants with advice and asked them to make a decision. In the first six studies, using research participants recruited from Amazon Mechanical Turk:

...participants were asked to predict the outcomes of a series of sporting events on the day on which the games were played. Participants in Studies 1 and 2 predicted NBA games, and participants in Studies 3–6 predicted MLB games...

For each of the games that participants were asked to forecast, we told them that, “You will receive advice to help you make your predictions. For each question, the advice that you receive comes from a different person.” Importantly, participants always received objectively good advice, which was based on data from well-calibrated betting markets. For each game, we independently manipulated the certainty of the advice, and, in all but one study, we also manipulated the confidence of the advisor.

Gaertig and Simmons presented research participants with advice, some of which was certain, and some of which was uncertain, with uncertainty expressed in a variety of way, including probabilistically. Research participants were then asked about the quality of the advice they received, and they made an incentivised choice, where they were paid more if they correctly predicted the outcome of the sporting event (the specific outcomes to be predicted varied across the studies). They found that:

As predicted, and consistent with past research, these analyses revealed a large and significant main effect of advisor confidence... Advisors who said “I am not sure but . . .” were evaluated more negatively than advisors who expressed themselves confidently.

More importantly, participants did not evaluate uncertain advice more negatively than certain advice...

Thus, these studies provide no evidence that people inherently dislike uncertain advice in the form of ranges.

Moving on to probabilistic statements of uncertainty, Gaertig and Simmons found that:

As in the previous analysis, there was a large and significant main effect of advisor confidence in all regressions... Advisors who said “I am not sure but . . .” were evaluated more negatively than advisors who expressed themselves confidently. We also found, in Study 6, that advisors who preceded their advice by saying, “I am very confident that . . .” were evaluated more positively than advisors who did not express themselves with such high confidence...

Participants evaluated exact-chance advice (e.g., “There is a 57% chance that the Chicago Cubs will win the game”) more positively than certain advice (e.g., “The Chicago Cubs will win the game”)...

Participants also evaluated approximate-chance advice (e.g., “There is about a 57% chance that Chicago Cubs will win the game”) more positively than certain advice...

In Study 5, we introduced a percent-confident condition, in which participants received confident advice in the form of “I am X% confident that . . . ” We found that participants evaluated this advice the same as certain advice...

The results of the “probably” condition were different, as participants did evaluate advice of the form “The [predicted team] will probably win the game” more negatively than they evaluated certain advice...

Gaertig and Simmons then go on to show that the research participants were no less likely to follow the uncertain advice in their predictions than the certain advice. They also found similar results in a laboratory setting (in Study 7), which also showed that:

...people’s preference for uncertain versus certain advice was greater when the uncertain advice was associated with a larger probability.

In other words, people prefer for advice that demonstrates uncertainty, when the probability is very high (or very low), but not so much when the uncertainty says that the chances of an event are 50-50. That makes some sense. However, it doesn't really accord with the finding of a preference for uncertain advice over certain advice - do decision-makers really prefer advice that says something is 95% certain than saying something is 100% certain? Perhaps they feel that predictions expressed with 100% certainty lack credibility?

Finally, in the last two studies Gaertig and Simmons asked research participants to choose between two advisors, one of whom provided certain advice while the other provided uncertain advice. They found:

...a large and significantly positive effect of the uncertain-advice condition, indicating that more participants preferred Advisor 2 when Advisor 2 provided uncertain advice than when Advisor 2 provided certain advice. This was true both when the uncertain advice came in the form of approximate-chance advice and in the form of “more-likely” advice. When one advisor provided certain advice and the other approximate-chance advice, 82.4% of participants chose Advisor 2 when Advisor 2 provided approximate-chance advice, but only 16.2% of participants chose Advisor 2 when Advisor 2 provided certain advice...

Gaertig and Simmons conclude that:

Taken together, our results challenge the belief that advisors need to provide false certainty for their advice to be heeded. Advisors do not have a realistic incentive to be overconfident, as people do not judge them more negatively when they provide realistically uncertain advice.

It seems that I may have misjudged decision-makers' preferences for certainty. They don't prefer certain advice; they prefer advisors who are certain about their uncertainty. 

*****

[*] Here I'm using uncertainty in its everyday broad sense. Financial economists distinguish between uncertainty that can be quantified (which they refer to as risk), and uncertainty that cannot be easily quantified.

[**] This is an exaggeration of course, in two ways. First, planners recognise that there is uncertainty. However, explaining that uncertainty to elected decision-makers is difficult, so having a single number makes their job easier in that way as well. Second, planners don't only want a single number for the total population. They usually want to know a bit more detail about the age distribution, etc.

Tuesday, 8 March 2022

The state of social science research in New Zealand, and a parting shot from Superu

Last year, the Government released a green paper entitled Te Ara Paerangi - Future Pathways, which set off a process of consultation on the future of New Zealand’s research system, including research priorities, funding, institutions, workforce, infrastructure, and Māori research. You can read Universities New Zealand's submission on the green paper here. I don't have too much to say on it, but it did prompt me to look back at a report that has been sitting on my (virtual) to-be-read pile for some time. That's this report by David Preston, written just as Superu (the relatively short-lived Social Policy Evaluation and Research Unit, which succeeded the Families Commission) was being disestablished.

Social science research has always been the poor sibling in research priority and funding in New Zealand, in comparison with physical and biological sciences, and medical research. There are some ironies to that, which I'll return to later in this post. Preston's report catalogues the failures (and modest successes) of publicly funded social science research institutions in New Zealand. In particular, he notes the demise of the New Zealand Planning Council (1977-1991), the Commission for the Future (1977-1982), the Social Sciences Research Fund Committee (1979-1990), the New Zealand Institute for Social Research and Development (1992-1995), and finally Superu (2014-2017, previously the Families Commission from 2004-2014). The short-lived BRCSS (Building Research Capacity in the Social Sciences) initiative (2004-2010), which I was on the steering group of as it wound down and became eSocSci, also rates a mention. The only two successes that Preston notes are the New Zealand Council for Education Research (NZCER) and the Health Research Council (HRC), both of which were set up in the 1930s are persist to this day.

Based on a study of documentary evidence from annual reports and other sources, as well as interviews with key informants, Preston identifies a number of factors that are associated with the success or failure of these institutions for social science research. In terms of the successful institutions:

Some of the factors which are present in successful social research bodies are those which are common to all successful professional organisations. These include competent professional staff, good management, and adequate resourcing and critical mass for the size of the task they faced. All Long Life institutions, however, also display four other characteristics:

1. a clearly defined field of research

2. well identified research priorities

3. a stable long term funding model, at least for base line funding

4. effective relations with the departmental policy and social service delivery agencies.

And in terms of the unsuccessful institutions:

For government departmental social research units which diminished or vanished in the course of public sector restructuring, no common factor other than the restructuring itself has been identified. They were located in ’less core’ parts of the public sector. Apart from this, no particular pattern is evident.

For social research and advisory bodies outside of main departments, the factors associated with a short existence are one or more of the following:

1. attempting to cover too many different areas of social research

2. not providing the type of information and advice wanted by the government of the day, or providing information or advice at odds with their policy direction

3. lack of an adequate long term base funding arrangement.

Probably the key thing that stands out from this report (aside from the fact that it is clearly a parting shot from Superu, which funded the report), is the highly political nature of social science funding. For example, Preston notes the problems associated with multi-sector research institutions that sit outside of core government services:

While this position outside of government proper gives the institution more independence, it also makes the entity more vulnerable to unfavourable reactions from the government of the day. This is especially so if it is providing advice or information on politically sensitive issues. The government cannot do without its core government departments, but it can do without particular advisory bodies or research institutions.

A related example is the closure of the Social Policy Journal of New Zealand, which had been set up and run by the Social Policy Agency, part of the Department of Social Welfare:

No official reason was ever given for the closure of the Journal. However, informal sources commented that an article about to be published included information which indicated that a statement made by a Minister was inaccurate. Publication of the issue was delayed until public interest in the topic died down and it was decided to cease publication of the Journal, apparently to avoid future difficulties with Ministers.

The development indicates the difficulties of maintaining the ability to publish research findings within a politically sensitive environment in a government department.

All of this suggests that social science research institutions are always in a precarious position, reliant on short-term funding sources and beholden to the political whims of government. Preston summarises the various reviews of social science research that have been undertaken since the 1970s (of which there have been many). One common theme across those reviews is the need for a social science research institution with a sufficient level of baseline funding to maintain core research activity. Preston uses the example of the Brookings Institution from the U.S., which admittedly is not funded by the government, but has had a lasting impact on policy development and is generally well respected.

New Zealand needs quality social science research capacity. At present, there is some capacity in various research units in government, including the Social Welfare Agency (previously the Social Investment Agency), Ministry of Social Development, and elsewhere. However, the main capacity lies in academic institutions and in the private sector, both of which have incentives that do not necessarily lead to good policy-relevant research. In the case of academics, the incentives are to publish in high quality international journals, which tend not to be interested in policy-focused New Zealand research. In the case of the private sector, their incentive is to cater to the needs of paying clients, such as market research and economic research.

The irony is that the government is increasingly focusing research funding and infrastructure on physical and biological sciences, agriculture, and engineering (broadly defined). There is a clear focus on research that provides economic returns. It might seem odd for an economist to argue that we need to refocus resources in a different direction than that, but there are important societal challenges where the natural sciences and engineering are of limited assistance. Improving social wellbeing, and better understanding the impacts of inequality, can't be achieved by ignoring the important contributions of social science. Climate change research also can't ignore social science, because ultimately it is the behaviour of people that will determine the success or otherwise of climate change initiatives. Engineers and natural scientists often have a 'build it and they will come' mindset to the solutions they design, which ignores the complex motivations that real people have.

If New Zealand truly has a goal of improving wellbeing for all New Zealanders, then social science research has an important role to play, and needs appropriate funding and infrastructure to support that role. This is the one thing that most needs to be picked up in the consultation on the government's green paper.