Showing posts with label Causation vs. correlation. Show all posts
Showing posts with label Causation vs. correlation. Show all posts

Wednesday, 3 June 2026

This research doesn’t convincingly show that biodiversity is good for business

I was interested to read this article in The Conversation last month by Paul Griffin (University of California, Davis) and Martien Lubberink (Victoria University of Wellington), mainly because of statements like this:

...firms operating in areas with richer biodiversity are measurably more productive.

I thought, that's interesting. This might be a good example to use in class next trimester to illustrate the difference between correlation and causation. After all, the authors may be correct that firms operating in areas with richer biodiversity are more productive (correlation), but that doesn't mean that biodiversity increases productivity (causation).

And then I read the paper that The Conversation article was based on. And at that point, I decided that I shouldn't use this as an example of the difference between correlation and causation, because even the correlations that they find are shaky at best.

The approach that Griffin and Lubberink take is to look at the relationship between measures of business output and measures of biodiversity. Their measure of business output is sales or gross profit, taken from Stats NZ's Longitudinal Business Database. They generally interpret this as a measure of productivity. And that is the first problem with the paper. Sales can be interpreted as gross revenue, and in some contexts sales may be used as a rough measure of gross output. But sales are not a good measure of productivity, and is not a good measure of the economic value created by a business. The more appropriate measure would be value added, or at least something closer to profit. To see why, consider two firms that both produce a product that sells for $1,000 per unit, and both firms sell 1,000 units per month. Both firms have sales of $1 million per month. Firm A buys the product wholesale at a cost of $800 per unit, then adds a mark-up. The value added of Firm A is $200,000 per month. Firm B buys raw materials of $200 per unit, adds labour of $300 per unit, and then sells the product. The value added of Firm B is $500,000 per month. Firm B creates a lot more economic value than Firm A, and yet measured by sales they are the same. Sales are therefore a poor measure of productivity. Gross profit is less problematic, because it subtracts at least some intermediate input costs, but even gross profit is not a pure measure of value added or productivity.

As a measure of biodiversity, or more accurately as a set of proxies for biodiversity-related conditions and pressures, Griffin and Lubberink use a variety of indicators that they call 'biodiversity abundance markers' (which for some reason they use the acronym BDAs to represent). They aggregate data from a range of sources for their various BDAs (which I will discuss in further detail below), with the data at the SA2 level (SA2s are geographical areas approximately the size of suburbs in urban areas, and larger in rural or remote areas). They note that:

For each SA2, we define a vector of “biodiversity abundance markers” (BDAs), where each ranges from 0 to 100. We denote these ranks as BDA1, BDA2, … , BDAm. We then assign them to an SA2 and, therefore, to the businesses and employees in the same SA2. For a given BDA in an SA2, BDAm = 0 means complete biodiversity loss (high pressure from biodiversity loss) for marker mBDAm = 100 (low pressure from biodiversity loss) is equivalent to an SA2 with an undisturbed or fully intact natural state.

So far, so good. The only issue with that approach is that the measures of biodiversity don't have a natural interpretation, because they are just an index. But we often work with indices - you just need to be cautious about how you interpret the magnitude of the effects. Griffin and Lubberink start by showing the correlation between each of their BDA measures and their measures of business output.

However, then they want to create an overall index of biodiversity, and to do this they:

...multiply each BDA by its SA2 land area and denote the result as an empirical proxy for the natural capital (n) of an SA2 applicable to the businesses operating therein.

Remember that the BDA is an index, bounded between 0 and 100, and it has no natural interpretation in terms of magnitude. So, multiplying the index by the land area of the SA2 is not meaningful, because the BDA is not a measured biodiversity stock per square kilometre. I guess it might make sense if you wanted to calculate a weighted average index, where the weights are based on SA2 land areas, but that isn't what Griffin and Lubberink are doing. Their approach is problematic because it mechanically causes the measured biodiversity to be higher in rural areas ceteris paribus (holding all else equal), where SA2s are larger, and lower in urban areas, where SA2s are smaller. Within urban areas, ceteris paribus it causes higher measured biodiversity in industrial and commercial areas, where SA2s are larger, and lower in residential areas, where SA2s are smaller.

Griffin and Lubberink then aggregate their index-multiplied-by-land-area measures in various ways. The aggregation approach they adopt is fine, but when you aggregate numbers that are not individually meaningful, the result is not meaningful either.

But let's take a step back, because there is another problem. Griffin and Lubberink pitch their analysis as based on a Cobb-Douglas production function. That is fine - a Cobb-Douglas function is a way of relating inputs to output. We already know that their measure of output is faulty. Their inputs are also faulty. Their three-factor Cobb-Douglas function includes inputs of financial capital, human capital, and natural capital.

Griffin and Lubberink measure human capital as the number of employees working in business units in an SA2. That is really a measure of labour input, not human capital. To measure human capital (as well as labour), it would be better to also consider the education level of those employees, since more educated (not to mention more experienced) employees have more human capital. So, their measure is unlikely to pick up the important variation in human capital across SA2s, but it will pick up differences in labour input. But as a measure of combined labour and human capital, their measure will bias downwards measured human capital in urban areas, where education levels are highest, and bias upwards measured human capital in rural and remote areas, where education levels are lowest.

Griffin and Lubberink measure financial capital by the number of business units operating in an SA2. That is not financial capital. That is business density. The relationship between the number of firms and financial capital is not straightforward. An SA2 might have lots of small firms that have low aggregate financial capital, or one large firm that has a lot of financial capital.

Finally, we come back to natural capital, which is measured as noted above. However, some of the measures of biodiversity that Griffin and Lubberink use are better suited than others as a measure of natural capital. The definition of capital is important here - capital is stored up resources that can be used to produce things. Financial capital is stored up savings that can be used in the future. Human capital is stored up education and experience that can be used in the future. So, capital is a stock. It is not a flow.

Now, let's consider the BDA measures one-by one. The first (BDA1 - Land Use) is "1 - the ratio of the number of agriculture and forestry business (primary industry) units in an SA2 to the total number of business units in an SA2". This is not really a measure of land use, because it isn't measured in terms of land. The relative size of the businesses is not taken into account, so many small farms would increase this measure compared to fewer large farms. It is also difficult to see how this is a measure of biodiversity.

The second measure (BDA2 - Infrastructure) is "1 - the rank of the number of business units in an SA2 to the land area in km2 of an SA2 divided by the total number of SA2 observations". It is difficult to understand why this BDA is measured as a rank, whereas BDA1 was not. It is also difficult to see how the number of firms is a measure of infrastructure, or how it relates to biodiversity. This measure will tend to be lower in urban areas, where many small businesses are clustered, than in rural areas. So, this is likely just a measure of urbanicity, not a measure of infrastructure or biodiversity.

The third measure (BDA3 - Mining) is "1 - ratio of the number of mining business units in an SA2 to the total number of business units in an SA2", Like BDA1, this doesn't account for the size of the mines. If you have a small quarry, that counts the same in this measure as the enormous Martha Mine in Waihi. It is more plausibly a measure of (negative) biodiversity than the other measures though. Or at least it would be, if the size of the businesses were taken into account.

The fourth measure is climate change in two forms (BDA4a - Climate Change, and BDA4b Heat Spell Anomaly), which are measured as "the sum of the presence of a heat spell, cold spell, rain spell, or wind spell in an SA2 divided by 4" and "the rank of the heat spell anomalies in an SA2 divided by the total number of SA2 observations". They measure heat spells, cold spells, rain spells, and wind spells as the number of days on which the measured variable (temperature, rain, or wind) falls above (or below, for cold spells) the 'rolling mean 95th percentile' (it isn't clear what the term 'rolling mean 95th percentile' actually means). It isn't clear why adding those four up makes any sense, but perhaps you could just label them weather anomalies. In the second form of this measure, like BDA2 it isn't clear why the rank is used when the actual number of heat spells could be used instead. Again, this isn't really a direct measure of biodiversity, but to the extent that weather anomalies impede biodiversity, it may be a reasonable proxy.

The fifth measure (BDA5 - River Diversity) is "River condition × 100, where River condition = Percentage of insect and related species in an SA-located river compared to all possible species". This is probably the clearest actual biodiversity measure in the paper. However, it is still a narrow one, because although it captures the presence of insect and related species in rivers, it doesn't capture biodiversity more generally. It also doesn't consider the abundance of species. 

The sixth measure (BDA6 - Drinking Water) is "An indicator of the average improvement (higher BDA) or deterioration (lower BDA) in drinking water quality in a region based on periodic water testing". This measure is not a stock, it is a flow. It is a change over time, which gives no indication of the stock available for businesses to use in production. Since Griffin and Lubberink are interested in natural capital as a stock, it would have been better to use the level of drinking water quality, rather than the change in drinking water quality over time. This measure also has problems of reverse causality. Griffin and Lubberink use their measures as if they are business inputs. However, water quality is likely an output of business. Consider a dairy farm that reduces the water quality in a nearby stream. They have the causal relationship backwards when this variable is included in the analysis.

The seventh measure (BDA7 - Plant Diseases) is "1 - percentage of plant diseases in an SA-unit compared to all possible plant diseases". Let's put aside the impossibility of measuring "all possible plant diseases". This might be a useful measure of (the lack of) biodiversity, but it would be better to directly measure plant biodiversity, rather than proxying for it by plant diseases.

The eighth measure (BDA8 - Matauranga) is "Percentage of SA2 population of Māori descent". This is a socio-cultural proxy for relationships with nature, not a measure of biodiversity.

The ninth measure (BDA9 - Population Density) is "1 - the rank of the population density in an SA2 divided by the total number of SA2 observations". Again, it isn't clear why the rank is used here, rather than actual population density. Also, like BDA2 this is a measure of urbanicity, not biodiversity.

The tenth measure (BDA10 - Possum Count) is "1 - the rank of the possum count in an SA2 divided by the total number of SA2 observations". Again, it isn't clear why the rank is used here, rather than some standardised measure of the actual possum count, or possums per land area. It is an indicator of biodiversity though, since more possums would typically mean fewer of other species.

Finally, the eleventh measure (BDA11 - Non-Drought Probability) is "1 minus the ratio of the number of drought weather events in an SA divided by the sum of the number of drought plus non-drought weather events in an SA2". It's not clear what a 'non-drought weather event' is, or why this is a sensible measure. This measure is probably correlated with the climate change measures in BDA4 in any case.

So, across the eleven (or twelve, if you treat the two BDA4 measures as separate) BDA measures, there are only three that are really measures of biodiversity, and there are a few that are likely to meaningfully correlated with biodiversity. The issue is not that every variable must be a perfect direct measure of biodiversity. Empirical research often relies on proxy measures. The issue is that the interpretation should match the proxy. A variable that measures urbanicity, business density, ethnicity, or weather anomalies may be related to biodiversity, but it is not itself biodiversity. If those variables are then combined into a single measure of 'natural capital', the interpretation becomes difficult. The estimated relationship may reflect biodiversity, but it may also reflect a mix of urbanicity, industry mix, infrastructure, climate, or demographic composition. Conflating urbanicity with biodiversity is an especially clear problem for Griffin and Lubberink's analysis, given that they multiply their BDA measures by SA2 land area when constructing their overall measure of natural capital, as I noted earlier.

Finally, Griffin and Lubberink attempt to exploit what they describe as a quasi-natural experiment. The idea is that a number of government policy changes in 2016 and 2017 were intended to improve the environment. If these policies successfully increased biodiversity, then the relationship between biodiversity and business output should become stronger after those policies were implemented. However, this is not a particularly convincing identification strategy. The policies were national, so there is no obvious untreated control group within New Zealand. The test is essentially asking whether the relationship between natural capital and business output changed after 2016 or 2017. But many other things could also have changed around the same time, including macroeconomic conditions, industry conditions, investment decisions, business confidence, and local economic trends. Moreover, the policies themselves may have affected firms through channels other than biodiversity, not least through expectations about future policy changes. That makes it difficult to interpret any post-2016 or post-2017 change as evidence that biodiversity caused higher business productivity. This part of the analysis instead shows that the estimated association between natural capital and business output is not stable over time, and that might be due to policy changes or any number of other reasons.

There are other issues that I could pick out as well, such as not including SA2 fixed effects in their analysis (so that time-invariant differences between SA2s are not controlled for). To be fair, including SA2 fixed effects would absorb much of the cross-sectional variation in biodiversity that the authors are trying to use. But that is exactly the problem, because without SA2 fixed effects, the estimates may reflect other time-invariant differences between SA2s, and not differences in biodiversity.

The overall takeaway from this paper is not that correlation is not the same as causation, it is that if you want to demonstrate correlation, you first need to use the right data in the right way. Biodiversity might be good for business. Business might be good for biodiversity. This research doesn't convincingly estimate the relationship between biodiversity and business output.

Monday, 23 March 2026

The relationship between obesity of politicians and corruption is correlation, not causation

Not every correlation between two variables represents a causal relationship. Even if we can tell a compelling story about why a change in one variable might cause a change in another, that doesn't make the relationship causal. Sometimes a correlation actually results from something other than the story you tell. Sometimes the correlation is just random noise (a spurious correlation). So, we should be cautious when interpreting correlations.

I was reminded of this when reading this 2021 article by Pavlo Blavatskyy (University of Montpellier), published in the journal Economics of Transition and Institutional Change (sorry, I don't see an ungated version online). The article even generated a small debate, with a comment by György Márk Kis, and then a reply by Blavatskyy, appearing in the same issue of the journal.

In the original article, Blavatskyy looks at the relationship between the body mass index (BMI) of politicians in a country and the Corruption Perceptions Index by Transparency International. The data Blavatskyy uses is for 2017, and the sample of countries is limited to 15 post-Soviet countries (Armenia, Azerbaijan, Belarus, Estonia, Georgia, Kazakhstan, Kyrgyzstan, Latvia, Lithuania, Moldova, Russia, Tajikistan, Turkmenistan, Ukraine, and Uzbekistan). The argument for why this correlation matters is explained in Blavatskyy's reply to Kis:

One common form of corruption/lobbying is inviting governmental officials to lavish banquets with excessive consumption of food and drinks... Corrupt politicians frequenting such banquets might risk gaining extra weight. This ‘hedonic theory of corruption’ postulates the existence of a positive relationship between median body mass index of public officials and the level of grand political corruption in society.

So, Blavatskyy is able to tell a good story for why greater corruption would cause higher BMI among politicians. However, that doesn't mean that the relationship is causal. Even though the correlation between perceived corruption and median politician BMI is clear, from Figure 1 in the original paper:

Low numbers in the Corruption Perceptions Index represent higher levels of perceived corruption. So, this figure shows that countries where the politicians have higher median politician BMI have higher levels of perceived corruption.

Kis took issue with a number of things in the paper. First, why those 15 countries? Why not all countries? Kis shows that if you separate the 15 countries in Blavatskyy's sample by their geographic location, you get different relationships within each subsample. However, the broader question is not what happens when you look at subsamples, but does this relationship hold if you add more countries to the sample? Neither Blavatskyy nor Kis answer that question. We should also wonder whether there is something special about 2017 that leads to this correlation. Does it hold in other years?

In his reply, Blavatskyy doesn't really address those two points (narrow sample, and a single year) in a convincing way. Instead, he narrows the sample even further to look at changes in politician BMI and perceived corruption for just one of the countries in his sample, Ukraine. In that analysis, he again shows a correlation between corruption perceptions and politician BMI, in this case over time for Ukraine. However, that simply raises the question of: why Ukraine? Why didn't he look at other countries in his sample in that way? And just because Ukraine shows a correlation over time, that still doesn't demonstrate a causal relationship.

Kis also takes issue with the machine learning algorithm that Blavatskyy uses to estimate the BMI for politicians in his sample. Kis notes that the accuracy of the algorithm is quite dubious (my words, not Kis's), with:

...errors of at least 5.5 in 21.1% of the time.

That's an error in the estimated BMI of 5.5 in over 20 percent of cases. That extent of measurement error would be problematic. To that, I would add that it is unclear whether the training sample that the machine learning algorithm was trained on included people from post-Soviet countries. The relationship between facial features and BMI could well be ethnic-specific, in ways that systematically bias the results. We have no way of knowing. And Blavatskyy didn't address this point in his reply.

Now, the point of this post is to focus on correlation or causation. From what I have seen, this seems a likely candidate for confounding. There are any number of variables that might increase politician BMI and increase corruption, without corruption being a cause of higher politician BMI. As one example, a country with high inequality might simultaneously have high corruption (with petty officials willing to take bribes to supplement their low incomes) and high politician BMI (since politicians would likely be among the wealthy class in society). Blavatskyy doesn't consider confounding variables such as inequality, or differences in age distribution, or differences in average BMI in the population, or regional differences in diet, in his analysis.

Now, to be fair to Blavatskyy, he doesn't adopt a causal interpretation of his results (except in his response to Kis, as I quoted above). Instead, Blavatskyy argues that, if BMI and perceived corruption are correlated, then we might infer how much corruption is being experienced in a country by looking at the median BMI of its politicians. However, even that inference is problematic, and Blavatskyy should know why. He gives the example of Swiss watches in China as a proxy for corruption, but then notes that:

...the rise of social media and Internet anti-corruption platforms in 2011–2012 made it no longer possible to measure grand political corruption through visible luxury Swiss watches. Luxury Swiss watches could still be a popular expenditure of corrupt governmental officials, but these officials are now more careful not to reveal their Swiss watches to the general public.

When politicians realised that their Swiss watches were giving away their corruption, they stopped showing off their Swiss watches. If politicians realised that their expanding waistlines were giving away their corruption, wouldn't they invest more in personal trainers (or liposuction)? As soon as this correlation was used for inference, the correlation would likely start to break down. This again illustrates the limited usefulness of such proxies.

Correlation does not imply causation. And sometimes, correlation today does not imply correlation in the future. We need to be much more cautious when considering analyses like this one.

Saturday, 7 March 2026

Perceptions of inequality and satisfaction with democracy

Last week, my ECONS101 class covered (among many other things) the faulty causation fallacy. This occurs when we observe two variables that appear to be related to each other (they are correlated), but a change in one of the variables does not actually cause a change in the other variable (there is no causal relationship). We might observe a relationship between two variables (call them A and B), and it might be because a change in A causes a change in B, in which case the relationship is causal. But even if we can tell a really good story explaining why we think a change in A causes a change in B, that in itself doesn't make it true. We might observe that relationship because a change in B causes a change in A (we call this reverse causation). Or, we might observe that relationship because a change in some other variable causes a change in both A and B (we call this confounding). Or, the two variables might be completely unrelated, and the observed relationship happens by chance (we call this spurious correlation).

To illustrate this, I'm going to use the example of the research in this 2024 discussion paper by Nicholas Biddle and Matthew Gray (both Australian National University). They also wrote a non-technical summary of their paper on The Conversation. Biddle and Gray look at the relationship between perceptions of income inequality and faith in democratic institutions. To be fair to them, they do say in the paper that "This does not, however, demonstrate a causal relationship from views on inequality to views on democracy". However, most of their interpretations and their policy recommendations assume that the relationship is causal. For example, they conclude that:

The fundamental issue identified in this paper is that the Australian population has identified the income distribution in Australia as being unfair, and that this appears to be impacting views on democracy.

First though, let's take a step back and look at the research. Biddle and Gray use data from Waves 5 and 6 (from 2018 and 2023 respectively) of the Asian Barometer Survey (with a sample size of over a thousand in each wave for Australia), as well as from the ANUPoll surveys, which is a quarterly survey of public opinion run by the Social Research Centre at ANU. For the ANUPoll, they use the January 2024 data, which includes data from over 4000 respondents.

First, from the Asian Barometer, Biddle and Gray find that there is substantial concern about inequality:

In both waves 5 and 6 of the survey, respondents were asked ‘How fair do you think income distribution is in Australia?’... more Australians think that the income distribution is unfair or very unfair (60.5 per cent) than think it is fair or very fair. This gap has widened slightly since 2018, particularly in terms of those who think the distribution is very unfair as opposed to just unfair.

Second, in the ANUPoll data, they find that:

Combined, 30.3 per cent of Australians were not at all or not very satisfied with democracy in January 2024 (compared to 34.2 per cent in October 2023). This is still well above the January 2023 levels of dissatisfaction (22.9 per cent) and even more so the March 2008 levels (18.6 per cent).

So over time, Australians' perceptions of inequality have gotten worse (they think the income distribution is less fair), and they are less satisfied with democracy. It is reasonable, then, to ask whether those concerns about inequality affect people's faith in democratic institutions. Biddle and Gray next look at that relationship, using to the ANUPoll data, and find that:

There is a very strong relationship between views on income inequality in Australia and views on democracy...

Their model (shown in Table 1 in the paper [*]) shows that the most negative views of the income distribution are associated with negative satisfaction with democracy, while the more positive views of the income distribution are associated with positive views of democracy.

So, there is a strong correlation between perceptions of inequality and satisfaction with democracy. But is that just a correlation, or is there a causal relationship? We can tell a good story here (and Biddle and Gray do that). People who are less satisfied with the income distribution may lay some blame on government, and therefore their satisfaction with democracy falls.

Before we conclude that this relationship is causal though, let me lay out some alternatives. First, perhaps people who are less satisfied with democracy become less satisfied in general with many aspects of society, including the income distribution. In this case, there could be reverse causality. Second, perhaps people who are less satisfied with life in general express less satisfaction with many aspects of life and society, and so they answer more negatively when asked about the satisfaction with democracy, and they answer more negatively when asked about their views of the income distribution. In this case, there would be confounding. Third, perhaps satisfaction with democracy is declining over time for some reason, and views about the income distribution are becoming more negative for some completely different reason. But they look like they are related because they are both trending downwards. In this case, there would be a spurious correlation between perceptions of inequality and satisfaction with democracy.

It isn't straightforward to see two variables that appear to be related, and assume that a change in one of those variables causes a change in the other variable. Economists and other researchers have developed a number of statistical tools and experimental methods to try and tease out when a correlation really is demonstrating a causal relationship. Biddle and Gray haven't done that. It might be that negative perceptions of inequality reduce satisfaction with democracy. By itself, this research doesn't allow us to conclude that.

[HT: The Conversation]

*****

[*] Table 1 in the paper actually has an error. The explanatory variable in the table is labelled as satisfaction with democracy, when that is actually the dependent variable. It is perceptions of inequality that is the explanatory variable.

Saturday, 12 July 2025

Is poverty a driver of crime?

This week my ECONS101 class covered, among other things, the difference between causation and correlation. When two variables appear to move together, the relationship might be causal. We might even be able to tell a good story of why the relationship is causal. However, there may be other explanations for the relationship.

For example, take the relationship between poverty and crime. It is well established that poor people commit more crime. Is that because poverty causes people to commit more crime? Many people think so. Being poor means that people lack access to resources, and they may try to obtain those resources through crime. Alternatively, perhaps being poor leads to anger, frustration, or resentment, which leads poor people to commit more crime.

On the other hand, perhaps there is reverse causation - people who commit more crime may make themselves poorer. Being caught and punished through fines or imprisonment will reduce a person's financial resources, making them more likely to be poor. Or, perhaps there is some confounding (or a common cause), and both poverty and crime are related to some third variable. Education is a possibility, since people with more education tend to earn more (and be less at risk of poverty), and also commit less crime. The local unemployment rate might also be a confounder, since when unemployment is higher, poverty will also be higher, and unemployment is also associated with crime.

Putting all of that together, it isn't clear that there is a causal relationship between poverty and crime. We would need some careful research to try and identify whether the relationship is causal. Fortunately, we have this 2023 NBER working paper by David Cesarini (New York University) and co-authors, to provide us with some evidence. Cesarini et al. look at the impact of winning the lottery in Sweden on criminal convictions. They are fortunate in two ways. First, they have data from the register of criminal convictions on all convictions between 1975 and 2017. And second and more importantly, they are able to match the conviction data to four samples of lottery players. Their sample includes over 350,000 lottery wins by over 280,000 individuals. They also look at the effects on children (of their parents winning the lottery), where they have a sample of over 100,000 children.

The cool thing about this analysis is that Cesarini et al. can look at what happens to criminal behaviour, comparing people who are otherwise similar but win the lottery (and therefore are less poor) with those who did not win the lottery (and therefore are just as poor as before). This analysis should establish the causal impact of financial resources on crime (and therefore by extension also the causal impact of income or wealth on crime).

For the adult analysis, Cesarini et al. find:

...a positive but statistically insignificant effect of lottery wealth on criminal behavior. The point estimate of our main outcome of interest — conviction for any type of crime within seven years of the lottery event — suggests 1 million SEK (about $150,000) increases conviction risk by 0.28 percentage points (10.2%). The 95% confidence interval allows us to reject reductions in conviction risk larger than 0.16 percentage points (5.8%). We find no clear evidence of differential effects across types of offenses.

In other words, lack of financial resources does not cause crime in this sample. If it did, then the increase in financial resources arising from the lottery win would lead to less crime. Or, lack of financial resources does cause crime, the effect is very small. Turning to the effect on children, Cesarini et al. find:

...an effect of parental financial resources on child delinquency close to zero, but non-trivial effects in either direction cannot be ruled out. The 95% confidence interval for the effect of 1 million SEK ranges from a 1.36-percentage-point reduction (12.9%) to a 1.54-percentage-point (14.6%) increase in conviction risk.

Again, the central estimate of the effect of financial resources on crime is zero, albeit with less certainty in the result. The takeaway is that a lack of parental financial resources does not cause crime among children. Both results point to a lack of a causal impact of financial resources on crime. Cesarini et al. conclude that:

Our results therefore challenge the view that the relationship between crime and economic status reflects a causal effect of financial resources on adult offending.

It is likely, then, that the observed correlation between poverty and crime arises as a result of reverse causation (in my view possible, but unlikely), or confounding. That has clear implications for policy, because it suggests that focusing on reducing poverty is unlikely to have any impact on crime. Now, this study was conducted in Sweden, and these results might not hold in other contexts. However, they should make us question more strongly the prevailing view that poverty is a driver of crime.

[HT: Marginal Revolution, back in December 2023

Thursday, 9 January 2025

The impact of Fox News on American democracy

In yesterday's post, I noted a number of opportunities for research on the economics of social media. At least one of those opportunities intersected with the impact of traditional media. So, I was interested to read this new article by Elliott Ash, Sergio Galletta, Matteo Pinna (all ETH Zurich), and Christopher Warshaw (George Washington University), published in the Journal of Public Economics (open access). They look at the impact of Fox News Channel on political ideology of voters, and voting outcomes in the US. Given the well-known right-wing nature of Fox News, it is reasonable to wonder whether it is having a political impact.

Identifying the causal impact of a television channel is somewhat challenging, because where viewership is higher that might be because there are more right-leaning voters. So, the causality might run from voters to viewership, rather than the other way around (reverse causality). However, Ash et al. make use of the fact that the channel number assigned to Fox News varies across markets, and that assignment is random (or, at least, it isn't related to the partisanship of the population in a particular television market). So, Ash et al. use channel position as an instrument for viewership (I'll come back to this point later). They then look at the political preferences of voters using data from:

...the 2000 and 2004 National Annenberg Election Survey (NAES) and the 2006–2020 Cooperative Congressional Election Study (CCES) surveys. Overall, we have data on the preferences of approximately 661,000 Americans.

They also look at the effect on presidential and down-ballot (senate, gubernatorial, and house) elections using a variety of election data sources. They find that:

...from 2006–2008 onwards, a lower Fox News Channel (FNC) position correlates with an increase in self-identified Republican viewers, significant at the 10% level initially and 5% in recent years. A one-standard-deviation drop in FNC’s channel position corresponds to a roughly one-percentage point rise in Republican self-identification... A one-standard-deviation decrease in FNC’s channel leads the average American’s ideological position to shift .03-.04 standard deviations to the right in recent years...

Looking across time, the results are not statistically significant (at the 5 percent level) until 2009-2012, or later, depending on the measure. Turning to presidential elections, Ash et al. find that:

Initially, in the 2000 and 2004 elections, FNC’s impact was minimal, likely due to its growing viewership. By 2008, a one-standard-deviation decrease in FNC’s channel position correlated with a 0.32 increase in the Republican vote share...

And on down-ballot elections:

Overall, the effects of FNC in down-ballot elections are qualitatively similar to, though less precise than, those in presidential elections. In House elections, we see a positive (Pro-Republican) coefficient on FNC for Republicans’ two-party vote share in 2012, and it becomes statistically significant starting 2018. Since around 2012, counties with one-standard-deviation lower FNC’s channel position have about .6 to 1.6 percentage point greater shares in Republican vote.

In Senate elections, we find a positive coefficient starting in 2006, which becomes statistically significant starting in 2014.23 Since then, there have been relatively consistent year-to-year effects of around .6 to .73 percentage points. In other words, a one standard deviation shift in FNC’s channel position increases Republican Senate candidates’ vote share by over half a percentage point.

In gubernatorial races, FNC had a small and statistically insignificant effect until the latter half of the 2010s. In recent elections, however, the effect of FNC in gubernatorial races is similar to Senate races — with a Republican vote share about .5 percentage points higher...

So, it is clear from the results that Fox News Channel is driving a rightward shift in the voting public, and this is having a clear effect on both presidential and down-ballot elections. Ash et al. conclude that:

Given the estimated effect sizes on presidential elections, for example, Fox News could have easily tipped the scales for Donald Trump in 2016.

Perhaps Donald Trump should think himself lucky that his 2022 feud with Fox News didn't escalate too much?

However, there is some reason for scepticism about these results. Although Ash et al. make a good case for their use of channel position as an instrument for viewership, it turns out that they didn't actually run a full instrumental variables analysis:

Ideally, we would provide first-stage and two-stage least-squares (2SLS) results for all years in our analysis. However, there are limitations in estimating and interpreting the 2SLS results. First, we only have data on both the endogenous regressor (FNC ratings) and channel positions for 2005, 2006, 2008, and 2020. Second, the first stage provides evidence that the typical variation induced by the channel positioning is somewhat limited, suggesting that the size of the 2SLS coefficients would be, by construction, unreasonable.

...our main analysis focuses on the reduced form, where the outcomes (e.g., vote shares) are regressed directly on the instrument (FNC channel position)...

So, while they have a good instrument, a lack of data prevents them from making the most of it. And so, I think we would need further evidence before we can conclude that these results demonstrate a causal effect of Fox News Channel on political outcomes in the US. It seems to me that the main holdup to doing a full instrumental variables analysis is not having all the viewership data from Nielsen. So perhaps some wealthy research institution needs to buy that data?

Wednesday, 18 December 2024

Onshore windfarms vs. birds

Several times recently, I've had conversations with others about environmental objections to windfarms, and specifically about their impacts on birds. I've expressed surprise that anyone could believe that large, slow-moving wind turbines could be a threat to birds. It turns out, there is research that supports the negative impacts of wind turbines on birds (see here or here), but that research doesn't actually demonstrate that wind turbines cause a decrease in bird populations. The problem, of course, is that it isn't feasible to run a randomised controlled trial with wind turbines due to cost - placing wind turbines at random across some areas and not others, and comparing the effect on bird populations in both areas. The cost of such an experiment would be enormous.

Fortunately, there are statistical methods that we can use to try and estimate the causal effects. And that is what this new article by Meng et al., published in the Journal of Development Economics (ungated earlier version here) attempts to do. Specifically, they look at the effect of onshore windfarms on bird biodiversity at the county level in China. They have two measures of biodiversity:

Bird abundance is the average number of birds of a given species per checklist observed at a county-month-year-species level. Species richness is the total number of unique species observed in a given county at the month-year level, which better reflects the diversity of the bird populations.

To establish causality, they use a difference-in-differences (two-way fixed effects) model, which essentially compares the difference in bird biodiversity before and after a windfarm is installed, between counties with and without windfarms. However, Meng et al. go a step further, using an instrumental variables approach, instrumenting for the location of windfarms by the interaction between national-level growth in windfarms interacted with county-level average windspeed at 100 metres. That instrumental variables approach should mitigate issues arising from the correlation of wind turbine location and bird biodiversity.

The novelty of this paper is not just in the methods, but in the data that Meng et al. employ. To measure bird biodiversity, they make use of data:

...from the China Birdwatching Report (CBR, similar to the eBird Reference Dataset), a citizen science dataset consisting of reports from users, including information on individual bird trips and associated characteristics, such as the specific date and time, location of a specific trip, as well as species and quality of birds encountered...

Their dataset covers the period from 2015 to 2022, and includes data collated from over 33,000 checklists. They also control for a variety of other variables:

...including average bird observed duration, average temperature, average visibility, average wind speed, total precipitation, average ozone, percentage of natural park areas in the county, average population density, and average night light value.

Using this data and the two-way fixed effects approach, Meng et al. find that:

A one standard-deviation increases in wind turbines (approximately 84 turbines)... in a given county leads to a 9.75% decrease in bird abundance per checklist from the mean value of 5.38...

...while a one standard-deviation increases in wind turbines (approximately 84 turbines) in a given county decreases the number of unique bird species by 17.67% from the mean value of 66...

Meng et al. also find evidence that the impacts are greater on migratory birds than on resident birds (important given that China is a major migration pathway for migratory birds), and that the effects are larger in forested and urban/farmland than for grassland. There is also evidence that the impact is greatest for the largest bird species.

Finally, Meng et al. show that there are effects of windfarms on neighbouring counties (as well as the counties in which the windfarms are located), and that those effects are somewhat smaller in size. That made me wonder why those analyses were not the primary results in the paper, since it seems obvious that birds may move across county borders.

So, it does appear that windfarms might cause a decrease in bird biodiversity. Meng et al. even address a bunch of concerns that jumped out at me as I was reading the paper, especially that birdwatchers, anticipating that there would be fewer birds near windfarms, do less birdwatching in those locations. On that point, Meng et al. note that:

We do not find a significant impact of wind turbine installations on birdwatcher behaviors regarding the submitted number of checklists...

And they further support that with detailed mobile phone GPS data, showing that there were not fewer trips made to the areas of windfarms, relative to areas further away. However, a couple of concerns do remain, but they are rather technical. First, I wondered why Meng et al. used the interacted variable (national growth in windfarms interacted with windspeed), rather than just windspeed alone. They describe this as a "Bartik-like variable", but we should be cautious about whether Bartik instruments are appropriate (see here). Also, two-way fixed effects models have also come in for criticism recently (see here and here). I'm not going to drag you into the technical details (read the links if you're interested). But suffice to say, this will not be the last word on whether windfarms negatively impact bird biodiversity. However, the best quality study we have so far seems to suggest they do.

Sunday, 28 July 2024

Unemployment and trans-Tasman migration

The New Zealand Herald reported earlier this month:

Record numbers of people leaving New Zealand to work in Australia could have a negative affect on the workforce over the medium-term.

A report by economic think tank Infometrics shows Australia’s rate of unemployment was lower than New Zealand’s in the first quarter of this year, which was a break from the average rate between 2014 and 2018 when Australia’s rate was 0.7 percentage points higher than New Zealand’s.

“There is a definite correlation between transtasman migration and the relative labour market performances in New Zealand and Australia,” Infometrics director Gareth Kiernan said in the report.

Correlation doesn't necessarily mean causation. The New Zealand Herald article's title is therefore misleading: "‘Drain’ leaves NZ’s unemployment higher than Australia". Now, there are two problems with the New Zealand Herald article here, especially in terms of the title. First, there could be reverse causation - higher unemployment in New Zealand, and lower unemployment in Australia, causing more migration, not migration causing changes in unemployment. To see why, consider the incentives for workers in New Zealand. If unemployment in Australia is lower than New Zealand, then if wages were similar, Australia would more a more attractive option. Workers would start moving to Australia. Wages are not similar though - they are higher in Australia. That increases the incentives to move from New Zealand to Australia even further. The takeaway is, though, that unemployment differences may be causing migration, not the other way around.

The second issue is that, based on a simple supply and demand model of the labour market, migration could affect unemployment in both countries, but the effect would be in the opposite direction to what the New Zealand Herald suggests. To see why, consider the diagrams below, which show the labour markets of Australia on the left, and New Zealand on the right. In both labour markets, the market wage (W1 in Australia, and WB in New Zealand) is above the equilibrium wage (W0 in Australia, and WA in New Zealand). This means that there is excess supply of labour in both countries. There are more people wanting to work than there are jobs available. That is, there is unemployment in both countries. This excess supply of labour is the difference between QS1 and QD1 in Australia, and the difference between QSB and QDB in New Zealand.

Now consider what happens as workers more from the New Zealand labour market to the Australian labour market, as shown in the diagrams below. Supply of labour decreases in New Zealand from SLA to SLC, and at the market wage, the quantity of labour supplied decreases to QSC. This decreases the excess supply of labour in New Zealand (to the difference between QSC and QDB), so unemployment decreases. In the Australian labour market, the supply of labour increases from SL0 to SL2, and at the market wage, the quantity of labour supplied increases to QS2. This increases the excess supply of labour in Australia (to the difference between QS2 and QD1), so unemployment in Australia increases. So, the migration of workers from New Zealand to Australia should have the effect of decreasing unemployment in New Zealand, and increasing unemployment in Australia, not the reverse.

Now, there are many alternative models of the labour market, aside from the model based on supply and demand for labour. However, I don't think those alternatives would suggest decreases in labour supply would increase unemployment. For example, in a search model of the labour market, fewer available workers in New Zealand might mean that job vacancies remain unfilled for longer, since it would take employers longer to find a suitable worker, but unemployment would be unaffected (on the other hand, wages would increase, because with fewer workers available, each worker has slightly higher relative bargaining power).

So, there may be a correlation between unemployment differences between Australia and New Zealand, and trans-Tasman migration. But that doesn't mean that the migration will make unemployment differences worse.

Monday, 22 July 2024

Peter Gray on the social media-mental health debate

The debate about whether social media has a causal negative impact on mental health, recently inflamed by Jonathan Haidt's book The Anxious Generation (see here and here), continues to rage. In the latest contribution, Peter Gray posted:

When I read, at Jon’s request, a pre-publication draft of the book, I told him I could not support it, and I explained why. I had at that time already looked quite broadly and deeply at the research pertaining to questions about effects of screens, Internet, smartphones, and social media on teens’ mental health and found that, despite countless studies designed to reveal such harmful effects, there was very little evidence for such effects. If you survey the research literature selectively, with an eye toward finding studies that seem to show the effects you are looking for, and if you don’t analyze them critically, you can make what will seem to readers to be a compelling case.  But people who really know the research and have examined it fully and critically will see through it.

Gray summarises several previous posts he has written on the topic, which seem to go against Haidt's evidence that social media caused an increase in teen mental health problems. Gray then focuses on one key part of the evidence base of Haidt's book, which is the sole randomised controlled experiment that Haidt uses:

On pages 147-148 of The Anxious Generation, Jon claims that random assignment controlled experiments have shown that social media is a cause (not just a correlate) of teen suffering. He cites just one example of such an experiment, so I looked it up and read the article. The reference, if you want to look it up, is this: Melissa Hunt et al (2018), “No more FOMO: Limiting social media decreases loneliness and depression. Journal of Social and Clinical Psychology, 37, pp 751-768.

The research participants in this experiment were 143 undergraduate students randomly assigned to either limit Facebook, Instagram and Snapchat use to no more than 10 minutes per day, per platform, or to use social media as usual for a three-week period. Self-report questionnaires were used to assess various indices of subjective well-being, namely their sense of social support, fear of missing out, loneliness, anxiety, depression, self-esteem, autonomy, and self-acceptance before, during, and after the three-week period.  The researchers also assessed the participants’ actual social media usage throughout the study by requiring them to submit screen shots showing accumulated use of these media.

The most damaging flaws with this study, which should be obvious to any social scientist, is there are no controls for demand effects or placebo effects. I’ll describe these separately.

Social scientists have shown repeatedly that when research participants can guess the purpose of an experiment and guess the researchers’ hypothesis, they are generally motivated, consciously or unconsciously, to support that hypothesis. In other words, they are likely to believe, or at least claim, they are experiencing what they assume the researcher expects them to experience, to prove the hypothesis correct. This is called the demand effect...

Now the placebo effect. This refers to the simple fact that when people believe they are doing something that will make them feel better, that belief by itself makes them feel better.

Gray points out that the sole study that Haidt relies on to show causal effects of social media on mental health is highly likely to be subject to both the demand effect and the placebo effect, and therefore cannot be relied on as strong evidence of what it claims to be showing. The debate on this book and its thesis is looking increasingly like resolving against the book's claims, or at the very least against the strength of its claims.

[HT: Marginal Revolution]

Read more:

Sunday, 19 May 2024

Mike Masnick on the social media-mental health debate

There's an ongoing debate about whether social media has a causal negative impact on mental health. The latest iteration of this debate was triggered by the release of Jonathan Haidt's book The Anxious Generation. I wrote briefly about the debate between Haidt and Candice Odgers last month. Around the same time, Mike Masnick wrote a long article on the Daily Beast clearly against Haidt's perspective. Here's one important part of the article:

Reading Haidt’s book, you might think the evidence supports his viewpoint, as he presents a lot of it. The problem is that he’s cherry-picking his evidence and often relying on flawed studies. Many other studies by those who have studied this field for many years (unlike Haidt), find little to no support for Haidt’s analysis. The American Psychological Association, which is often quick to blame new technologies for harms (it did this with video games), admitted recently that in a review of all the research, social media could not be deemed as “inherently beneficial or harmful to young people.”

Two recent studies from the Internet Institute at Oxford used access it had obtained to huge amounts of data that showed no direct connection between screen time and mental health or social media and mental health. The latter study there involved data on nearly 1 million people across 72 countries, comparing the introduction of Facebook with widely collected data on mental health, finding little to support a claim that social media diminishes mental health.

To get around this unfortunate situation, Haidt seems to carefully pick which data he uses to support his argument. For example, Haidt mentions the increase in depression and suicide among teen girls from 2000 to the present. The numbers started rising around 2010, though they are still relatively low.

What’s left out if you start in 2000 is what happened earlier. Prior to 2000, the numbers were on par with what they were today in the late 1980s and early 1990s, when no social media existed. Across the decades, we see that the late ’90s and early 2000s were a time when depression and suicide rates significantly dipped from previous highs, before returning recently to similar levels from the ’80s and ’90s.

It’s worth studying why it dropped and then why it went up again, but by starting the data in 2000, Haidt ignores that story, focusing only on the increase, and leading readers to the false conclusion that we are in a unique and therefore alarming period that can only be blamed on social media.

Masnick also highlights that suicide rates (which are indicative of extreme negative mental health) have not seen an uptick in all countries since 2010, or even in all Western countries, pointing to these data:

I felt a bit obliged to include the figure, since it shown the overall downward trend in youth suicide rates in New Zealand. It doesn't break the data down by gender, and part of Haidt's argument is that the negative effects are concentrated among young women. However, if you look at data for young women for those same countries, you would have to squint really hard to see any uptick in suicide rates starting around 2010:

Masnick concludes that:

In the end, neither the data nor reality support his position, and neither should you. Kids and mental health is a very complex issue, and Haidt’s solution appears to be, in the words of H.L. Mencken: clear, simple, and wrong.

Clearly, there is more to come in this debate. I remain agnostic, but very cautious about claims on both sides that are not supported by clearly causal evidence.

[HT: Marginal Revolution]

Read more:

Wednesday, 15 May 2024

Homebrewing as the gateway to craft brewing

Despite some fluctuations and concerns about reaching the peak, one of the key trends in the brewing industry (both in New Zealand and in most Western countries) has been the rise of craft brewing. However, craft brewing remains quite concentrated in some areas rather than others. What might explain the regional concentration of craft brewing?

That is essentially the question that this 2019 article by Michael McCullough (California Polytechnic State University), Joshua Berning (Colorado State University), and Jason Hanson (History Colorado), published in the journal Contemporary Economic Policy (sorry, I don't see an ungated version online). Specifically, they look at the effect of legalising the homebrewing industry on brewing across states in the U.S. As they explain:

Amendment XXI, ratified in 1933, repealed Prohibition and made the commercial production of beer and other alcoholic beverages legal again in the United States, although it left it to the states to allow and regulate brewing, vinting, and distilling within their borders. Importantly, the amendment solely omitted homebrewing, the brewing of beer at home for personal consumption, from the list of legal activities...

From 1933 to 1978, 13 states affirmed the right to homebrew in spite of the federal ruling... In 1978, President Carter signed H.R. 1337 which legalized homebrewing, although federal law deferred to state statutes. At that time, only an additional nine states opted in to legalize homebrewing. The remaining 28 states gradually legalized homebrewing over the next 35 years, with Alabama and Mississippi being the last in 2013.

McCullough et al. look at how the date that a state legalised homebrewing affected the commercial brewing industry, hypothesising that:

...states that legally restricted homebrewing may have hindered the development of future brewmasters and therefore the expansion of their own brewing industry...

However, there are some challenges here, because states that legalised homebrewing earlier may have done so because of high demand for beer, so any relationship between brewing and legalisation of homebrewing would arise because both are driven by beer demand (a common cause, or confounding). Or, large breweries might lobby for less restrictive laws on all brewing, thereby cultivating a homebrewing culture that would also lead to more demand for their products (reverse causation). So, a simple model that looks at the relationship between legalisation of homebrewing and brewing would not demonstrate a causal relationship, just correlation.

McCullough et al. solve this problem using an instrumental variable model, which involves finding an instrument that is correlated with homebrewing legalisation, but which would have no effect on commercial brewing more generally. They argue that the number of years since each state repealed their antimiscegenation laws (laws prohibiting marriage between different races) is such an instrument, because it represents "that measures a state’s willingness to pass legislation in favor of individual rights", and because these laws were pure-and-simple racism, they aren't related to brewing [*].

Using data from 1970 to 2012, McCullough et al. find that:

...the legalization of homebrewing has a positive effect on the average number of breweries per capita. The estimate suggests roughly 7.1 breweries per 1 million people.

McCullough et al. also show that there is an increase in the growth rate of brewing after homebrewing is legalised. Moreover, when comparing how legalisation of homebrewing affected breweries of different sizes, they find that:

...the change in homebrewing laws had a significant effect on the number of small breweries. There were roughly 5.6 more breweries per 1 million people. Furthermore, the number of breweries is growing over time... Looking at medium-sized breweries... we find that the effect of legalization is smaller and not growing significantly over time...

There is no significant change in the number of larger breweries per million people following changes in homebrewing laws...

These results are consistent with their hypothesis, because if homebrewing leads to the development of brewmasters, you would expect a greater number of small breweries to develop, since that's what the brewmasters would create first (to become a middle-sized or big brewery, you probably have to start out as a small brewery first). McCullough et al. also show that there is a statistically significant effect on craft beer production, where:

...craft production increases significantly following the legalization of homebrewing. The estimated impact is roughly 85,000 barrels per million people.

So, it seems clear that homebrewing is the gateway to craft brewing. As McCullough et al. conclude:

While one cannot draw the conclusion that the mere legalization of homebrewing was the main driver for the existence of the beer brewing industry as it is today, one can say that it would not exist in its current fashion without such political action.

*****

[*] However, as McCullough et al. partially note in the paper, states that are more religious and conservative may be more likely to maintain antimiscegenation laws, and more likely to be in favour of temperance. McCullough et al. wave this away by saying that they control for alcohol laws like Sunday sales bans, as well as state fixed effect, but including those variables is only likely to partially allay concerns about the instrument.

Wednesday, 24 April 2024

Jonathan Haidt and Candice Odgers debate the relationship between social media and mental health

Does social media worsen mental health for young people, especially young women? It has become an article of faith for many that it does. And there is bountiful anecdotal and research evidence that supports the view. Take, for example, the furore that erupted back in 2021 around Frances Haugen's leaking of internal Facebook research showing the negative impacts of Instagram on young women.

I've written on this topic several times before (most recently here, but see the list of links at the bottom of this post as well). My take is that much of the research on social media and mental health, or social media and subjective wellbeing, shows correlation, but not causation. The challenge here is that perhaps people with mental health issues (or people with lower wellbeing) are more likely to use online social networks, in which case there is reverse causality (the causality runs from mental health to social media, not from social media to mental health).

So, I was interested to read this recent article in Nature by Candice Odgers, reviewing the new Jonathan Haidt book The Anxious Generation (which I have yet to read, but it is currently on my Amazon Wish List). Odgers really takes Haidt to task, claiming that all that Haidt is demonstrating is correlation, not causation:

The plots presented throughout this book will be useful in teaching my students the fundamentals of causal inference, and how to avoid making up stories by simply looking at trend lines.

Hundreds of researchers, myself included, have searched for the kind of large effects suggested by Haidt. Our efforts have produced a mix of no, small and mixed associations. Most data are correlative. When associations over time are found, they suggest not that social-media use predicts or causes depression, but that young people who already have mental-health problems use such platforms more often or in different ways from their healthy peers...

Odgers then suggests some alternative explanations:

There are, unfortunately, no simple answers. The onset and development of mental disorders, such as anxiety and depression, are driven by a complex set of genetic and environmental factors. Suicide rates among people in most age groups have been increasing steadily for the past 20 years in the United States. Researchers cite access to guns, exposure to violence, structural discrimination and racism, sexism and sexual abuse, the opioid epidemic, economic hardship and social isolation as leading contributors...

The current generation of adolescents was raised in the aftermath of the great recession of 2008. Haidt suggests that the resulting deprivation cannot be a factor, because unemployment has gone down. But analyses of the differential impacts of economic shocks have shown that families in the bottom 20% of the income distribution continue to experience harm... 

Haidt responded, initially on X, but then in thorough detail in this post on After Babel. He starts by pointing to the range of published evidence:

Zach Rausch, Jean Twenge, and I began to collect all the studies we could find in 2019, and we organized them by type: correlational, longitudinal, and experimental. We put all of our work online in Google Docs that are open to other researchers for comment and critique. You can find all of our “collaborative review” documents at AnxiousGeneration.com/reviews

The main document that collects studies on social media is here:
Social Media and Mental Health: A Collaborative Review

Then notes that:

In that document, we list dozens of correlational and longitudinal studies...

In that document, we also list 22 experimental studies, 16 of which found significant evidence of harm (or of benefits from getting off of social media for long enough to get past withdrawal symptoms)...

In that document, we also list nine quasi-experiments or natural experiments (as when high-speed internet arrives in different parts of a country at different times), eight of which found evidence of harm to mental health, especially for girls and women...

I am not saying that academic debates are settled by counting up the number of studies on each side, but bringing so many studies together in one place gives us an overview of the available evidence, and that overview supports three points about problems with the skeptics’ arguments.

First, if the skeptics were right and the null hypothesis were true (i.e., social media does not cause harm to teen mental health), then the published studies would just reflect random noise... and Type I errors (believing something that is false). In that case, we’d see experimental studies producing a wide range of findings, including many that showed benefits to mental health from using social media (or that showed harm to those who go off of social media for a few weeks). Yet there are hardly any such experimental findings. Most experiments find evidence of negative effects; some find no evidence of such effects, and very few show benefits. Also, if the null hypothesis were true, then we’d find some studies where the effects were larger for boys and some that found larger effects for girls. Yet that’s not what we find. When a sex difference is reported, it almost always shows more harm to girls and women. There is a clear and consistent signal running through the experimental studies (as well as the correlational studies), a signal that is not consistent with the null hypothesis.

Haidt supports this with a further footnote:

Yes, there could be a “file drawer problem” if researchers on one side are systematically discouraged from publishing, so the missing “positive” studies are all sitting in file drawers in researchers’ offices. But because findings of benefits would be unusual and newsworthy, I don’t believe that there is a strong or consistent bias against the skeptics. 
However, simply asserting that there is no file drawer problem is not the same as showing that there isn't. That's where meta-analysis comes in. Haidt could easily conduct a meta-analysis with these studies to demonstrate what the overall effect is, and whether there is evidence of publication bias. In fact, he even cites some meta-analyses that have already been conducted (such as this one, which found "mildly significant" publication bias in one of two tests of bias, with the other being statistically insignificant).

Haidt then goes on to address Odgers' suggested alternative explanations, focusing on her assertion that the Global Financial Crisis explains the sudden change in adolescent mental health. Haidt concludes that:

Odgers has pointed to an alternative causal explanation that A) does not fit the timing in the U.S., B) does not fit the social class data in the U.S., and C) does not fit the international scope of the crisis.

Having satisfied himself that he has rebutted Odgers' critique, Haidt then reiterates some solutions from the book:

In contrast, if leaders and change makers were to embrace my account of the “great rewiring of childhood,” in which the phone-based childhood replaced the play-based childhood, what policy implications follow? That we should roll back the phone-based childhood, especially in elementary school and middle school because of the vital importance of protecting kids during early puberty. More specifically, we’d try to implement these four norms as widely as possible: 

  1. No smartphones before high school (as a norm, not a law; parents can just give younger kids flip phones, basic phones, or phone watches).
  2. No social media before 16 (as a norm, but one that would be much more effective if supported by laws such as the proposed update to COPPA, the Kids Online Safety Act, state-level age-appropriate design codes, and new social media bills like the bipartisan Protecting Kids on Social Media Act, or like the state level bills passed in Utah last year and in Florida last month).
  3. Phone-free schools (use phone lockers or Yondr pouches for the whole school day, so that students can pay attention to their teachers and to each other)
  4. More independence, free play, and responsibility in the real world.

Note that these four reforms, taken together, cost almost nothing, have strong bipartisan support, and can be implemented all right now, this year, if we agree to act collectively.

Even if Haidt is wrong about the causal relationship here, I agree that these reforms are relatively low-cost, and the precautionary principle suggests that they might be appropriate. However, I have argued previously that we should be cautious about regulation that allows parents discretion over their children's social media use. Odgers even partially agrees at the end of her review:

Many of Haidt’s solutions for parents, adolescents, educators and big technology firms are reasonable, including stricter content-moderation policies and requiring companies to take user age into account when designing platforms and algorithms. Others, such as age-based restrictions and bans on mobile devices, are unlikely to be effective in practice — or worse, could backfire given what we know about adolescent behaviour.

It will be interesting to see how this debate progresses. Odgers clearly needs to step things up, because Haidt was very well-prepared for her critique, and had clearly anticipated the points that she (and other skeptics) would raise. I look forward to reading the book after I place my next book order.

[HT: Marginal Revolution for the Odgers article, and Haidt's initial response on X]

Read more:

Saturday, 16 March 2024

Causation vs. correlation in the relationship between ultra-processed food and mental health

The New Zealand Herald reported earlier this week:

According to an annual global report, if you’re after mental wellbeing and a flourishing life, you should pay attention to those who live in the Dominican Republic, Sri Lanka, Tanzania and Panama.

The report, by Sapiens Labs, is available here. However, it was this bit of the article that caught my eye:

According to Sapien Labs, adults’ risk of mental health challenges is four times lower if you have close family relationships - but wealthier countries were least likely to say they were close with many adult family members, at just 23 per cent...

Similarly, there is a strong body of research on the impact of processed food and a growing number of studies around technology use.

“We found that over half of those who eat ultra-processed food daily are distressed or struggling with their mental wellbeing, compared to just 18 per cent of those who rarely or never consume ultra-processed food,” the report stated. This is almost a three-fold increase.

I just talked about the difference between causation and correlation with my ECONS101 class a couple of weeks ago. Everything that Sapiens Labs has found is correlation. Sure, you can tell a plausible story about how ultra-processed foods lower mood and lead to worse mental wellbeing. However, there is also a plausible story going in the other direction (reverse causation) - people with worse mental health might comfort eat, thereby consuming more ultra-processed foods. Just because we observe a correlation between higher ultra-processed food consumption and lower mental wellbeing (a negative correlation), it doesn't mean that ultra-processed food consumption causes decreases in mental wellbeing.

Even worse than that, the report itself (but not the New Zealand Herald article) tries to suggest a link between higher consumption of single-use plastics and lower mental wellbeing. I'm not even sure that you can tell a plausible story linking those two variables in that direction - how would plastic straws, forks, and grocery bags make our mental health worse? This could well be spurious correlation. However, I'd be surprised if there is even a correlation there at all. Many countries (including New Zealand) have recently banned single-use plastics. Have we seen an immediate improvement in mental health in those countries? I thought not.

Just because two variables are moving together (either in the same direction, or opposite directions) that doesn't mean that changes in one variable are causing changes in the other one. No matter how much you might want them to, or how much you are looking for a simple explanation. Correlation is not the same as causation.

Wednesday, 28 February 2024

Challenges in establishing causality in the relationship between alcohol outlets and crime

In my ECONS101 class this week, among other things we discussed the 'faulty causation fallacy'. That occurs when you observe two variables (A and B) that appear to be moving together (either in the same direction or opposite directions), and you assume that a change in Variable A is causing a change in Variable B. You might even be able to tell a really good story about why it is that changes in Variable A cause changes in Variable B. But that doesn't mean that your observation and story about causality is true.

What we observe when we see two variables moving together is correlation. When the two variables move in the same direction, that is positive correlation. When the two variables move in opposite directions, that is negative correlation. [*] Sometimes, when we observe correlation, there really is a causal relationship between the variables. When I push down on the accelerator in my car, my car goes faster. Pushing the accelerator (Variable A) causes a change in the car's speed (Variable B).

However, not all correlations that we observe arise because a change in Variable A causes a change in Variable B. Sometimes, it is the other way around - a change in Variable B causes a change in Variable A. We call this reverse causation. Sometimes, there is some third variable (Variable C), and it is a change in that variable that causes both a change in Variable A and a change in Variable B. We refer to Variable C as a confounder (or a confounding variable). Alternatively, we can say that Variable C is a common cause for both Variable A and Variable B. Finally, the correlation that we observe might be entirely by random chance. In that case, we would say that we have observed a spurious correlation (as in the excellent Tyler Vigen website spurious correlations, which offers up a new classic in the form of correlation #2,204: The number of global pirate attacks is highly correlated with the number of downloads of the Firefox browser - perhaps pirates use Firefox?).

Anyway, I want to illustrate these with an example related to my own research, on the relationship between alcohol outlets and crime. I've published articles on this here and here, with another report here. That research establishes a generally positive correlation between the number (or density) of alcohol outlets and crime. The strength of the correlation varies depending on context - it is different for different locations, and different for different types of alcohol outlets. However, the correlation suggests that where there are more alcohol outlets, there is more crime.

Is this relationship causal though? My earlier research doesn't establish this. However, we can tell a good story, using what is termed availability theory. Availability theory suggests that alcohol consumption depends on the 'full cost' of alcohol - which is made up of the price of alcohol, plus the travel cost of getting to and from the alcohol (such as driving to the alcohol outlet and home). When there are more alcohol outlets in an area, they may compete more vigorously on price, meaning that the first part of the full cost of alcohol is lower. And, when there are more alcohol outlets in an area, consumers don't have to travel as far to get the alcohol, lowering the second part of the full cost of alcohol as well. When there are more alcohol outlets in an area, the full cost of alcohol is lower. And when the full cost of alcohol is lower, people will drink more. And when people drink more, then the amount of crime increases (either because there are more alcohol-impaired victims, or more alcohol-affected offenders). So, this observed relationship could be causal.

On the other hand, there could be reverse causation here. In areas where there is more crime, commercial property rents are lower, and there may be more vacant storefronts. Retailers (including alcohol retailers) looking to set up a store are looking for a vacant storefront, and they will tend to be attracted to low rents. So, perhaps an increase in crime would cause an increase in alcohol outlets, as the crime forces other businesses out of an area?

On the third hand, there could be confounding here. As I noted here, social disorganisation theory is the idea that differences (or changes) in family structures and community stability are a key contributor to differences (or changes) in crime rates between different places (or times). Areas that are more socially disorganised have more crime. Areas that are more socially disorganised are also less able to act collectively to prevent alcohol outlets from opening (or remaining open) in their area. So, social disorganisation might be a confounding variable in the relationship between alcohol outlets and crime, because social disorganisation causes more outlets and more crime.

Finally, the observed relationship could just be a spurious correlation, but spurious correlations tend to arise when you have two variables that are trending over time. In this case, the number of outlets doesn't have an obvious time trend (in some areas it is increasing, and in others it is decreasing), and similarly for crime. So, it seems like there is something more than random chance that leads alcohol outlets and crime to be correlated.

So, we can tell a good story for a causal relationship. However, we can also tell a good story for reverse causation, and a good story for confounding. It requires some careful statistical analysis to disentangle these potential explanations, and that is something that researchers (including myself) will continue to work on. I had an article published in the journal Addiction last year (open access, and I blogged about it here) that shows that at least one potential confounding variable, retail density, doesn't explain the relationship. I also presented at the NZAE Conference a couple of years ago on some further analysis which tentatively suggests that the causal relationship is statistically insignificant (although that research is somewhat hampered by the low quality of alcohol outlets data in New Zealand). There will be more to come on this topic in the future.

*****

[*] This is just one way of conceptualising a correlation between Variable A and Variable B (and I think it is the easiest way). There are other ways we can conceptualise a correlation. For example, if we ignore changes over time, we can observe correlations by looking at variables across different individuals or different areas. In that case, if individuals (or areas) with higher values of Variable A also have higher values of Variable B, that is a positive correlation. And if individuals (or areas) with higher values of Variable A have lower values of Variable B, that is a negative correlation.