Showing posts with label Online data. Show all posts
Showing posts with label Online data. Show all posts

Thursday, 28 May 2026

Try this: Taxed

Today was Budget Day in New Zealand. The government revealed its forecasts of future revenue and its spending plans. There is a good summary of this on The Conversation (disclaimer: I wrote the blurb at the top of that summary).

The problem with the Budget is that the numbers are large, and it is difficult to get a good sense of the relative magnitudes. How do you interpret $1.18 billion in spending on rail network renewal and upgrades?

One of my recent students, Tyler Dunseath, created the Taxed website, that uses your income to work out how much tax you pay (weekly, fortnightly, monthly, or annually), then apportions that tax to the various categories of spending from the government accounts. So, for example, if your weekly income is $1000 before tax, and you don't adjust for ACC, KiwiSaver, or student loan repayments, you pay $165.77 in tax. Of that, $56.79 goes to social security and welfare, $37.40 goes to health, $24.25 goes to education, and so on. The results give you a better sense of how taxes are distributed.

Of course, there are a number of caveats, the biggest of which is that government services are a bundle, and while Taxed might make it seem like you could in theory say, "I don't want to pay $0.32 per week for international peacekeeping", it doesn't work that way. Moreover, a lot of government spending is on services that are public goods and therefore non-excludable, so even if you could opt out of paying for them, you would still receive the benefits of them.

Second, government receives some income that is earmarked for particular purposes. For example, the fuel excise tax is earmarked for the National Land Transport Fund. So, your income tax isn't distributed in exact proportion to the government's spending on different categories, because less of your income tax goes towards transport.

Third, the site doesn't account for the taxes we pay on goods and services (GST, or excise taxes on alcohol, tobacco, or fuel), or the user charges we pay.

With those caveats in mind though, Taxed is a pretty cool way of showing how the government's spending is distributed, and in a way that most people are more likely to understand than the millions or billions of dollars cited in the budget.

Enjoy!

[HT: Tyler Dunseath]

Sunday, 25 January 2026

The Census Tree project

An exciting (and new-ish) dataset offers us an unprecedented opportunity to explore research questions using historical US Census data. When I posted about what's new in regional and urban economics last year, one of the things that was raised was the linking of historical Census records over time. That was based on the work of Abramitzky et al., known as the Census Linking Project (CLP). However, in a recent article published in the journal Explorations in Economic History (open access), Kasey Buckles (University of Notre Dame) and co-authors report on an alternative Census linking dataset that has far larger coverage than the CLP. As they explain:

In the Census Tree project, we use information provided by members of the largest genealogy research community in the world to create hundreds of millions of new links among the historical U.S. Censuses (1850–1940). The users of the platform link data sources—including decennial census records—to the profiles of deceased people as part of their own family history research. In doing so, they rely on private information like maiden names, family members’ names, and geographic moves to make links that a researcher would never be able to make using the observable information...

The result is the publicly-available Census Tree dataset, which contains over 700 million links among the 1850–1940 censuses...

The article describes the creation of the Census Tree dataset, which can be accessed for free online. Buckles et al. also demonstrate the use of the dataset, in a particular application in comparison with the CLP data of Abramitzky et al.:

...who show that the children of immigrants were more upwardly mobile on average than the children of the U.S.-born in the late 19th and early 20th centuries. We replicate this result using the Census Tree, and are able to increase the precision of estimates for each sending country. Furthermore, the Census Tree includes sufficient numbers of links to produce estimates for an additional ten countries, including countries from Central America and the Caribbean. We find that the sons of low-income immigrants from Mexico had significantly worse outcomes on average than sons of fathers from other countries, including U.S.-born Whites. We further extend [Abramitzky et al.] by analyzing the mobility of women in a historical sample, and compare these results to historical estimates for men and modern estimates for women. While the patterns for daughters and sons are broadly similar, differences in marriage patterns contribute to gender gaps in mobility in some countries.

As I noted in this post last year, the ability to link people over long periods of time (including between generations) has opened up a wealth of new research questions. Buckles et al. offers a peek at the range of research that has already been done using the Census Tree dataset (see Appendix B in the paper for a bibliography).

Now, the coverage isn't perfect, and there is still some ways to go. You can evaluate the quality of the dataset based on what Buckles et al. report in their article, but it is clearly better than previous efforts. And importantly:

...we plan to update the Census Tree every two-to-three years to incorporate new information added by FamilySearch users, to include new links... and to implement methodological advances in linking methods that we and others develop.

This seems like a really important resources for researchers in economics, sociology, regional science, and other fields, and not just for those interested in economic history. 

Tuesday, 6 January 2026

Try this: The Opportunity Atlas

It's hard to believe that, in over twelve years of blogging, I have never blogged about any of Raj Chetty's research. That's not because I haven't read it. If anything, it's because it is so detailed that it defies a short blog take. For example, we read two related papers published in the journal Nature (here and here, both open access) in the Waikato Economics Discussion Group back in 2022. Ordinarily, I would follow up with a blog post, but they are so in-depth that I couldn't find the time to summarise them effectively [*]. Three years later, they are sitting in a virtual pile of read-but-not-yet-blogged-about papers [**].

Anyway, Chetty and co-authors have suddenly made it much easier for me to summarise their extensive research on social mobility in the US. That's because you can see the data in action for yourself now, at The Opportunity Atlas. This very cool online tool allows you to see social mobility in action. Social mobility is effectively how much a child's socioeconomic position in adulthood depends on their socioeconomic position when they were growing up.

On The Opportunity Atlas, you can choose from a range of outcomes in adulthood, and see where the mean outcome is, if they grew up in a household at different rankings of parental income (1st, 25th, 50th, 75th, or 100th percentile). You can also look separately by gender and by race (Black, White, Hispanic, Asian, Native American). The interface is quite intuitive to use. For example, here's the basic map of expected (mean) income at age 35, for children who grew up in households at the 25th percentile of parental income:

The red areas, such as the South, have lower social mobility, because children who grew up there in households at the 25th percentile have lower incomes as adults. In contrast, the blue areas (in the north and west) have higher social mobility, because children who grew up there in households at the same 25th percentile have higher incomes as adults.

The tool is very flexible. It's very easy to switch to looking at other outcome variables, and for other percentiles of parental income, as well as zooming in on particular areas. For example, here's the teenage birth rate for Black women who grew up in households at the lowest (1st) percentile of parental income in Los Angeles:

The greyed-out Census tracts are those where there are too few Black women who grew up in the lowest income households for the data to be reported. However, the map shows a band of high teenage birth rates for mothers who grew up in the lowest income households, that stretches from South Central to Compton.

Importantly, the underlying data can be downloaded from the Opportunity Insights website. The cool thing about the data underlying the Atlas is that it is based on the census tract where the child grew up, not the census tract where they live as an adult. That means that the Atlas is showing you the adult outcomes for children who grew up in a particular area, not the adult outcomes of adults who live there today. That is explained in this new article by Chetty (Harvard University) and co-authors, published in the journal American Economic Review (ungated earlier version here).

That article outlines the methods underlying the dataset. In short:

...we use de-identified data from the 2000 and 2010 decennial censuses linked to data from federal income tax returns and the 2005–2015 American Community Surveys to obtain information on children’s outcomes in adulthood and their parents’ characteristics. We focus in our baseline analysis on children in the 1978–1983 birth cohorts who were born in the United States or are authorized immigrants who came to the United States in childhood...

We construct tract-level estimates of children’s incomes in adulthood and other outcomes, such as incarceration rates and teenage birth rates by race, gender, and parents’ household income level—the three dimensions on which we find children’s outcomes vary the most. We assign children to locations in proportion to the amount of their childhood they spent growing up in each census tract. In each tract-by-gender-by-race cell, we estimate the conditional expectation of children’s outcomes given their parents’ household income using a univariate regression whose functional form is chosen based on estimates at the national level to capture potential nonlinearities.

Chetty et al. then go on to show why it matters that we look at social mobility based on the place where children grew up, rather than contemporary poverty rates or adult outcomes, and finally give some short use cases for the dataset. I won't go into detail on those (you should read the paper), but one of the things that Chetty et al. do show is that because the effects change slowly over time, looking at outcomes today for children who grew up in a particular census tract in the 1980s still provides meaningful information that can be used for targeting social programmes today.

It's important to note that the Opportunity Atlas by itself doesn't show us causal estimates of adult outcomes. However, Chetty et al. establish how much of the effect is causal using a couple of different methods: (1) using data from the Moving to Opportunity experiment; and (2) a quasi-experiment that looks at how the effects differ depending on how many years a child was 'exposed' to a particular Census tract). Both methods both imply that roughly 62%) of the observed variation across census tracts reflects causal neighbourhood exposure effects, not just higher-opportunity families sorting into better places.

In the conclusion, Chetty et al. highlight a number of applications where the Opportunity Atlas data has already been used:

For researchers, the Opportunity Atlas data provide a new tool to study the determinants of economic opportunity. For example, recent studies have used the Opportunity Atlas data to analyze the effects of lead exposure, pollution, neighborhood redlining, and the Great Migration on children’s long-term outcomes (Manduca and Sampson 2019; Colmer, Voorheis, and Williams 2019; Park and Quercia 2020; Aaronson, Hartley, and Mazumder 2021; Derenoncourt 2022). Other studies use the Atlas statistics as inputs into models of residential sorting (Aliprantis, Carroll, and Young 2024; Davis, Gregory, and Hartley 2019) and to understand perceptions of inequality (Ludwig and Kraus 2019). The ongoing American Voices Project (https://americanvoicesproject.org/) is interviewing families in neighborhoods with particularly low or high levels of upward mobility to uncover new mechanisms from a qualitative lens.

I can see a number of use cases for this as well. For instance, there is probably a lot of value in using the Opportunity Atlas data alongside the data on racial diversity and segregation from the Mixed Metro project (which also offers data down to the Census tract level). Also related is this from a footnote in the Chetty et al. paper:

Understanding how neighborhood effects change with the composition of the neighborhood is an important question that warrants further work...

This also makes me think (again) that we need more detailed work on social mobility in New Zealand, building on the work of my colleagues Niyi Alimi and Dave Maré (see here). One of the amazing things about Chetty's research is that it is now looking at the neighbourhood (Census tract) level, and that sort of spatial disaggregation offers a lot of opportunity for detailed follow-up research and policy action. And with StatsNZ's Integrated Data Infrastructure, we have the basic framework necessary to do this sort of work in New Zealand as well. We could use that to build our own Opportunity Atlas for New Zealand.

*****

[*] So, in lieu of a separate blog post, here's the short summary of those two papers. In the first paper, Chetty et al. use billions of Facebook friendship links to measure local social capital, especially "economic connectedness" (cross-class friendships). They find that places with higher economic connectedness have much higher upward social mobility. In the second paper, the same group of authors show that cross-class friendship gaps come from both who people are exposed to (whether schools, neighbourhoods, or groups) and "friending bias" (less cross-class befriending even when exposed).

[**] In case you're wondering, there are currently 45 papers in that virtual pile, and it seems to be growing. I'm reading research faster than I'm blogging about it. I might have to start blogging about multiple papers in a single post to keep from falling further behind!

Wednesday, 10 December 2025

People care about whether their data are shared, but not so much where their data are stored

There has been a substantial policy movement in favour of the localisation of data storage over the past decade (for example, see here). Policymakers often justify data localisation policies by appealing to consumers' supposed preference for having their data stored locally. In particular, they refer to privacy concerns, lack of trust in data handling practices in other countries, and preference for supporting local data storage firms. However, the evidence that consumers have strong preferences for data localisation is very thin. In fact, this new article by Jeffrey Prince (Indiana University) and Scott Wallsten (Technology Policy Institute), published in the journal Information Economics and Policy (ungated earlier version here), may represent the first attempt to really evaluate consumers preferences for data storage.

Prince and Wallsten use a discrete choice survey to evaluate preferences for localisation for different types of data. Specifically:

We constructed five different survey structures, one each centered on the respondent’s smartphone, financial institution, healthcare app, smart home device, and social media. The data types we consider include home address, phone number, income, financial activity, health status and activity, biometrics, music preferences, location, networks, and communications. Across the five survey structures and range of data types, we measure the relative value of full privacy (no data sharing) versus sharing only domestically (localization), sharing domestically and internationally (no localization), and sharing domestically and internationally excluding China and Russia (no localization but with limits). We administered each of these five different surveys across seven different countries: the United States, the United Kingdom, South Korea, Japan, Italy, India, and France

Their sample, drawn from Dynata's online panel, is 11,375 respondents, with 325 completed surveys for each of the five survey structures, for each of the seven countries. Each respondent was shown ten different discrete choice questions. In each question, respondents would have been shown hypothetical alternative scenarios about how their data could be stored and shared, and had to pick their preferred alternative. However, the article doesn't make clear how many alternatives the respondent was choosing from in each choice task, nor whether they simply chose the best of the alternatives, or provided a full ranking of all of the alternatives. Those are issues that are consequential for the analysis, but probably don't bias the results in any way.

In terms of data localisation, Prince and Wallsten distinguish between data not being shared at all, and data being stored and shared domestically only, internationally, or internationally while excluding China and Russia. The latter is included because consumers may be more concerned about their data being stored in China or Russia than being stored in other countries. First, Prince and Wallsten find that:

...virtually all of our parameter estimates are highly significant. As these are estimates of (dis) utility from sharing data in one of three ways (domestically only, internationally, internationally except China and Russia) versus not sharing, the consistent, negative and statistically significant estimates imply that respondents across all of our countries are averse to sharing their data.

In other words, people really don't like their data being shared, regardless of how or where it would be shared. However, in terms of data localisation, Prince and Wallsten find that:

...it is evident that there are a just handful of data types for which we find any notable data localization premium: bank balance, facial recognition, home address, and phone number, all with multiple instances, and voiceprint, with one instance.

Interpreting these results, Prince and Wallsten note that:

...the data types for which we find a data localization premium are also the data types for which citizens find the most value in having no sharing of any kind... citizens across our seven countries, by and large, place little to no value in data localization requirements, despite placing value on full privacy for these data (i.e., no domestic or international sharing)...

In addition, there are no differences between sharing internationally, and sharing internationally while excluding China and Russia. If anything, there is some weak evidence that respondents in South Korea and Japan preferred to have their data shared with China and Russia. That’s striking given how often policymakers highlight the dangers of data flowing to China and Russia. Prince and Wallsten conclude that:

Our findings have several implications. First, they suggest that the use of privacy concerns as motivation for data localization laws may be overstated, although there may be some gross welfare gains for some types of data. Our findings also indicate that if international sharing is allowed, restricting prominent authoritarian countries such as China and Russia appears to have little impact on consumer value, at least for a number of highly populated countries...

...our findings do provide a counterweight to any claim that citizens find value from imposing constraints on international data sharing.

It may still be worthwhile for policymakers to insist on data localisation. Of course, this is just one study (albeit the first study) using survey data from an online panel, so we should be cautious about overgeneralising. Nevertheless, based on this study, the argument that data localisation reflects consumers' preferences for data storage does not hold up to scrutiny. If they want to keep pushing data localisation, policymakers will need to lean on geopolitical or protectionist arguments instead.

Sunday, 26 January 2025

Try this: Treasury's Income Explorer shows effective marginal tax rates

A new Treasury Analytical Note by Meghan Stephens, Yvonne Wang, and Liam Barnes presents data on effective marginal tax rates for different families (more on that in a moment). However, one cool thing that the note points to is Treasury's Income Explorer tool, which allows you to graph effective marginal tax rates (EMTRs) based on different historical tax schedules (from 2014 to 2024, and forecast tax schedules for 2025 to 2028. You choose the taxpayer's hourly wage, whether they are partnered (and the partner's work hours and pay rate), the number of children, and whether they are a homeowner or renting (which affects their eligibility for the accommodation supplement). This allows you to look at EMTRs for a whole variety of taxpayers in different situations.

As a quick reminder, the effective marginal tax rate for a taxpayer is the proportion of the next dollar earned that is lost to taxation and to decreases in government transfers, rebates, or subsidies. These rates can get quite high - sometimes over 100 percent (in which case, the taxpayer would be worse off in net terms by earning another dollar). As an example of a high EMTR, consider this graph (made in the Income Explorer tool) based on a single parent (with two children aged 0 and 2) earning $40 per hour in their job, and paying weekly rent of $450:

Notice that the EMTR varies depending on the number of hours worked (shown along the top x-axis) and annual income (shown along the bottom x-axis). There are points in the distribution where EMTRs are low, but others where EMTR is very high. And for this taxpayer, working between 38 and 42 hours leads to an EMTR that is 107.6 percent. The Income Explorer tool even breaks that EMTR down: it is made up of 34.6 percent wage tax (including ACC levies), 27 percent Working for Families abatement, 21 percent Best Start abatement, and 25 percent accommodation supplement abatement. Notice that most of the EMTR (73 percentages points out of 107.6) is made up of reductions in entitlements to government assistance for that taxpayer. This would clearly affect the incentives to work additional hours (at least, between 38 and 42 hours).

Fortunately, the results are not so bad across the board. Stephens et al. show that less than six percent of all taxpayers have EMTRs greater than 50 percent, as summarised in this table:

The majority face an EMTR between 25 percent and 50 percent, and for most people, their EMTR is equal to their marginal income tax rate (plus ACC levies).

Understanding EMTRs is important for understanding the incentive effects of the tax and transfer system. Treasury's Income Explorer is a great tool for visualising EMTRs (as well as replacement rates, participation tax rates, and other complementary measures). Try it out for yourself!

[HT: Les Oxley for the analytical note]

Monday, 3 June 2024

The consequences of low-quality alcohol licensing data

I've been meaning to post on this for a little while now. Back in April, Eric Crampton pointed to Police and the Medical Officer of Health using incorrect data on the number of licences in Wellington as part of their submissions opposing alcohol licences. Sadly, for those of us who know about the alcohol licensing data in New Zealand, this sort of outcome won't come as a surprise.

Current alcohol licence data is available for free from the Ministry of Justice website. However, the quality of that data is somewhat questionable. It doesn't take long to realise that there are problems. For example, I downloaded the latest data, and looked at Waikato District (which I know reasonably well, as I am a Commissioner of the District Licensing Committee there). Looking only at on-licences, McGinty's bar in Huntly appears in the list twice. Fortunately, it looks like that's the only duplicate in the Waikato data for on-licences, but it would not surprise me at all to find out that there are many duplicates for Wellington. So, simply getting the free data and counting the number of licences in that data is going to over-state the number of licences.

I think that this particular issue arises because, when a new licence is granted to new owners of an existing premises, the previous licence remains in the dataset until it expires. Once you know that it is an issue, identifying duplicates and removing them is straightforward (although not helped by the dataset having only street names and not street numbers in the address fields).

However, there is a bigger issue with the dataset. It is only updated when the local council updates the Alcohol Regulatory and Licensing Authority (ARLA, which is part of the Ministry of Justice). If the local council doesn't send updates very regularly (or at all), then the dataset can quickly get woefully out of date.

So, as one example, according to the latest dataset, there is only one licensed premise in the entirety of the South Taranaki District, and only two in the entirety of Waitaki District. Obviously, the data are incorrect there, and almost certainly because those districts are not updating ARLA when new licences are issued, or existing licences are renewed. So, the licences drop out of the dataset as they 'expire', when they are really being renewed and ARLA just doesn't know.

This is essentially the reason that I have hit 'pause' on research on alcohol outlets across the country as a whole (while maintaining some research in Hamilton City and South Auckland, where I know the quality of data is high). We really need a more reliable source of data than the dataset that ARLA maintains and makes available through the Ministry of Justice website. To be clear, I don't think it's ARLA's fault here at all. I strongly suspect that the dataset is sub-par because ARLA aren't being supported by the councils actually updating them, as is mandated by the Sale and Supply of Alcohol Act. A large part of the problem is probably that there doesn't appear to be any sanction against councils that fail to update ARLA.

One day, when I have a lot of spare resources (especially time, but also funding), we may be able to pull together a reasonable dataset that could avoid all of these problems. It would likely require a whole bunch of LGOIMA requests of councils that don't make their licensing decisions available online (again, as requires by the Sale and Supply of Alcohol Act), and a lot of effort. Some day.

In the meantime, we should spare a thought for those that want to use data on the number of alcohol licences (in Wellington, or anywhere else). The data that they are being asked to work with is not fit for purpose.

Wednesday, 22 May 2024

Try this: Long-term data series for New Zealand at Data1850

NZIER (New Zealand Institute of Economic Research) has a Public Good Fund that supports awards and scholarships, as well as the Data1850 project. Data1850 collates data for New Zealand going back as far as records allow, which is 1850 for some data.

The Data1850 site allows you to create some cool graphics of the data, like this graph of the unemployment rate from 1956 to 2023 (I've suppressed the separate female and male rates, although it is a little annoying that they still show on the legend):

Jason Shoebridge posted on the Asymmetric Information substack, offering some other examples. I should point out that you can't interpret the graph of life expectancy and household income as showing anything important, because two variables with underlying time trends will always look like they are correlated (this is spurious correlation). However, that post was useful in that it reminded me that this excellent resource exists. [*]

Importantly, you can download the underlying data series (in the Resources tab), which are organised into five datasets: (1) Economic activity; (2) People; (3) Prices; (4) International linkages; and (5) Government. Each dataset has data for several different variables.

Try it out for yourself!

*****

[*] Although it would have been good to remember this before my BUSAN205 students began their group research projects. Having said that, projects using time series data are trickier for students at that level to do well.

Monday, 3 July 2023

Have large language models killed online data collection?

Data is the lifeblood of empirical social science research. Whether it be quantitative or qualitative data, or both, you couldn't do empirical research without it. Self-evidently, the quality of data matters. As the saying goes, garbage-in-garbage-out. You want high quality data to analyse. So, this new working paper by Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West (all Ă‰cole Polytechnique FĂ©dĂ©rale de Lausanne) should be causing some disquiet, especially among those who use Amazon mTurk and similar sources for generating data, because in the paper the authors:

...quantify the usage of LLMs by crowd workers through a case study on MTurk, based on a novel methodology for detecting synthetic text. In particular, we consider part of the text summarization task from Horta Ribeiro et al. (2019), where crowd workers summarized 16 medical research paper abstracts. By combining keystroke detection and synthetic text classification, we estimate that 33-46% of the summaries submitted by crowd workers were produced with the help of LLMs.

Yikes! Between one-third and one-half of mTurk workers are already using large language models (LLMs) like ChatGPT to complete their work. It is easy to see that using mTurk for collecting data from experiments, surveys, etc. has just become untenable. At least, it is untenable if researchers want data collecting from real humans, rather than from LLMs masquerading as humans.

It gets worse though. It isn't just mTurk where this is likely to be a problem. Any online survey is now vulnerable to being completed by a LLM, rendering most online data collection fraught. Journal editors and reviewers will no doubt become aware of this in the future (if they aren't already), so publishing research based on data collected from humans using online methods is going to become a whole lot harder to get published in future.

It's not going to end there. Since LLMs are now generating a non-trivial proportion of online content, a lot of online data is going to lose credibility. And, to top it all off, if future LLMs are being trained on internet-sourced data, they will effectively be being trained on data that is partially generated by today's relatively-low-quality LLMs. There doesn't seem to be much of a way around this.

Anyway, getting back to the Veselovsky et al. article, they aren't as negative in their conclusions as I am above:

All this being said, we do not believe that this will signify the end of crowd work, but it may lead to a radical shift in the value provided by crowd workers.

I guess it depends on what you want the crowd workers to do. As I said above, they won't be contributing much of value to researchers in the future (unless the researchers are researching LLMs). Part of the lifeblood of social science research is bleeding away.

[HT: Marginal Revolution]

Sunday, 23 April 2023

Take care with ILO's broken labour force data series

One of the first rules of working with real-world data is to graph it. That allows us to see where the data has weird inconsistencies, such as those described in this blog post by Kathleen Beegle over on the Development Impact blog:

Has the share of women in the labor force in Rwanda fallen from 84% in 2014 to 52% in 2019? I highly doubt it. More likely, we are seeing the consequence of a major change in the internationally agreed-upon statistical definition of employment. Yes, really. It is quite likely to be a change of which mainly statistical-type-super-data-nerd economists would be aware. But it is one that all of us might want or need to be aware of, especially if you want to properly use country statistics and/or benchmark your own surveys with estimates from national statistical agencies.

In 2013, the 19th International Conference of Labour Statisticians (ICLS) redefined several key labor statistics. It’s taken a few years for these new definitions/concepts to get integrated into questionnaires, into survey efforts, and, as I suspect above, to show up in country statistics. The ICLS19 made several changes but here I focus on one specific change: employment is now work for pay or profit. But wasn’t that was it was before? Not quite. A key feature to this change is that work that is mainly intended for “own-use production” is now excluded, where it was counted as employed before. Before ICLS19, production of primary products, whether for market or household consumption, was counted as employment. The new definition means that someone farming mainly for family consumption (i.e. subsistence farmers) is no longer “employed”. (Though they could still be employed by the new definition if they have another job that qualifies, and in the labor force if they being available and actively searching for work in the form of pay or profit). So subsistence farmers (or those otherwise growing crops mainly intended for home consumption) are “working”, but not “employed”. I put quotes to emphasize these specific terms.

Beegle identifies an interesting phenomenon, and one that we should be careful about - statistical agencies changing the definition of variables that we use. Our analyses would be confounded by these definitional changes if we don't account for them in some way. Now, if all statistical agencies applied the new definition at the same time, that would be easy, but it appears they haven't. Looking at the World Development Indicators, here's the data for five countries: (1) Rwanda; (2) Niger; (3) Papua New Guinea; (4) Benin; and (5) Cameroon (original data are here).

I deliberately chose these five countries as they all seem to experience a similar transition from a high steady-state labour force participation rate to a lower steady-state labour force participation rate. The problem is that all five countries make this transition at different times. So, the usual ways that we would deal with a break in the time series (such as by using a dummy variable for before/after the change in definition, or a before/after dummy variable interacted with a dummy variable for each country) simply isn't going to work well.

This graph also highlights something else, which is no doubt a feature of the underlying ILOSTAT data. Each transition is far too smooth. It is like a straight line has been drawn from the initial steady state series to the start of the new series. No doubt this period of linear transition is actually masking missing data between two consecutive labour force surveys.

We would need to take great care when using the annual data, because there is a great degree of measurement error created by the change in definition, as well as the way the data series have been smoothed between each labour force survey. This is a timely reminder of why the first step in any data analysis should be to graph the data, to help us understand what we are working with.

Wednesday, 18 August 2021

Global inequality and the Subnational Human Development Index

The Human Development Index (HDI) is a widely used summary measure of the level of development of countries. It improves on a simple ranking of GDP per capita or income per capita, because it takes into account health and education. Specifically, it is made up of four indicators: (1) life expectancy at birth; (2) mean years of schooling of adults (aged 25 years and over); (3) expected years of schooling of children aged 6 (which is based on current age-specific enrolment rates at each level of schooling); and (4) gross national income per capita (adjusted for purchasing power parity).

However, one of the problems with the HDI is that it aggregates across each country as a whole. If you want to know anything about the relative levels of development in rural and urban areas of a country, or coastal and landlocked areas, or between different states, the HDI doesn't provide much assistance. However, help is at hand. There is now a Subnational Human Development Index (SHDI) available, and published by the Global Data Lab. The SHDI covers 1625 regions in 161 countries going from 1990 to 2019. Interestingly, for New Zealand it provides index values for all 16 regions, which might be useful for research (because there seems to be a reasonable amount of variation both between regions and over time).

I was alerted to the SHDI's existence by this 2020 article by Inaki Permanyer (Centre d'Estudis DemogrĂ fics) and Jeroen Smits (Radboud University), published in the journal Population and Development Review (ungated version here, and useful summary here). Permanyer and Smits use the SHDI to characterise changes global human development inequality since 2000, which marks a change from considering inequality purely in terms of income (or wealth). Here's the 2018 distribution of the index (Figure 1 from the paper):

Permanyer and Smits note that:

...one can observe clear geographic patterns within countries (e.g., north–south divides in Belgium, Germany, Italy, and Spain). Some countries exhibit large regional variations (e.g., China, India, or Colombia) while others are quite homogeneous (e.g., Australia). Very often, the region where the capital city is located exhibits the highest human development levels and remote rural regions the lowest.

Not all of those are visible in the figure of course, due to the scale. Then, looking at inequality as measured by the Gini coefficient, they find that:

...inequality in the global SHDI distribution has monotonically decreased from 0.14 in 2000 to 0.11 fifteen years later. 

That's consistent with the overall trend observed in income (see here, for example). There are similar trends in the components of the SHDI (health, education, and income). However:

...we observe substantial differences in the magnitudes and speed of the decline. According to the Gini index, differences in the life expectancy index across world regions are smaller than differences in the education index.

Countries are converging much quicker in terms of health than in terms of either education or income. Looking at whether global inequality is mostly within or between countries, they show:

...the very high contribution of within country inequality to total inequality in the groups of countries at low- and intermediate levels of development (where as much as 70 percent of the world population lives). In these groups of countries, about half of inequality in SHDI is within-country inequality.

That is quite a different result from the analysis of Branko Milanovic, who showed that only 10-20 percent of global income inequality was within-country inequality (see this post). Permanyer and Smits explore their results a little further, finding that for the least developed countries:

...within-country SHDI inequality is mostly due to variation in education... In the high developed countries, standard of living surpasses education as the most important explanatory factor for within country SHDI variation...

So, if you only consider variation in per capita income, as Milanovic does, you potentially miss a large contributor to within-country inequality in the least developed countries, which is the variation in education.

All of this helps to paint a more complete picture of global inequality in living standards. Looking forward, if we want greater equality in human development, there clearly needs to be a greater focus on education in developing countries.

Read more:

Friday, 16 July 2021

The ethnographic atlas isn't 'tabulated nonsense'

I've seen a number of papers over the years that have made use of data from George Murdock's Ethnographic Atlas (see here for a gated summary, or here for the data), often as an instrument for some other variable (as one example, see here). The Atlas summarises ethnographic data from over 1200 pre-industrial societies, including a variety of characteristics such as political organisation, social organisation, norms, and agricultural practices. A search on Google Scholar reveals that it has been cited over 6700 times, so it is widely used.

Given widespread use of the Ethnographic Atlas , I was very interested when I saw the title of this new article by Duman Bahrami-Rad, Anke Becker, and Joseph Henrich (all Harvard University), published in the journal Economics Letters (ungated earlier version here): "Tabulated nonsense? Testing the validity of the Ethnographic Atlas". It turns out that I didn't need to worry too much, and nor should researchers using data from the Ethnographic Atlas. Bahrami-Rad et al. compare data from the Ethnographic Atlas with comparable variables from more recent Demographic and Health Surveys for the same ethnic groups. They find:

...positive associations between the historical information reported by ethnographers and the contemporary information reported by a large number of individuals. Importantly, the associations between historical ethnicity-level measures and contemporary self-reported data do not only hold for dimensions that would have been easy to observe for an ethnographer, such as how much a society relies on agriculture, or whether marriages are polygynous. Rather, they also hold for dimensions that are more concealed, such as how long couples abstain after birth, or whether people prefer sons.

Clearly, no cause for concern. The title of the paper is clickbait (the quote "tabulated nonsense" is attributed to the British anthropologist Sir Edmund Leach), and clearly effective, since it got me to read the paper.

Wednesday, 16 December 2020

Is Twitter in Australia becoming less angry over time?

A couple of days ago, I posted about Donald Trump's lack of sleep and his performance as president. One of the findings of the research I posted about was that Trump's speeches and interviews were angrier after a night where he got less sleep. However, if you're like me, you associate Twitter with the angry side of social media, so late-night Twitter activity would seem likely to make anyone angry, sleep deprivation or not.

One other thing that seems to make people angry is the weather, particularly hot weather. So, it seems kind of natural that sooner or later some researchers would look at the links between weather and anger on social media. Which is what this recent article by Heather Stevens, Petra Graham, Paul Beggs (all Macquarie University), and Ivan Hanigan (University of Sydney), published in the journal Environment and Behavior (sorry, I don't see an ungated version, but there is a summary available on The Conversation), does.

Stevens et al. looked at data on emotions coded from Twitter, average daily temperature, and assault rates, for New South Wales for the calendar years 2015 to 2017. They found that:
...assaults and angry tweets had opposing seasonal trends whereby as temperatures increased so too did assaults, while angry tweets decreased. While angry tweet counts were a significant predictor of assaults and improved the assault and temperature model, the association was negative.

In other words, there were more assaults in hot weather, but angry tweets were more prevalent in cold weather. And surprisingly, angry tweets were associated with lower rates of assault. Perhaps assault and angry Twitter use are substitutes? As Stevens et al. note in their discussion:

It is possible that Twitter users are able to vent their frustrations and hence then be less inclined to commit assault.

Of course, this is all correlation and so there may be any number of things going on here. However, the main thing that struck me in the article was this figure, which shows the angry tweet count over time:

The time trend in angry tweets is clearly downward sloping (see the blue line) - angry tweeting is decreasing over time on average. Stevens et al. don't really make a note of this or attempt to explain it. You might worry that this is driving their results, since temperatures are increasing slowly over time due to climate change. However, their key results include controls for time trends. Besides, you can see that there is a seasonal trend to the angry tweeting data around the blue linear trend line.

I wonder - is this a general trend, or is there something special about Australia, where Twitter is becoming more hospitable? The mainstream media seems to suggest that Twitter is getting angrier over time, not less angry. Or, is this simply an artefact of the data, which should lead us to question the overall results of the Stevens et al. paper? You can play with the WeFeel Twitter emotion data yourself here, as well as downloading tables. It clearly looks like anger is decreasing over time, but it may be a result of changing trends in the use of Twitter (or who is using Twitter) over time, especially in the change of language use over time.

I would want to see some additional analysis on other samples, and using other methods of scoring the emotional content of Twitter activity, before I conclude that Twitter is angrier when it is colder, or that Twitter anger is negatively associated with assault. On the plus side, the WeFeel data looks like something that may be worth exploring further in other research settings, if it can be shown to be robust.

Thursday, 23 April 2020

The gender gap in U.S. economics education

The gender gap in economics is a recurrent theme on this blog (see the list of links at the end of this post). This is for good reason - it is pervasive and has a number of negative effects, as noted in those earlier posts. The data outlining the gender gap is becoming more visible, including from this 2019 article by Amanda Bayer (Swarthmore College) and David Wilcox (Federal Reserve Board), published in the Journal of Economic Education (seems to be open access, but just in case there is an earlier ungated version here).

Bayer and Wilcox summarise the differences in the proportions of students studying economics by gender and ethnic group at U.S. universities. If you believe that these proportions should be anywhere near similar, it makes for depressing reading:
Women and students from historically underrepresented race/ethnicity groups graduate with a major in economics at distinctly lower rates than do their counterparts. The pattern is observed both in aggregate and within gender and race/ethnicity categories. For example, among whites, 3.0 percent of men graduate with a major in economics, whereas only 0.8 percent of women do. Among underrepresented minorities, 2.2 percent of men graduate with a major in economics, compared with 0.6 percent of women. Similarly, among both men and women, whites major in economics at higher rates than do [underrepresented minority] students.
Looking individual at each university, they find that:
At every institution in the nation where more than about 3 percent of white men graduate with a major in economics, white women graduate with a major in economics at a lower rate. URM women are similarly underrepresented at almost every institution. The underrepresentation of URM men is less stark than it is for either white women or URM women, but still notable.
Bayer and Wilcox then present a measure of inclusion that they calculate for each institution. However, while their measure is intuitive, I don't believe that it stands up to much scrutiny. They would have been much better off using a proper diversity index like Shannon's evenness index. All of the data that they use is available for you to play with at the New York Fed website, so in principle anyone can go in and calculate a 'better' index based on the data.

Probably the best contribution of this article though, is the recommendations that Bayer and Wilcox make for teachers (to "recognize their sway over the situation"; and to "think intentionally about the implications for diversity and inclusion of the mentorship that they provide"), for textbook authors and publishers (to "commission critical reviews of their own materials, with the goal of identifying how those materials can be made more inclusive along gender, race/ethnicity, and socioeconomic lines"), and for department chairs (to "give careful consideration to maximizing demographic balance among instructors, especially at the introductory level"; to "help recruit and train a diverse set of student teaching assistants"; and to "work actively to improve the culture of their departments, expressed both in formal policies and in the everyday practices of faculty and students"). They also make recommendations for university and college administrators, employers, foundations (e.g. those that fund scholarships), and for the American Economic Association.

Finally, in the conclusion they provide an interesting counterpoint to the argument that differences in choice of major by gender or ethnicity simply represents the optimising behaviour of rational students (including consideration that students act on the basis of comparative advantage):
It is counterproductive to hold an unexamined assumption that the choice of major in college or university is just an example of consumer sovereignty.
It would be nice to see our assumptions about student choices examined in more detail. I don't think we have a very good understanding at all about why many capable students do not choose to follow through on an initial interest in economics.

Read more:

Wednesday, 5 February 2020

A long-run measure of country-level subjective wellbeing

Thanks to the Maddison project, we have long-run measures of GDP that go back to 1820 for many countries, and all the way back to 1 C.E. for some countries. However, it is widely acknowledged that GDP is an imperfect measure of wellbeing - at which point, everyone quotes Robert Kennedy's speech at the University of Kansas in 1968:
The gross national product does not allow for the health of our children, the quality of their education, or the joy of their play. It does not include the beauty of our poetry or the strength of our marriages; the intelligence of our public debate or the integrity of our public officials. It measures neither our wit nor our courage; neither our wisdom nor our learning; neither our compassion nor our devotion to our country; it measures everything, in short, except that which makes life worthwhile.
So, what do we do if we want an alternative measure of wellbeing? Many researchers have begun making use of measures of subjective wellbeing (e.g. life satisfaction, or happiness), notwithstanding recently identified problems with these measures (e.g. see this recent post). But the problem is that these measures have only been collected across a few countries, and only since the 1970s (e.g. see the World Database of Happiness).

A recent article by Thomas Hills (University of Warwick), Eugenio Proto (University of Glasgow), Daniel Sgroi (University of Warwick), and Chanuki Illushka Seresinhe (Alan Turing Institute at the British Library), published in the journal Nature Human Behaviour (ungated earlier version here), attempts to fill this gap. They use data from around 8 million books in the Google Books corpus, published in the U.S., U.K., Germany, and Italy. They analysed the sentiment of words in these books:
We use the words published in these books to compute subjective wellbeing at a given time by using affective word norms to derive sentiment from text. Affective word norms are ratings provided by groups of individuals who examine a list of words and rate them on their valence, indicating how good or bad individual words make them feel.
They then validate their data by showing that it correlates with life satisfaction data from the Eurobarometer survey since the 1970s, and that it seems to pick up key expected trends in life satisfaction over the whole period from 1820. These includes decreases in life satisfaction in all four countries during World War I, for instance.

They then demonstrate some other results using their data, such as the following (based on a regression of life expectancy and GDP growth on their National Valence Index measure):
...one extra year of life expectancy is worth as much as 4.3% annual growth in GDP per capita.
There is a problematic issue that I can see with this data. The meaning of words changes subtly over time, and no doubt the sentiment of words also changes over time. So, measuring the sentiment over nearly two centuries, using word norms from modern times, has the potential to lead to bias. However, it is a measure we didn't have before, and all measures have their limitations. It seems to me that there is a lot of potential for using this measure in some interesting research, and the index can be downloaded from Github here.

[HT: The Economist]

Sunday, 15 December 2019

Alcohol consumption worldwide

Our World in Data recently updated their excellent article on alcohol consumption, with interactive graphics showing cross-sectional differences between countries, and trends over time. The data are pretty clear - New Zealand is neither the least, nor the most, negatively affected by alcohol consumption. Here are a few of the graphs that stuck out for me (though I encourage you to browse through the whole article and have a play with the graphs yourself). First, total alcohol consumption per capita over time:


Notice that New Zealand is towards the bottom of this group of Western countries all the way through the series from 1890 to 2014. Next, consumption by type of beverage for New Zealand, over time:


There's only a few data points in this one, but the massive increase in wine consumption since the 1970s is apparent. Finally, an illustration of why alcohol is such a concern for many public health researchers:


The big four mortality risk factors are high blood pressure, smoking, high blood sugar, and obesity. Then there's a big gap, but alcohol leads the rest of the risk factors in terms of the number of deaths, contributing to over 1200 deaths per year. Once you factor in the other costs of alcohol-related harms, it's easy to see why this is a focus.

Anyway, Our World in Data is an excellent site, and this is just one of many articles on topics as diverse as income inequality and renewable energy. It's a great place to visit to visualise some data, and the best part is that the graphs are (somewhat) customisable and the data are freely accessible as well. Enjoy!

Friday, 8 February 2019

Is language a source of gender bias?

Unconscious bias has been implicated as one of the drivers of gender inequality. Where does this unconscious bias arise from? Given that it is unconscious, it must be part of our identity or our culture - things which are difficult to change (at least in the short term). That might seem like a cop-out, but it does explain why, despite the hype, most interventions to prevent unconscious bias don't work (although, if you follow the link, you'll see that some may do).

One important aspect of culture is language, and languages differ in their treatment of gender. So-called 'gendered languages' attach genders to objects (even inanimate objects). In Spanish, think of el toro (the bull, a masculine noun), or la casa (the home, a feminine noun). English doesn't attach genders to objects like that (although we do sometimes refer to objects, like ships, as a particular gender). So if unconscious bias arises (in part) from culture, and a key aspect of culture is language, then this raises the question: do the words we use, or how we use them, affect gender bias?

This interesting question was addressed in a recent working paper by Pamela Jakiela (University of Maryland) and Owen Ozier (World Bank). They pulled together a dataset on over 4300 languages, spoken by over 99 percent of the world's population. [*] They then used their dataset to explore whether gendered language was associated with women's educational attainment, women's labour force participation, and gender attitudes among men and women. They found:
...a robust negative relationship between grammatical gender and female labor force participation. Our preferred specification suggests that grammatical gender is associated with a 12 percentage point reduction in women's labor force participation and an almost 15 percentage point increase in the gender gap in labor force participation... Taken at face value, our coefficient estimates suggest that gender languages keep approximately 125 million women around the world out of the labor force...
We find a far more muted cross-country relationship between grammatical gender and women's educational attainment. This may be due to the fact that the average within-country gender gap in educational attainment is much smaller than the gender gap in labor force participation | since many wealthy countries have no gender gap in educational attainment, particularly at the primary school level. The prevalence of gender languages is negatively associated with the gender gap in primary school completion after controlling for continent fixed effects, but the estimated relationship is only marginally statistically significant.
Using data from the World Values Survey (WVS), we show that grammatical gender predicts support for traditional gender roles. The coefficient estimate is large in magnitude, suggesting that differences in language could explain the entire gap in gender attitudes between Ukraine (at the 55th percentile of WVS countries in terms of support for gender equality) and Trinidad and Tobago (at the 80th percentile).
So, not only did gendered language explain some of the differences in gender attitudes and female labour force participation between countries, the size of the effects is meaningful. Jakiela and Ozier also showed that the results were similar when looking within countries with language heterogeneity (where some local languages are gendered and others are not), including Kenya, Nigeria, Niger, Uganda, and India. The results are correlations so fall a little short of demonstrating causality, although the paper does include a section that shows that a causal explanation is likely. Jakiela and Ozier conclude that:
Our results are consistent with research in psychology, linguistics, and anthropology suggesting that languages shape patterns of thought in subtle and subconscious ways.
The obvious policy implication to draw from their results, if they are indeed causal, is that gendered languages (e.g. Spanish, German) should immediately drop the gendered treatment of nouns. It couldn't be that simple though, could it?

[HT: Development Impact last June, although they referred to an earlier version of the same working paper]

*****

[*] As an aside, this data seems like it would be a fantastic resource for all sorts of other research, particularly where an instrumental variable for gender bias is required.

Saturday, 21 July 2018

The ancestral characteristics of modern populations

Economic development is remarkably persistent. There is plenty of research that demonstrates that historical patterns of development are predictive of current patterns of development (for example, refer to the research by Daron Acemoglu and James Robinson, as detailed in their book Why Nations Fail (which is on my long list of books-waiting-to-be-read).

Paola Giuliano (UCLA) and Nathan Nunn (Harvard) have a new dataset that, as far as I can see, has enormous potential for looking at a wide range of questions in development, as well as providing a host of candidate variables for use as instruments in otherwise-unrelated analyses. The development of the dataset is described in an article published earlier this year in the journal Economic History of the Developing Regions (ungated version here). The dataset itself is available from Nathan Nunn's website here.

The journal article by Giuliano and Nunn explains:
We contribute to this line of research by providing a publicly accessible database that measures the economic, cultural, political, and environmental characteristics of the ancestors of current population groups... Specifically, we construct measures of the average pre-industrial characteristics of the ancestors of the populations in each country of the world. The database is constructed by combining preindustrial ethnographic information for approximately 1,300 ethnic groups with information on the current distribution of approximately 7,500 language groups measured at the grid-cell level.
Giuliano and Nunn then go on to describe the dataset, as well as providing illustrations of the data. What particularly caught my eye was a brief analysis they did of the relationship between their historical geographic characteristics (meaning the average ancestral characteristics of populations living in current countries) and current GDP. They find that:
Not surprisingly, being further from the equator is positively associated with real per capita GDP. However, what is more surprising is that the ancestral measure appears to be much more strongly correlated than the contemporary measure. This is particularly striking since we would expect the ancestral measure to be more imprecisely measured than the contemporary measure.
They find similar results for ancestral ruggedness of the land, and ancestral distance from the coastline. The reason these results caught my eye was that it suggests to me that these variables might be suitable instruments for GDP in other analyses (such as when GDP would be endogenous in the particular model you are trying to run. If that was a bit too pointy-headed for you, don't worry. It just suggests that these variables have a lot of potentially cool uses for economists.

[HT: Marginal Revolution]

Thursday, 26 April 2018

Facebook as a measure of social connectedness

Economists are often maligned for not recognising the importance of social relations in research. That is somewhat unfair, since the importance of social connections or networks is well recognised in the research on migration and trade, not to mention the growing literature on the importance of social capital. However, the biggest problem with including social connections in economics research is that they are notoriously difficult to measure. So, I was quite excited to read this 2017 NBER Working Paper (ungated version here) by Michael Bailey (Facebook), Ruiqing Cao (Harvard), Theresa Kuchler, Johannes Stroebel (both New York University), and Arlene Wong (Princeton). In the paper, the authors demonstrate a new Social Connectedness Index (SCI), derived from Facebook friends data:
Specifically, the SCI corresponds to the relative frequency of Facebook friendship links between every county-pair in the U.S., and between every U.S. county and every foreign country.
The paper then goes on to demonstrate the usefulness of the SCI:
We use these data to document important geographic patterns of social networks. We also show that the SCI data can be informative about the role of social connectedness for the large number of social and economic outcomes that can be measured at various levels of geographic aggregation, such as trade, migration, and patent citations. To facilitate further research along these dimensions, the SCI data can be made accessible to members of the broader research community...
We find that the intensity of friendship links is strongly declining in geographic distance, with the elasticity of the number of friendship links to geographic distance ranging from about -2.0 over distances less than 200 miles, to about -1.2 for distances larger than 200 miles. Conditional on distance, social connectedness is significantly stronger within states than across state lines. We also show that, conditional on geographic distance, the social connectedness between two counties is increasing in the similarity of these counties along important social and economic characteristics...
After aggregating the SCI to the state level to match available interstate trade data, we document that state-pairs with higher social connectedness see larger trade flows, even after controlling flexibly for geographic distance...
We also find that when counties are more connected, they are likely to have more cross-county patent citations...
Finally, we find that more connected county-pairs see more migration and labor flows, highlighting the potential of social networks to overcome frictions involved in moving across the United States...
Overall, the findings presented in this paper suggest that social connectedness plays a large role in explaining social and economic interactions, both within and across counties.
It seems to me that there is a huge amount of potential in using the SCI data. Better still, the dataset is available to researchers, as Bailey et al. note in a footnote:
Researchers are invited to submit a one-page research proposal for working with the SCI data to sci_data@fb.com. The data will be shared for approved research projects under the terms of an NDA between Facebook and approved researchers.
The Bailey et al. analysis suffers from being correlation rather than causal, but the depth and coverage of the SCI data means that there are a lot of research questions that it could be useful for, especially in studies of migration (where social networks matter in terms of migrants' or potential migrants' decisions about where to move), immigrant assimilation (where local social networks facilitate immigrants' adaptation to their new location), entrepreneurship (where social networks may impact on business success), idea or norms diffusion (where social networks are important mechanisms for promotion), and for any application where the measurement of social capital is important.

The biggest problem may be: is this dataset still available, given the current climate surrounding Facebook and data? There is no individual data in the SCI dataset (it is made up of county-level and country-level data only), so one would hope so.

[HT: Marginal Revolution, in July last year]

Tuesday, 30 January 2018

Boy racers can't do statistics

Matthew Hansen wrote in the New Zealand Herald today:
Let's get this out of the way early - "boy racer" is a ridiculous and outdated term. Much of the country's modified car culture is propped up by the middle-aged, and by women. We're a world leader for female involvement in motorsport.
The "boy" aspect isn't exactly prevalent in the New Zealand Transport Agency's numbers for road deaths either. In the past 12 months, 379 drivers have been killed, and the three biggest age groups represented are those from 25-39 (103), 60-plus (91), and 40-59 (91). By contrast, deaths for those aged 15-19 number 25, and 52 for 20- to 24-year-olds.
I'm no rocket scientist, but those numbers are smaller. So why empower an outdated, incorrect term like "boy racer"?
Yes, those numbers are smaller, but that's often what happens when you compare the numbers of events happening to people in a five-year age group (15-19 or 20-24 years) with the comparable numbers for a 15-year age group (25-39), a 20-year age group (40-59), or a 40+-year age group (60 years and over). Before comparing the number of road deaths between age groups, you need to adjust for the relative number of people in each age group, to work out the incidence of road deaths. [*]

The number of road deaths this year so far are available from the NZTA road toll website. There are some small differences with the numbers that Hansen uses (probably because the statistics reported there are for the 12 months that end on the day you access the website), so for comparability I will use Hansen's numbers. The numbers of people in each age group are available from Statistics New Zealand (NZ.Stat) for 30 June 2017 (which is close enough to the mid-point for the year ended on some day in January).

While there have been 103 road deaths among people aged 25-39, there are 962,550 people in that age group. That works out to 1.07 road deaths per 10,000 people. Compare that with 52 road deaths and 355,830 people aged in the 20-24 age group, which works out at 1.46 road deaths per 10,000 people. For completeness, the other values are in the table below.


Clearly, the highest risk group is the 20-24 year age group when it comes to road deaths. I'm no rocket scientist either, but the numbers for other groups are smaller. In some cases close to half of the incidence for the 20-24 year age group. Maybe boy racers should stick to cars, not statistics? And the Herald should send its reporters to a course on some basic statistical literacy.

*****

[*] Even better would be to adjust for the number of vehicle road miles travelled by members of that age group, which would be a 'more correct' measure of risk exposure (for vehicle travellers, so probably pedestrian and cyclist road deaths should be excluded first too).

Tuesday, 3 October 2017

Trade and the Atlas of Economic Complexity

Last week in ECON100 we covered the gains from trade. The simple model we employ is essentially a model based on Ricardian trade, which assumes that each country specialises in producing (and exporting) goods that they have comparative advantage in producing (goods that they can produce at a lower opportunity cost), and imports goods that they have comparative disadvantage in producing (goods that they produce at a high opportunity cost). However, the real world is significantly more nuanced than this simple model, as Noah Smith noted recently:
Most academic models of international trade are pretty simplistic. Some of these models are surprisingly effective for making certain types of predictions -- for example, economists are very good at predicting how much different countries will trade with each other. But they’re not so good at predicting what kind of things the countries will specialize in, which country will have a trade deficit or surplus, how trade will affect growth, or which workers and businesses will benefit from trade...
Now, a number of economists are working on new empirical approaches that take into account the huge variety and complicated connections between the products and services that get traded across international borders.
Two such economists are Ricardo Hausmann of Harvard’s Kennedy School and Cesar Hidalgo of Massachusetts Institute of Technology. They and their research team have a theory that the more different products a country makes, the better positioned it is to grow. This idea runs counter to the conventional wisdom -- and the predictions of many standard models -- that different countries hyperspecialize in only a few goods and services. According to Hausmann and Hidalgo, countries are better off when they can make a multitude of things. Countries such as Saudi Arabia that rely on a single product will perform worse, all else equal, than countries such as Japan that can make almost anything they want.
The economists claim that their so-called economic complexity index is much better at predicting long-term economic growth than other forecasting methods based on things like the level of regulation or the amount of investment in education. They recently put out a report predicting that China’s growth will slow over the next decade, while India’s will remain rapid.
Hausmann and Hidalgo's Atlas of Economic Complexity is well worth looking at. There is a wealth of trade data, and excellent visualisations (if you click on 'Visualizations' in the top bar). For instance, here's New Zealand's exports by category for 2015 (it's much easier to see at the website):


I was surprised that raw aluminium was as much as 1.7% of exports. And here's a similar visualisation of export destinations (again for New Zealand in 2015; here's the direct link):


No surprises about China, Australia, the United States and Japan being the biggest export destinations, but Algeria (1.2%) and Egypt (1.0%) were a bit surprising to me. Anyway, there's lots more to explore on the site, and lots of surprises (try playing the 'which country is the biggest exporter of *some random product*?' game with your family or friends). Minutes of fun, guaranteed. Enjoy!