Friday, 18 April 2025

This week in research #71

Here's what caught my eye in research over the past week:

  • Bertola and Lo Prete (open access) find large propensities to guess in Italian data from a survey on financial literacy and resilience, and show that truly financially literate respondents are more likely than those who guessed and randomly picked the correct answers in the financial literacy test to make ends meet at the end of the month and to cope with unexpected expenses
  • Feyzollahi and Rafizadeh investigate the likely use of large language models in the writing of papers published in 25 top economics journals, and find a 4.76 percentage point increase in LLM-associated terms during 2023–2024, and that the effect more than doubles from 2.85 percentage points in 2023 to 6.67 percentage points in 2024, suggesting rapid integration of language models in economics research (are you also wondering if their paper was partly written by an LLM?)
  • Hatton (open access) presents new data on the voyage times and travel costs for emigrants from the UK traveling to the US and to Australia from 1850 to 1913, showing that the voyage time from Liverpool to New York fell from 38 days to just 8 days (or 79 per cent), and the voyage time to Sydney fell from 105 days to 46 days (or 56 per cent)
  • Brakman, Kohl, and van Marrewijk (open access) use a gravity model to link long-run changes of the demographic dividend to geographical changes in world trade for the 21st century, and show that, compared to the current situation, North America and Europe will no longer be the centre of global trade in 2100 due to their aging populations, while South Asia and Sub-Saharan Africa will experience a substantial increase in their share of world trade, and China will experience a substantial decrease
  • De Haro investigates whether declining drug revenues in Mexico incentivised cartels to target the avocado sector, and finds that the decline in the demand for heroin increased homicide rates, including those of agricultural workers, as well as truckload thefts in avocado-growing municipalities
  • Chenarides et al. find that dollar store presence in a county corresponds with higher employment levels within the general merchandise retail sector, and decreases in average weekly earnings (although these reverse in urban areas over time)

Thursday, 17 April 2025

What's new in regional and urban economics?

The journal that I edit, the Australasian Journal of Regional Studies, is about to release its latest issue (more on that in a future post). That makes it timely to think about what's new in regional and urban economics. Actually, it's probably always a good time to think about what's new, but this is a particularly useful time because we can rely on this new NBER Working Paper (ungated here) by Ran Abramitzky (Stanford University), Leah Boustan (Princeton University), and Adam Storeygard (Tufts University).

The paper mainly covers new data sources that have come into more regular use in recent years, and provides a good survey of the literature that has developed using each source. Abramitzky et al. also identify some new use cases for some of the data sources, which points to new research directions or extensions of existing work. For new (or experienced) researchers looking for inspiration for their next research project in regional and urban economics, this paper is a good one to read.

To save you a little bit of time though, here are some of the key data sources that Abramitzky et al. discuss (some of which I have grouped together differently than they do). The first is historical (US) Census records:

The US Census is far from a “new” data source, having provided the backbone of empirical research in urban economics and other applied fields for decades. Yet advances in record linkage have allowed researchers to convert (historical) census data into large panel datasets that follow individuals over time. This longitudinal data opens up a set of new research questions on spatial topics, including the determinants of geographic mobility, the long-run effect of childhood exposure to environmental conditions or economic shocks, and the causes and consequences of neighborhood change within cities.

Complete census records, including an individual’s name and detailed location information, becomes available to the public 72 years after the Census is taken; the 1950 Census was just released in 2022.

Sadly, this is not a data source that is available for many countries (including New Zealand, where Census unit records prior to 1966 were destroyed). However, the ability to link people over long periods of time (including between generations) has opened up a wealth of new research questions. Second, Abramitzky et al. discuss digitised historical maps and directories:

Digital spatial data in Geographic Information Systems (GIS) is indispensable for a variety of modern urban applications but, until recently, historical maps were not compatible with this tool. In recent years, economic historians and other social scientists have digitized a wide range of historical maps, including census geography, and environmental and land management maps. These efforts have opened up study of historical neighborhoods and the effects of proximity to relevant geographic features like administrative boundaries, industrial sites, religious and cultural institutions and the epicenters of natural disasters.

I have a project in progress (which is, unfortunately, somewhat stalled due to non-map-related data issues) that has made use of digitised boundary maps for the electoral boundaries from past New Zealand elections (more on that in a future post, if that project ever gets re-started). However, the key point is that there is a wealth of information stored in historical maps and archives that are currently underutilised. On a related note, Abramitzky et al. note that:

Beyond mapping the location of households or firms, GIS is also useful for reconstructing historical transportation infrastructure via waterways, roads or railroads.

Given that past infrastructure patterns, including transportation and other networks, affect the patterns observed today, these seem like important sources. Third, Abramitzky et al. talk about a range of remote sensing data, including night lights (from satellite imagery), and physical attributes like air quality, weather variables, and building footprints and heights. These data are often available at small spatial resolutions, allowing analysis at very fine-grained spatial levels. However, it is worth reading the paper (and the references) carefully, as they also identify issues to be aware of with remote sensing data.

Fourth, Abramitzky et al. very briefly discussed picture and video data, including Google Streetview, and CCTV camera data. There are definitely some interesting use cases for these atypical data sources, and you can expect to see a growing use of them in future research. Fifth, Abramitzky et al. discuss mobile phone or smartphone data, including data derived from particular smartphone apps:

Mobile phones provide information about the location of the people who use them, and sometimes the vehicles they drive. Broadly, there are two kinds of cell phone data. Call data records (CDRs), provided by network operators, report the location of the phone at the time a call was made or received, as triangulated from the network of cell towers. In some cases, the counterparty to the call can also be identified...

For research purposes, CDRs have been mostly replaced by data from smartphones, whose apps collect more accurate GPS-based locations at all times (regardless of whether a call is placed)...

Researchers have used location data from individual apps with which they have developed relationships. Most prominently, Uber has provided data on its trips to several groups of researchers...

Similarly, they discuss transportation data derived from vehicle location trackers, transit cards, or electronic tolling stations, or electronic payment systems for transit riders. All of these sources are useful for identifying transportation and commuter flows, which have high policy relevance.

Finally, Abramitzky et al. give a rapid-fire selection of other data sources that are only beginning to be used, including e-commerce and payment card transactions data, posted prices and listings (often scraped from websites), routinely collected administrative data (which in my experience will generally require a lot more data cleaning), and text as data.

Clearly there are lots of new and emerging data sources in use in regional and urban economics. However, Abramitzky et al. are clear that developing skills with these data sources and the appropriate methods for dealing with them is not feasible for everyone. They do, however, suggest a solution:

We encourage urban economists, both young and old, to familiarize themselves with these data sources and to become conversant in some of the methods needed to build new data from textual corpora, digital traces, and images and video of the world around us, including large language models and deep learning more broadly. We emphasize the word “conversant” because we do not think that all of us need to become experts in these techniques. Rather, we anticipate and encourage interdisciplinary collaboration with scholars around the university in data science, computational linguistics, computer science, geography and the natural sciences who know these methods well and can thus complement the research focus and conceptual framework specific to urban economics.

So, we should definitely expect the current trend for larger, more interdisciplinary, research teams to continue into the future.

[HT: Marginal Revolution]

Tuesday, 15 April 2025

Mexico's lake of tequila

What happens when prices fail to adjust to a new equilibrium? Mexico found out the hard way at the end of last year, as the Financial Times reported in December (paywalled):

Mexico is sitting on more than half a billion litres of tequila in inventory, almost as much as its annual production, as the fast-growing industry reckons with slowing demand and the prospect of tariffs on exports to the US under Donald Trump.

By the end of 2023, the industry had 525mn litres of tequila in inventory, either ageing in barrels or waiting to be bottled, according to data shared with the Financial Times by the Tequila Regulatory Council. Of the 599mn litres of tequila produced last year, about one-sixth remained in inventory, according to the figures.

“Much more new spirit is being distilled than is being sold, and inventories are starting to accumulate,” said Bernstein analyst Trevor Stirling, attributing the build-up to falling demand and new distillery capacity that has recently begun operating in Mexico. “The tequila industry is set for a very turbulent 2025.”

Consider the market for tequila, as shown in the diagram below. The market was originally in equilibrium, where the supply curve S0 meets the demand curve D0, with an equilibrium price of P0 and Q0 units of tequila being traded. Then demand decreased to D1, and supply decreased to S1. The market should move to the new equilibrium, where the supply curve S1 meets the demand curve D1. However, say that the price remained at the original price P0 for a little while. What would happen?

If the price remained P0, the quantity of tequila supplied would increase to QS (because with the supply curve S1, the quantity supplied at the price P0 is equal to QS). The quantity of tequila demanded would decrease to QD (because with the supply curve D1, the quantity demanded at the price P0 is equal to QD). The difference between QS and QD is the quantity of tequila that remains unsold - a surplus, or excess supply. Or, as the FT article refers to it, a 'tequila lake'.

What happens next? In a market with excess supply, we would expect the price to adjust. Since tequila distilleries can't sell all of their tequila inventory, and it is costly to store it, they would start to lower the price. This 'bidding down of the price' by sellers would continue until the excess supply is eliminated. On the diagram above, that happens when the market gets to the new equilibrium, at the lower price P1, where Q1 tequila is traded.

And the FT article even notes some evidence that this adjustment is happening:

Two of the largest tequila brands, Bacardi-owned PatrĂ³n and Casamigos, which is now owned by London-listed Diageo, have been cutting prices for more than a year in response to weaker consumer demand, according to research by Bernstein.

Eventually, there would be no more tequila lake. Which is sad, because it calls to mind some interesting imagery. Here's what ChatGPT thinks a tequila lake looks like:

That's one way to get rid of the tequila lake, I guess!

Sunday, 13 April 2025

Anthropic on how university students are using generative AI

This week Anthropic released a fascinating and important report on university students' use of generative AI (specifically, how they are using Anthropic's Claude AI). The report is based on anonymised data from over 570,000 conversations (by people with a university-affiliated email address, who the AI judged to be students, not staff) over an 18-day period. 

The report has a number of important insights, but I want to focus on two in particular. First, it tells us how students are interacting with Claude. Anthropic summarises this with the following taxonomy:

Anthropic then note that:

These four interaction styles were represented at similar rates (each between 23% and 29% of conversations), showing the range of uses students have for AI.

Now, most people are probably most interested in how students are interacting with Claude, in order to determine the extent of cheating in assessment that is going on. In terms of the four categories, I'd suggest that there is a clear hierarchy. Direct output creation is the most likely to be cheating, since asking AI to write an essay or project report is likely to fit in there. Next is direct problem solving, since asking AI to provide answers to take-home tests and multiple-choice quizzes is likely to be in that category. However, students asking direct questions, using generative AI in place of a search engine, would also be captured, and that likely isn't cheating and is likely to contribute to student learning (indeed, that is how we encourage students to use Harriet, our ECONS101 AI tutor). Third is collaborative output creation, since that may involve rewriting an essay to thwart plagiarism tools, or using AI to provide critiques of other output, or debugging code. However, there is no doubt a lot of genuine collaborative effort that is allowed as part of assessment guidelines that will fit into that category. Finally, collaborative problem solving is likely to be the least problematic. A 'Socratic tutor' or guided learning approach would fit in here, for example.

In relation to cheating, Anthropic notes that:

...nearly half (~47%) of student-AI conversations were Direct—that is, seeking answers or content with minimal engagement. Whereas many of these serve legitimate learning purposes (like asking conceptual questions or generating study guides), we did find concerning Direct conversation examples including:

  • Provide answers to machine learning multiple-choice questions
  • Provide direct answers to English language test questions
  • Rewrite marketing and business texts to avoid plagiarism detection

These raise important questions about academic integrity, the development of critical thinking skills, and how to best assess student learning. Even Collaborative conversations can have questionable learning outcomes. For example, “solve probability and statistics homework problems with explanations,” might involve multiple conversational turns between AI and student, but still offloads significant thinking to the AI.

They are absolutely right that how to best assess student learning is an important question. And as Justin Wolfers notes, any high-stakes at-home assessment is essentially a non-starter in terms of credibility. That certainly limits the options available to lecturers and teachers.

The second important insight from the report is the cognitive level at which students are engaging with Claude. This is the really worrying aspect (and the modes of interaction are worrying enough already), because:

We saw an inverted pattern of Bloom's Taxonomy domains exhibited by the AI:

  • Claude was primarily completing higher-order cognitive functions, with Creating (39.8%) and Analyzing (30.2%) being the most common operations from Bloom’s Taxonomy.
  • Lower-order cognitive tasks were less prevalent: Applying (10.9%), Understanding (10.0%), and Remembering (1.8%).

In other words, students were outsourcing tasks that were higher on Bloom's taxonomy. As I noted in this post last year:

Teachers might hope that generative AI is better at the lower levels - things like definitions, classification, understanding and application of simple theories, models, and techniques. And indeed, it is. Teachers might also hope that generative AI is less good at the higher levels - things like synthesising papers, evaluating arguments, and presenting its own arguments. Unfortunately, it also appears that generative AI is also good at those skills. However, context does matter. In my experience, and this is subject to change because generative AI models are improving rapidly, generative AI can mimic the ability of even good students at tasks at low levels of Bloom's taxonomy, which means that tasks at that end lack any robustness to generative AI. However, at tasks higher on Bloom's taxonomy, generative AI can mimic the ability of failing and not-so-good students, but is still outperformed by good students. So, many assessments like essays or assignments that require higher-level skills may still be a robust way of identifying the top students, but will be much less useful for distinguishing between students who are failing and students who are not-so-good.

It seems that either I was wrong in my assessment of the strengths of generative AI at different levels of Bloom's taxonomy, or that despite the weaknesses of generative AI at the higher levels, students still prefer to use it that way. That might reflect comparative advantage. Perhaps generative AI is better at both lower-level and higher-level tasks (it has absolute advantage in both), but it has comparative advantage in the higher-level tasks? In that case, students may find it more useful to have the generative AI work on the higher-level tasks, while they complete the lower-level tasks themselves (I feel like that would be a useful topic to explore in a future post). Anyway, Anthropic point to the worries when they note that:

...AI systems may provide a crutch for students, stifling the development of foundational skills needed to support higher-order thinking.

This was definitely an interesting report and provides some genuine insights into university student use of generative AI. I feel like there is much more to learn here, including how to steer students towards more collaborative modes of engagement with the AI, which are likely to lead to more learning. It also highlights (yet again, as if we need further reminders) the vulnerability of much assessment to students' use of AI.

[HT: Marginal Revolution]

Read more: