Monday, 3 November 2025

Generative AI and entry-level employment

Does technological change increase employment, or decrease employment? The answer depends on what you believe about the technological change. If the technological change primarily automates tasks that were previously done by human workers, then it may decrease the demand for labour and reduce employment [*]. On the other hand, if the technological change primarily makes workers more productive, so that they generate more value for their employers, then it may increase the demand for labour and increase employment [**].

Which of those two situations is AI creating? In reality, it's probably a bit of both. However, which effect is stronger? We can get a sense of where things are at from this recent working paper by Erik Brynjolfsson, Bharat Chandar, and Ruyu Chen (all Stanford University). They use data from ADP, the largest payroll processing firm in America, which contains records on between 3.5 and 5 million workers each month between January 2021 and July 2025. Using this data, Brynjolfsson et al. demonstrate six key facts about the labour market over that time:

  • First, we find substantial declines in employment for early-career workers in occupations most exposed to AI, such as software development and customer support.
  • Second, we show that economy-wide employment continues to grow, but employment growth for young workers has been stagnant.
  • Third, entry-level employment has declined in applications of AI that automate work, with muted effects for those that augment it.
  • Fourth, these employment declines remain after conditioning on firm-time effects, with a 13% relative employment decline for young workers in the most exposed occupations.
  • Fifth, these labor market adjustments are more visible in employment than in compensation.
  • Sixth, we find that these patterns hold in occupations unaffected by remote work and across various alternative sample constructions.

The first key fact is demonstrated by looking across all occupations. However, it is most clearly seen in Figure 1 from the paper, which shows how the number of workers (by age group) employed as software developers or customer service representatives (two of the occupations most cited in the media as being affected by AI) have changed compared with October 2022:

Notice how the blue line (employment of early career workers aged 22 to 25 years) trends downwards, starting from 2022, while employment of the most senior workers (especially those aged 35 years and over) continue to trend upwards.

The second key fact builds on this, to show the trend is apparent when you pool all workers together, as demonstrated in Figure 4 from the paper:

Note that the trends are not as wildly different as they are for software developers or customer service representatives, but remember that figure shows the trends across all workers, including those in jobs like nurses, welders, or baristas, whose employment is unlikely to be very impacted by AI (yet!).

Their third key fact relates most closely to the point I raised at the beginning of this post. On this, Brynjolfsson et al. find that:

...occupations with the highest estimated automation shares have experienced declining employment for the youngest workers...

...occupations with the highest estimated augmentation shares have not experienced a similar pattern...

The relevant figures are Figures 7 and 8 from the paper (which are too large for me to reproduce sensibly here). For their fourth key fact, Brynjolfsson et al. use a Poisson regression model within each age group, which allows them to control for firm-time-specific and firm-quintile-specific effects (where the quintiles are quintiles of exposure to AI, drawn from the paper that I discussed in this 2023 post). The firm-time effects will control for firm-specific shocks that affect all of the firm's workers, while the firm-quintile effects will control for different trends affecting a firm's workers that have similar exposure to AI. The analysis won't control for all of the relevant differences, and the analysis remains correlational rather than causal. Nevertheless, Brynjolfsson et al. find that:

For workers aged 22-25, estimates for higher quintiles are large and statistically significant, with a 12 log point decline in relative employment... Estimates for other age groups are generally much smaller in magnitude and not statistically significant.

A 12-log point change is about 11.3 percent, so relative to older workers working in the same firm and with the same exposure to AI, the youngest workers have suffered an 11.3 percent decrease in employment.

Brynjolfsson et al.'s fifth point is simply that the effects show up for employment, but not for wages. And their sixth point is that the effect is robust to various alternatives, including: excluding tech occupations; looking separately at jobs that are, or are not, amenable to remote work; extending the pre-period back to 2018; looking separately by gender; and using the Current Population Survey instead of the payroll data. Interestingly, looking differently at occupations depending on the education level of workers, the results show that:

Occupations with a high share of college graduates have declining employment overall, with muted differences between more-exposed and less-exposed occupations compared to our main results. In contrast, occupations with a low share of college graduates have rising overall employment, with the least AI-exposed occupations growing and the most exposed occupations declining in employment.

All of this is not great news for current university students. Think about all of the results taken together. Firms have been employing fewer entry-level workers, which are your typical graduates. Occupations that have a high share of college graduates are experiencing declining employment overall, in both occupations that are more exposed to AI and occupations that are less exposed to AI. And employment in AI-exposed occupations with a low share of college graduates has also been declining. It seems that the only groups of entry-level workers who aren't experiencing negative trends are non-college-educated workers in occupations that are not exposed to AI.

Brynjolfsson et al. try to finish on a positive note though, noting that:

The adoption of new technologies typically leads to heterogeneous effects across workers, resulting in an adjustment period as workers reallocate from displaced forms of work to new forms with growing labor demand... Past transitions such as the IT revolution ultimately led to robust growth in employment and real wages following physical and human capital adjustments, with some workers benefiting more than others...

Here's hoping that we don't have to wait too long for those adjustments.

[HT: Marginal Revolution

*****

[*] However, decreasing employment (and wages) in the automating industry may increase the supply of workers into other industries, increasing employment (but decreasing wages) in those other industries. The general equilibrium effect is not straightforward.

[**] Arguments that increasing productivity will reduce the demand for labour, since fewer workers are needed to complete the same amount of work, run into the 'lump of labour' fallacy. They forget that firms can expand production if they have more productive workers, and will want more workers if their workers are more productive and therefore more profitable to employ.

Read more:

  • ChatGPT and the labour market
  • More on ChatGPT and the labour market
  • The impact of generative AI on contact centre work
  • Some good news for human accountants in the face of generative AI
  • Good news, bad news, and students' views about the impact of ChatGPT on their labour market outcomes
  • Swiss workers are worried about the risk of automation
  • Saturday, 1 November 2025

    Accounting for free goods in GDP, and the economic value of generative AI

    In a recent Wall Street Journal article (ungated version here), Avinash Collis (Carnegie Mellon University) and Erik Brynjolfsson (Stanford University) report that:

    In late 2024, a nationally representative survey of U.S. adults revealed that 40% were regular users of generative AI. Our own survey found that their average valuation to forgo these tools for one month is $98. Multiply that by 82 million users and 12 months, and the $97 billion surplus surfaces.

    This estimate of $97 billion per year provides a measure of the economic impact of AI. That is about 0.33 percent of US GDP (which was about $29.3 trillion in 2024). That is somewhat less that I would have expected. However, to put that $97 billion in context, it is nearly the same size as the 'value added' (the industry-level equivalent of GDP) of the entire motion picture and sound recording industry (which was $119 billion in 2024).

    Comparisons such as that may not be sensible though, because we aren't quite comparing apples with apples. GDP is really a measure of production, whereas the value generated by ChatGPT estimated by Collis and Brynjolfsson is closer to a measure of consumer surplus. ChatGPT does contribute to GDP already of course, because each user who has a paid account pays for access. However, what they pay is far less than the value they receive from access to ChatGPT. And some consumers are paying nothing at all.

    This problem, and a potential solution, are outlined in this new article by Brynjolfsson, Collis, and co-authors, published in the American Economic Journal: Macroeconomics (ungated earlier version here). They explain the problem as follows:

    New, sometimes very specialized, goods appear with increasing rapidity... and digital goods (such as information and entertainment services) are increasingly available at zero price, reflecting their very low marginal costs of replication and distribution... the positive quantities of these goods that are consumed have a measured price of zero and measured value of zero in the conventional national accounts even if they create considerable consumption value for consumers. A related difficulty arises in the valuation of new goods, since there is no observed price for the period before their appearance. Despite the increasing relevance of new and free goods, the value to consumers is not reflected in standard statistical agency reports for GDP or derivative metrics like productivity, which are typically defined in terms of GDP.

    Brynjolfsson et al. provide a framework for estimating the welfare contribution from new zero-priced goods, as well as the quality improvement of existing goods. The framework itself is quite mathematical and not for the faint-hearted. However, the basic premise is that the welfare that is generated by a good that is offered for free can be estimated by estimating how much consumers would be willing to accept to give up access to that good for a period of time (notice that this is what Collis and Brynjolfsson talk about in their WSJ article in relation to generative AI).

    There are various ways that can be used to estimate what consumers would be willing to accept to give up a free service. However, the challenge is that the estimates need to be credible. Brynjolfsson et al. provide two examples of incentive-compatible experiments. In these experiments, the research participants might really have to give up the service they were being asked to value, which provides a strong incentive for them to provide their 'real' valuation of what they would accept. The first experiment was conducted online and used to value Facebook (this is research that I blogged about back in 2018). In this case:

    In the experiment, each participant was asked to make a single discrete choice between two options: (i) keep access to Facebook or (ii) give up Facebook for one month and get paid $E. We allocated participants randomly to 1 of 12 price points: E ∈ {1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1,000} —that is, we observe at least 200 participants per price point. Before participants made the decision, we informed them that their decisions were consequential such that we would randomly pick 1 out of every 200 participants and fulfill that person’s selection. We also informed them about how we can monitor their Facebook online status remotely. In order to check if the selected participants gave up Facebook and qualified for the payment, we monitored their online status on Facebook for 30 days...

    Research participants were only paid if they didn't log onto Facebook for the 30 days. In their second example, Brynjolfsson et al. conducted a lab experiment with students in the Netherlands. In this case:

    We asked participants to state the minimum amount of money they would request in order to give up their smartphone camera (both main camera and front camera) for one month. Participants were informed that this amount would serve as a bid in a lottery. If their minimum bid to forgo their camera would be higher than a random price, drawn from a uniform distribution, they could keep access to their smartphone camera but would not receive any cash...

    In order to induce incentive compatibility and make the answers consequential, we provided further information that 1 out of 50 participants would be selected for the lottery and that we would block their smartphone cameras with a special sealing tape if their bid was successful... The sealing tape would break if the participants tried to peel it off so that it was not possible to reapply it. We also signed the tape so that it was not possible to buy the same type of seal and reapply a seal. If, after the one-month period, the original seal was still intact participants were rewarded with the money and the seal could be removed.

    Notice that in both cases, the research participants might really have to give up what they are being asked to value. Collis and Brynjolfsson no doubt did something similar to derive their estimate of the value of generative AI.

    Anyway, coming back to the Brynjolfsson et al. paper, their purpose is to create a new measure, which they term GDP-B, that includes the benefits of goods and services that have high marginal value to consumers, but are offered at low or no cost to them. This is not limited to the particular examples that Brynjolfsson et al. cite (which include not just Facebook and smartphone cameras, but also Instagram, Snapchat, Skype, WhatsApp, digital maps, LinkedIn, and Twitter). Even that list is very incomplete of course. 

    And, given that every good or service generates consumer surplus, the approach could be extended to every good or service. No doubt that would be the ideal, so that GDP-B captures the consumption benefits of all goods and services. However, the data collection task required to estimate consumer value for every good and service would be enormous, dwarfing the efforts already taken to measure the Consumer Price Index (which doesn't survey the price of every good and service).

    Adopting an approach that only accounts for some goods and services seems to me to be suboptimal. And that is what makes the comparisons a bit problematic. Brynjolfsson et al. compare traditionally-measured GDP growth with growth based on their measure GDP-B, which includes the benefits from Facebook (or smartphone cameras). However, all they can really show there is that growth is higher when the benefits from Facebook (or smartphone cameras) are included. I would take the question of how much higher growth is with a grain of salt - they aren't accounting for the benefits of all of the other goods and services that they haven't valued in the same way (for some of which, the benefits might have declined, because consumers value them less, leading to lower growth).

    None of this is to say that we shouldn't be finding better ways to measure welfare than GDP, which has long been acknowledged as a very incomplete measure. Brynjolfsson et al. have provided a measure that goes beyond GDP, which is what many have been calling for, for some time.

    [HT: Marginal Revolution, here for the WSJ article, and here for the AEJ:Macro paper]

    Read more:

    Friday, 31 October 2025

    This week in research #99

    Here's what caught my eye in research over the past week (yet another very quiet week, it seems):

    • Miller et al. (with ungated earlier version here) use a conjoint survey experiment to examine the hiring preferences for lobbyists, finding that organised interests prefer lobbyists with policy-specific expertise and the necessary connections to get access to decision-makers, but they find little evidence that connections are more valuable than expertise
    • Li et al. find that an in-sample shift in de-seasoned weather from the coolest to the hottest semester reduces semester-long undergraduate student performance by 1.5 percent in Singapore, using data from 2005 to 2019

    Thursday, 30 October 2025

    How people use ChatGPT, for work and not

    The economic impacts of AI (referencing my previous post) are driven by who uses AI tools, and how they use them. We got an sense of this from this working paper by Handa et al., which used data from Claude. However, by far the most widely used generative AI tool is ChatGPT, so I read this new NBER working paper by Aaron Chatterji (Duke University) and co-authors with a lot of interest (see also this blog post by David Deming, one of the coauthors). Many of the co-authors are with OpenAI, which gave them privileged access to user data from ChatGPT. Having said that though, the authors are very clear and very detailed in pointing out the steps they took to ensure data privacy was maintained (and in this matter, this paper is a model for others to follow).

    Their main data is made up of:

    ...a random selection of messages sent to ChatGPT on consumer plans (Free, Plus, Pro) between May 2024 and June 2025.5 Messages from the user to chatbot are classified automatically using a number of different taxonomies: whether the message is used for paid work, the topic of conversation, and the type of interaction (asking, doing, or expressing), and the O*NET task the user is performing.

    Using this dataset (and some related datasets, including one that matches users to the demographic details, while maintaining confidentiality), Chatterji et al. document a lot of important descriptive facts about ChatGPT users (on consumer plans), as well as trends over time.

    First, they document the by-now well-known exponential growth of ChatGPT use over time, summarised in Figure 3 from the paper:

    Not only has ChatGPT use grown over time, but it has also grown within every cohort of users over time (where cohorts are defined by how long ago users first started using ChatGPT). The early adopters still use ChatGPT the most, but every subsequent cohort has increased use over time, Looking at Figure 5 in the paper, it looks to me like there is a noticeable spike in about March-April 2025, when the o3 model was released:

    Turning to the use of ChatGPT for work, Chatterji et al. report that:

    ...both types of queries grew rapidly between June 2024 and June 2025, however non-work-related messages grew faster: 53% of messages were not related to work in June 2024, which climbed to 73% by June 2025.

    Interestingly, later cohorts of users have a greater share of non-work messages than earlier cohorts. However, the differences between cohorts are relatively small, and a majority of ChatGPT messages are non-work-related for every cohort. Nevertheless, there is a lot of ChatGPT-related working going on!

    What sort of work? Chatterji et al. next look at the topics of conversations with ChatGPT, finding that for work-related messages:

    About 40% of all work-related messages in July 2025 are Writing, by far the most common Conversation Topic. Practical Guidance is the second most common use case at 24%. Technical Help has declined from 18% of all work-related messages in July 2024 to just over 10% in July 2025.

    'Writing' includes things like editing or summarising text, or translating. 'Practical Guidance' includes things like how-to advice, and tutoring or teaching. 'Technical Help' includes things like calculations, programming, or data analysis. Including non-work-related conversations:

    The three most common Conversation Topics are Practical Guidance, Seeking Information, and Writing, collectively accounting for about 77% of all ChatGPT conversations.

    'Seeking Information' is basically using ChatGPT as a replacement for web search (a use case that I have become particularly fond of ever since ChatGPT started routinely providing web links in its responses). Of interest to educators should be this:

    Education is a major use case for ChatGPT. 10.2% of all user messages and 36% of Practical Guidance messages are requests for Tutoring or Teaching.

    Of course, that won't count the ChatGPT conversations that relate to the completion of assessments, which are more likely to fall into the 'Writing' or 'Technical Help' categories.

    Chatterji et al. then look at user intent, based on a categorisation of messages into 'asking' (seeking information from ChatGPT), 'doing' (asking ChatGPT to complete a task), and 'expressing' (anything else). For work-related messages, they find that:

    Doing constitutes nearly 56% of work-related queries, compared to 35% for Asking and 9% for Expressing.

    That contrasts with what they see when looking at all messages, where 49% of messages were 'asking' and 40% were 'doing'. Interestingly, 'doing' messages are declining as a share over time, while 'expressing' are increasing. I would have thought that 'asking' messages would have increased, but there is only slight evidence for that (obviously I am extrapolating from my own experience!).

    The work activities results are quite detailed, so I won't discuss them in detail here. However, Chatterji et al. provide the following summary:

    We find that about 81% of work-related messages are associated with two broad work activities: 1) obtaining, documenting, and interpreting information; and 2) making decisions, giving advice, solving problems, and thinking creatively.

    Turning to the demographics of ChatGPT users and their use of ChatGPT, Chatterji et al. report that:

    ...a significant share (around 80%) of the weekly active users (WAU) in the first few months after ChatGPT was released were by users with typically masculine first names. However, in the first half of 2025, we see the share of active users with typically feminine and typically masculine names reach near-parity. By June 2025 we observe active users are more likely to have typically feminine names. This suggests that gender gaps in ChatGPT usage have closed substantially over time.

    We also study differences in usage topics. Users with typically female first names are relatively more likely to send messages related to Writing and Practical Guidance. By contrast, users with typically male first names are more likely to use ChatGPT for Technical Help, Seeking Out Information, and Multimedia (e.g., modifying or creating images).

     Looking at differences by age group:

    Among those who self-report their age, around 46% of the messages in our dataset are accounted for by users 18-25.

    A higher share of messages are work-related for older users. Work-related messages comprised approximately 23% of messages for users under age 26, with this share increasing with age.

    Perhaps younger users are more likely to disclose their age? Having said that, I don't think anyone would be surprised by those results. Nor would they be surprised by the results by level of education:

    Educated users are much more likely to use ChatGPT for work. 37% of messages are work-related for users with less than a bachelor’s degree, compared to 46% for users with exactly a bachelor’s degree and 48% for those with some graduate education. Those differences are cut roughly in half after adjusting for other characteristics, but they are still statistically significant at the less than 1 percent level. Educated users are more likely to send work-related messages.

    The results by occupation are more interesting. Chatterji et al. report that:

    ...the unadjusted work shares are 57% for computer-related occupations; 50% for management and business; 48% for engineering and science; 44% for other professional occupations; and only 40% for all non-professional occupations. Regression adjustment moves these figures around slightly, but the gaps by occupation remain highly statistically significant. Users in highly-paid professional occupations are more likely to send work-related messages.

    The 'regression adjustment' refers to using the results from a multiple regression model that controls for age, gender, education, and some other variables. Looking at user intent and conversation topics by occupation, Chatterji et al. find that:

    ...users in highly paid professional occupations are more likely to use ChatGPT for Asking rather than Doing... This is especially true in scientific and technical occupations. 47% of the work-related messages sent by users employed in computer-related occupations are Asking messages, compared to only 32% for non-professional occupations. These differences shrink somewhat with regression adjustment, but remain highly statistically significant...

    Writing is especially common for users employed in management and business occupations, accounting for 52% of all work-related messages. Writing is also relatively common in non-professional and other professional occupations like education and health care, accounting for 50% and 49% of work-related messages respectively. Technical Help constitutes 37% of all work-related messages for users employed in computer-related occupations, compared to 16% in engineering and science and only about 8% for all other categories.

    Chatterji et al. note that:

    Across all occupations, ChatGPT usage is broadly focused on seeking information and assistance with decision-making.

    People use ChatGPT for what it is best suited for. So, what does this all mean for the economic impact of AI (and specifically, the economic impact of ChatGPT)? Chatterji et al. conclude that:

    ...our findings suggest that ChatGPT has a broad-based impact on the global economy. The fact that non-work usage is increasing faster suggests that the welfare gains from generative AI usage could be substantial... Within work usage, we find that users currently appear to derive value from using ChatGPT as an advisor or research assistant, not just a technology that performs job tasks directly. Still, ChatGPT likely improves worker output by providing decision support, which is especially important in knowledge-intensive jobs where productivity is increasing in the quality of decision-making.

    None of that answers the stream of questions posed by Kevin Bryan (which I outlined in my previous post). Nevertheless, it is important to recognise both how widely used ChatGPT is (in case you've been living under a rock for the last three years), and how it is used, particularly by workers in their daily tasks. This research provides us with broad answers, which can now be supplemented with more detailed analyses of particular industries and occupations.

    [HT: Marginal Revolution]