Tuesday, 22 April 2025

The gender gaps in academia may not arise entirely from gender biases

I've written a lot of posts about the gender gap in academia, in economics and in other (mostly STEM) disciplines (see the list at the end of this post). However, this 2023 review article by Stephen Ceci (Cornell University), Shulamit Kahn (Boston University), and Wendy Williams (Cornell University), published in the journal Psychological Science in the Public Interest (open access), suggests that the gender gap may not be a substantial as previously believed (including by me). This article is quite credible, having arisen as an 'adversarial collaboration', meaning a collaboration between researchers who previously disagreed on the key conclusions from the literature. As Ceci et al. explain:

This article represents more than 4.5 years of effort by its three authors. By the time readers finish it, some may assume that the authors were in agreement about the nature and prevalence of gender bias from the start. However, this is definitely not the case. Rather, we are collegial adversaries who, during the 4.5 years that we worked on this article, continually challenged each other, modified or deleted text that we disagreed with, and often pushed the article in different directions...

Kahn has a long history of revealing gender inequities in her field of economics, and her work runs counter to Ceci and Williams’s claims of gender fairness.

Ceci et al. focus on seven questions of relevance to understanding the gender gap in academia:

In this article, we comprehensively examine evidence in six key evaluation contexts: (a) Are similarly accomplished women and men treated differently by academic hiring committees? (b) Are grant reviewers biased against female PIs? (c) Are journal reviewers biased against female authors? (d) Are recommendation-letter writers biased against female applicants for tenure-track positions? (e) Are faculty salaries biased against women? And, (f) are student teaching evaluations biased against female instructors? Claims of gender bias are omnipresent in all six of these domains... (We also review the literature in a seventh context, gender differences in publication rates, because publishing productivity can moderate evaluation in most of these six contexts.)

Ceci et al. focus on research published since 2000, which is more likely to represent the 'current' state of academia (although, arguably, they should weight more heavily more recent studies, which they don't). They also distinguish between:

...the most mathematically intensive fields—geosciences, engineering, economics, mathematics/computer science, and physical science (GEMP)—and less math-intensive fields—life sciences, psychology, and social sciences (LPS).

It is generally claimed that gender gaps are more prevalent in the GEMP fields than in the LPS fields (and a quick read through the links at the end of this post would suggest that is definitely true of economics, as one example).

The article is very thorough and has a lot of detail related to each of those contexts, so I'm just going to hit the headlines in relation to each of the seven research questions. If you are interested in any particular finding, the article is open access, so you can easily look at their review and unpack the details there. In relation to the first research question (are similarly accomplished women and men treated differently by academic hiring committees?), Ceci et al. conclude that:

The vast majority of findings—from (a) synthetic cohort analysis, (b) institutional hiring records, and (c) experiments—indicate that women are less likely than men to apply for tenure-track jobs, but when they do apply, they receive offers at an equal or higher rate than men do.

So, the news is both good and bad. On the positive side, there is little evidence of bias. However, where a gender gap in hiring persists, and is not because of bias in hiring, the gender gap must arise from differences in the rates of applying for academic positions between men and women. Indeed, Ceci et al. note that:

...women are more likely than men to give up their initial aspirations to become tenure-track professors while in graduate school, a finding primarily true of women with children or contemplating children. Undoubtedly, broad systemic factors are partly responsible, along with biological factors, for these women not applying for tenure-track positions.

That still suggests that there is further work to do (both in terms of research and in terms of addressing the problem), but this particular article skirts around that issue because it focuses on the gender biases for academics, and not the graduate student experience. In relation to the second research question (are grant reviewers biased against female principal investigators (PIs)?), Ceci et al. conclude that:

...pre-2006 evidence suggests that although some agencies evaluated men and women differently, on average they did not.

Ceci et al. then conduct their own meta-analysis of the literature since 2000 (including 39 studies), and conclude that:

Taken together, both the analytic dissection and our meta-analyses appear not to support the claim that the grant peer-review process has been rigged against women PIs during the past 20 years in the United States. This is particularly true when analyses controlled for PIs’ research productivity...

On the third research question (are journal reviewers biased against female authors?), Ceci et al. conduct both a review and meta-analysis and conclude that:

...overall, our meta-analyses and our dissection of key studies revealed no evidence of systematic bias against female authors, notwithstanding claims to the contrary.

For the fourth research question (are recommendation-letter writers biased against female applicants for tenure-track positions?), Ceci et al. conclude that:

On the basis of our analysis of the nine studies in this domain, we conclude that no persuasive evidence exists for the claim of antifemale bias in academic letters of recommendation.

In relation to the fifth research question (are faculty salaries biased against women?), Ceci et al. conclude that:

...the evidence supports the claim that women are paid less than men in tenure-track academia, although the magnitude of the gap is much smaller (60%–80% smaller) than often claimed in executive summaries and headlines, and in some situations has disappeared.

Ceci et al. also dig a bit deeper on the salary differences, noting that:

Some of the unexplained gender salary gap may be due to implicit bias (although this seems unlikely in biology, where starting salaries are higher for women), and some of it may be due to differences in willingness to negotiate and solicit outside offers... some of the remaining pay gap may be due to women’s work discontinuities for family leave... or to a desire to keep jobs flexible... Finally, some of the relatively small remaining pay gap may be due to women’s lower likelihood of negotiating higher salaries or their lower likelihood of pursuing more lucrative job offers. The lower likelihood of negotiating higher salaries may itself be due to bias... Without specific data on family leaves, past employment, and job pursuit, it is impossible to know how much, if any, of the less than 4% unexplained pay gap is attributable to bias.

The salary gap of 4% is small, but it is not zero. The fact that much of the gap can be explained by the factors outlined above would accord with research by 2023 Nobel Prize winner Claudia Goldin (whose work they cite, among others). However, that leaves open the question of how large the salary gap would be, after accounting for work discontinuities, preferences for flexibility, and negotiation? Again, a research question to be addressed in the future.

On the sixth research question (are student teaching evaluations biased against female instructors?), Ceci et al. conclude that:

...the evidence supports the claim that female instructors are penalized for being women, independent of the content and delivery of their lectures and independent of students’ actual learning. The effect sizes we calculated indicate penalties for women that ranged between small and moderately large (ds = 0.10–0.50). So, unlike the domains in which we were able to unequivocally reject claims of widespread gender bias, in this domain, we conclude that there is gender bias.

This is consistent with many research findings on student evaluations of teaching, including studies I have blogged about before (most recently here, and see the other links at the end of that post for more). Unfortunately, this seems to be a pervasive finding across all teaching contexts, and it doesn't appear to be getting any better.

Finally, in relation to the additional research context (gender differences in research productivity), Ceci et al. conclude that male researchers do have higher research productivity (more publications), and that:

...gender productivity differences are smallest in GEMP fields (with the exception of economics) and are largest (and possibly growing) in biology, psychology, and economics.

The overall takeaway from this work is that there are still gender gaps in academia, but that many of the gaps, or claimed gaps, don't seem to arise from gender bias. At least, that's what we should conclude from the research to date. That is true of all of the domains except salaries and teaching evaluations. None of this means that we should conclude that all is rosy for female academics (unlike the title of this post), and that is certainly not the case across all fields. Indeed, some of the changes in recent years might actually make things worse for female academics before they get better (as noted in yesterday's post). We still have some way to go, especially in the more technical fields, including STEM and economics.

[HT: Marginal Revolution, back in 2023]

Read more:

Monday, 21 April 2025

Co-authorship in economics in the aftermath of #MeToo

The #MeToo movement was a necessary corrective action recognising decades of toxic behaviour across many occupations. Economics was not immune (for example, see here or here). However, could the #MeToo movement have had an unintended consequence on the careers of female economists? If having a female co-author increases the chances of a male economist being called out for even minor indiscretions, does this meaningfully raise the cost of having female co-authors? And if the cost of having female co-authors meaningfully increases, we would expect to see fewer male-female collaborations (especially where the male economist is more senior).

That is the topic addressed in this new article by Noriko Amano-Patiño, Elisa Faraglia, and Chryssi Giannitsarou (all Cambridge University), published in the journal European Economic Review (open access). The use data on co-authorships in the nearly 27,000 working papers published in the NBER and CEPR working paper series between January 2004 and December 2020. They first note that:

The MeToo movement’s impact on the economics profession may have fostered a more respectful research environment, increased scrutiny of existing practices, and promoted greater diversity and inclusivity within the community. Conversely, it could have induced a chilling effect on collaborations, potentially causing researchers to become more hesitant in forming partnerships outside their established networks due to heightened concerns about trust and reputational risk.

Although the #MeToo movement started in 2017, Amano-Patiño et al. use the second quarter of 2018 as the effective date for the onset within the economics profession (that dates to the fallout arising from a series of studies, including a particularly notable study by Alice Wu, which I blogged about here). However, Amano-Patiño et al. vary the effective start date and find little difference in their results. So, comparing papers written before and after 2018, and controlling for a variety of author characteristics, Amano-Patiño et al. find that there was:

...a rise in the proportion of women coauthors for men, both overall and within junior and senior subgroups. Conversely, we find a decrease in the proportion of women coauthors for women, both overall and within corresponding seniority levels. Using a back-of-the-envelope calculation, these increases in mixed-gender collaborations, translate to an estimated 12.3% increase in women coauthors per 100 men-authored papers.

That seems to go against what we might expect, if the cost of having female co-authors has increased for male economists after #MeToo. However:

...we estimate decreases in the proportion of senior coauthors (especially senior women) for juniors, and symmetrically, in the proportion of junior coauthors (particularly junior women) for seniors. The decreases in collaborations between senior and junior economists we quantify, suggest a 3.0% decrease in the share of senior authors collaborating with junior coauthors.

 Amano-Patiño et al. interpret this as showing that:

...post-MeToo, authors have increasingly sorted their collaborations by seniority rather than by gender.

What might explain these findings? Researchers who are worried that they might get called out by female co-authors might respond by reducing their collaborations with female co-authors generally, as I noted at the start of this post. Or, they might reduce their collaborations with new co-authors, who they have not developed trust with, while continuing to collaborate with more senior authors that they trust. This is also consistent with Amano-Patiño et al.'s further results, where they note that:

...we find evidence of a general chilling effect on the expansion of economists’ professional networks. We estimate decreases in the share of new coauthors across all seniorities and genders, the share of new senior coauthors for juniors, and the share of new junior coauthors for seniors. Our estimates translate into 5.4% fewer new coauthorships per 100 papers. This trend is primarily driven by a substantial decrease in new coauthorships between senior and junior authors: for seniors, the share of new junior coauthors has dropped by 18.4%, with a particularly sharp 48% decrease in their share of new junior women coauthors.

Amano-Patiño et al. interpret their results as bad news, noting that if the results can be interpreted as causal:

First, authors may have prioritised increasing gender diversity in their collaborations. Second, senior authors have increasingly relied on their existing collaboration networks rather than forming new coauthorships. The latter trend, if persistent, could have long-lasting consequences for the career development of women economists and potentially exacerbate the already ‘leaky’ pipeline in the profession.

 Amano-Patiño et al. stop short of noting that this is a substantial negative unintended consequence of the #MeToo movement in economics. Although the environment for female economists may be improving, at least one aspect, being the opportunities for collaboration and mentoring from senior economists, appears to be declining. And that will be a difficult problem to address. Indeed, Amano-Patiño et al. aren't able to offer any concrete steps that could be implemented to solve this issue, concluding with some more general statements:

These results underscore the urgent need for sustained efforts to cultivate a supportive ecosystem and dismantle systemic barriers hindering the advancement of women and junior economists in the field. The economics profession must proactively continue to foster a safe, inclusive environment by evaluating, monitoring, and educating on relevant issues.

Sadly, I also can't offer anything concrete, and only hope that the current desire for change within the profession will ultimately lead to greater opportunities for female economists overall.

Sunday, 20 April 2025

Grumpy young economists

Academic writing has changed over time, as I noted in this post back in 2022. The research I referred to in that post identified an increase in the use of adjectives and adverbs over time, noting that as a result, research was becoming less readable over time. The research speculated on the reasons why research had become less readable, but one explanation that they didn't consider was that different generations of academics might express themselves in different ways.

And that is essentially what this 2024 article by Lea-Rachel Kosnik (University of Missouri-St. Louis) and Daniel Hamermesh (University of Texas at Austin), published in the Southern Economic Journal (ungated earlier version here). sets out to look at. Kosnik and Hamermesh look at a sample of all 15,138 articles published in the 'Top 5' economics journals between 1969 and 2018 (the 'Top 5' journals are American Economic Review, Econometrica, Journal of Political Economy, Quarterly Journal of Economics, and Review of Economic Studies). Once they restrict the sample to authors with at least five articles, the sample reduces to 1389 researchers (and 12,812 articles).

Kosnik and Hamermesh then apply sentiment analysis to the articles in the sample, resulting in three scores:

...a positive/negative score (POSN), a certain/tentative score (CERT), and a contemporaneity/past score (CONP).

The POSN score reflects the (positive or negative) emotive tone of the writing, CERT measures how certain or tentative the writing is, and CONP measures whether the writing is contemporary or focused on the past. The measures are normalised by subtracting the average score for all articles in the same field of economics (the same JEL group). Kosnik and Hamermesh look at how these normalised measures vary systematically across the sample of authors and over time, paying particular attention to how the measures are related to the number of years since each researcher completed their PhD. They find that:

Based on the fixed-effects estimates for the entire sample (the 1970s cohort), a one standard-deviation increase in age leads to changes of 0.07 (0.02), -0.03 (-0.01), and -0.05 (-0.03) standard deviations in POSN, CERT, and CONP, respectively.

In other words, older economists write in a more positive emotive tone. However, the effects for CERT and CONP are not statistically significant. Kosnik and Hamermesh then pivot to looking at the square of each normalised measure, contending that it represents the deviation from the norms. It isn't clear to me why they consider that 'more negative' and 'more positive' deviations in norms should be treated identically, and so I don't find that analysis particularly illuminating. It seems like an arbitrary approach, and they note in a footnote that using the absolute value of the normalised measure rather than its square makes the results less statistically significant. That should also give us pause.

However, the basic analysis does provide some other points of interest, including:

Natives write less positively, with less certainty, and with less present/future orientation than do leading economists whose mother tongue is not English. This is true, however, only for those native English-speakers who grew up in North America (57% of authors) or the United Kingdom (5% of authors), whose styles of writing economics are almost identical along the three measures we examine. The styles of the 2% of authors whose native English comes from elsewhere (Ireland, South Africa, Australia, or New Zealand), however, do not differ from those of non-native speakers.

And:

There are also significant differences across the five journals, with all of them being more positive and more contemporary-oriented than the AER, and all but the QJE being written in a more certain voice than the AER...

Clearly, the most dismal scientists get published in the American Economic Review. And:

Additional coauthors, however, do make writing styles more positive, more certain, and less present-oriented, both in the full sample and in the 1970s cohort.

Are sole authors more negative because they have to do all of the work themselves, I wonder? Anyway, Kosnik and Hamermesh then turn to looking at citations, finding that:

While positive deviations of all three measures of sentiment reduce citations significantly or nearly so, the more important question is how large these reductions are. Taking simultaneous one-standard deviation increases in sentiment scores... these increases reduce citations by 5 (2.5)%, or 0.015 (0.01) standard deviations. Writing in a more positive, more certain, or more present-oriented way than others publishing at the same time and in the same sub-field reduces the scholarly impact of one's articles, although the effects are quite small.

Decomposing the change in citations as authors get older, Kosnik and Hamermesh find that:

Scholarly recognition decreases with author's age, but only a small part of the decrease is due to changes in writing style with age.

Finally, Kosnik and Hamermesh look at the subset of Nobel Prize winners, and find that:

Nobelists' style exhibits significantly less certainty than that of other star authors. This example suggests that writing in a more tentative style distinguishes one's scholarship and might provide the scope for subsequent researchers to accord it the attention that helps to generate the distinction of a Nobel Prize.

What do we take away from all this? There is a lot of depth in the analysis, but when we put aside the analyses that rely on the squared measure (which, as I noted above, I don't have as much faith in), it seems that the only remaining result (in terms of age) is that older economists write in a more positive tone than younger economists. Fortunately, the impact of tone on research impact (as measured by citations) is fairly small, so I guess younger economists can afford to be grumpy.

Coming back to where I started this post, what does that imply for changing writing styles over time? If younger economists write in a more negative tone, then as the population (including the population of economists) ages, we might see more positively minded economics writing! Now, the question arises, do younger researchers in other disciplines also write in a more negative tone than older researchers?

[HT: Marginal Revolution, back in 2023]

Saturday, 19 April 2025

The 'rising stars' of economics education (and yes, it includes me!)

I've been meaning to post about this working paper by Wayne Geerling, Dirk Mateer (both University of Texas at Austin), and Jadrian Wooten (Virginia Polytechnic Institute and State University) for a while (it was highlighted in a 'this week in research' post back in November, and I read it then). The paper analyses citation data from articles on economics education published in economics journals between 2019 and 2023. It then ranks authors based on citation counts and i10-index (the number of published articles with ten or more citations).

The top 50 authors are shown in Table 1 in the paper. The top four will not come as a surprise to anyone familiar with the literature: William Walstad, Sam Allgood, William Becker, and KimMarie McGoldrick. However, if you scroll a little further down the list, there I am, ranked at #27, with 161 citations and an i10-index of three. And if you look carefully at the list, you'll see that almost all of the names above me are from US institutions. So, I am ranked #5 among non-US-based economics education authors! Sadly, I don't make it onto the list of the top authors based on i10-index, which they cut off at four. However, this is still welcome recognition of my research on economics education.

However, it is worth noting that my performance is almost entirely driven by this one article (ungated earlier version here). That article on financial literacy among high school students was co-authored with Richard Calderwood, Ashleigh Cox, Steven Lim, and Michio Yamaoka, all of whom also make it onto the authors list in joint 30th place, with the majority of their 145 citations coming from that one article, which has 118 (the rest of their citations will be from this other article (ungated) co-authored with me, from the same project). The article on financial literacy among high school students is actually the fifth most cited in the whole sample used by Geerling et al.

I have a number of economics education projects on the go. Most of them involve data drawn from my ECONS101 classes, where I am always trying (and evaluating) something new. This trimester. we've been trialling the use of a generative AI tutor, Harriet, which I posted about at the start of the year. We've also given students the opportunity to do practice multiple choice questions every day ('Question of the Day') on Moodle. Both of those initiatives will be evaluated in terms of their contribution to student performance in assessment. However, even having done the evaluation, that doesn't guarantee that I find the time to write the analysis up as a paper. But if I want to retain my ranking in Geerling et al.'s list, then I will have to make a more concerted effort to get my economics education research published!

[HT: Wayne Geerling]