Monday, 13 February 2023

Gender differences in reactions to editorial decisions in economics

There is a persistent gender gap in economics, which starts at undergraduate level and extends through graduate study, junior academic positions, tenure, and to the professoriate. The gender gap gets worse as you move up through each level, with women making up a smaller proportion as you go. Many studies have now looked into this, but one aspect is relatively under-explored - how different genders respond to setbacks in the publishing process.

That research gap is what makes this recent working paper by Gauri Kartini Shastry and Olga Shurchkov (both Wellesley College) most interesting. They presented a sample of around 1300 academic economists with a hypothetical scenario regarding the rejection of a paper. Specifically:

Respondents then read a hypothetical decision letter from an editor which begins by describing referee reactions to a paper the respondent hypothetically submitted for publication at a top general-interest journal. Our experimental manipulation randomizes respondents into three main treatments: a revise and resubmit decision (R&R), a reject and resubmit decision (RJR), and a flat rejection (FR); this randomization affects only the last few lines of the decision letter. A key feature of the design is that outcomes are measured following almost identical decision letters from an editor on a hypothetical paper, ensuring that all respondents are given the same information about the quality of this submission. We are interested in the differential effect of the more negative treatments (FR/RJR) relative to the baseline (weak R&R) on the respondents’ perceived likelihood of eventually publishing the paper in a highly-regarded journal.

Shastry and Shurchkov find that:

...getting an RJR reduces the perceived likelihood of publishing the paper in any leading journal relative to an R&R, but getting a rejection has the most negative impact. On average, negative decisions have similar effects on men and women. However, we observe heterogeneous effects by rank: female assistant professors who get a rejection perceive an 18 pp lower likelihood of publishing the paper in any leading journal than male assistant professors, as compared to the difference across genders for those who get an R&R. The gender gap is not present among tenured associate and full professors.

Shastry and Shurchkov try to tease out the mechanism explaining these results, positing that:

...female assistant professors attribute the negative feedback of a rejection to subpar quality of their work to a greater extent than men do, and that this is exacerbated by the time constraint of an upcoming review.

There is some evidence to support this explanation. However, the more interesting (to me) finding is that the gender gap closes after tenure. On this point, I strongly suspect that there is some survivorship bias. That is, economists who are less resilient to negative feedback may be less likely to achieve tenure, and if male economists are (because they are more confident, or some other reason) more resilient, then more male economists would survive to tenure than female economists. And both male and female economists who survive to tenure will tend to be similar in terms of how they respond to negative feedback. That's speculative on my part. Shastry and Shurchkov have no evidence in favour of it, but neither are they able to discount it as a possibility.

This paper is useful in filling in a bit more detail on how the gender gap in economics arises. However, it would be useful to build on this work in order to better understand the mechanisms that underlie it, because it is the mechanisms that are the source of the problem.

[HT: Marginal Revolution, last year]

Tuesday, 7 February 2023

What machine learning is telling us about the gender dynamics in economics seminars

Economics seminars have come in for a lot of criticism for apparent gender bias (see here, or for a broader view on gender in economics see the links at the end of this post). To some extent, past research on gender dynamics in economics seminars has been limited by the hand-coding of data, so limited detail on the interactions within the seminar are available to analyse. Not any more, thanks to machine learning, as explained in two new papers.

The first paper, by Amy Handlan and Haoyu Sheng (both Brown University), uses machine learning for audio analysis of presentations at the 2022 NBER Summer Institute. The audio classification algorithm that they use allows them to automatically classify speakers by gender, age, and the tone of their comments. They find that:

...women in the audience are more likely to ask female presenters questions than male presenters, and female presenters are more likely to be assigned female discussants.

This suggests that there is significant gender sorting (although Handlan and Sheng can't say anything about the wider audience in each presentation, because they only have data on audience members who spoke). Interestingly, next they find that:

Regarding interruptions, we find that there are similar number of interruptions for both male and female presenters.

That is somewhat at odds with earlier research (see here, for example). However, more consistent with the earlier research, Handlan and Sheng also find that:

...interruptions for female presenters last longer than those for male presenters...

But what about tone? In that respect, they find:

...gender differences in tone within speakers. On average, female speakers are more likely to sound positive or happy, while male speakers are more likely to sound negative and serious or angry. This holds whether we consider only presenters or only speaking audience members. Furthermore, when we consider how speakers may change their tone over time, we find that tone is highly persistent for both men and women. That is, if you sound happy and are uninterrupted, you are likely to continue sounding happy.

So far, so good. But what happens when speakers are interrupted, or responding to others? In that case:

...we find that speakers sound more negative when responding to women, whether the person is an audience member asking a question to a female speaker or a presenter responding to a female audience member. When we look at the interaction between presenter gender and audience gender, we find that female speakers respond more negatively to other women compared to how they respond to men. The gap in tone responding to men versus women is larger for female speakers than for male speakers.

This finding that women are harsher towards other women than men are towards women, is unfortunately looking increasingly common in economics. Handlan and Sheng suggest that this may be due to societal norms that lead to higher expectations of women, as well as because:

Women may speak more positively to men in an attempt to offset a larger negative bias from men compared to women.

I'd label those reasons speculative for now. Handlan and Sheng aren't able to test them, of course. However, the paper has some interesting insights, and not just about gender. For instance, people (presenters and audience members) in macroeconomics seminars are more likely to sound sad, but less likely to sound angry, than people in microeconomics seminars. Given that this was 2022, historically high inflation might make macroeconomists sad, and microeconomists angry? We need some causal analysis of tone!

Anyway, moving on. The second paper I want to discuss is this job market paper by Mateo Seré (University of Antwerp). His dataset is much broader than Handlan and Sheng's, covering all web-streamed (on YouTube) seminars that were part of a seminar series of a university economics department in the US or Europe, over the period from 2020 to 2022 (as well as the National Bureau of Economic Research (NBER), the American Economic Association (AEA), and the Centre for Economic Policy Research (CEPR)). Seré focuses specifically on interruptions to the seminar, and first notes that:

...the distribution of interruptions made during seminars presented by women is slightly shifted to the right compared to the distribution for males.

Specifically, female presenters receive between 0.9 and 1.5 more interruptions than male presenters, controlling for presenter characteristics and seminar characteristics. That is a large effect, given that the mean number of interruptions is 11. So, female presenters receive about 10 percent more interruptions than male presenters. That is inconsistent with Handlan and Sheng's results, but remember that their results come from the NBER Summer Institute only, whereas Seré is looking at a much broader set of seminar series.

When looking at who does the interrupting, Seré finds that:

...the proportion of female interruptions is significant to explain the overall number of interruptions in a seminar only when the presenter is a woman.

In other words, the additional interruptions that female speakers experience are mostly driven by female audience members. Seré notes that male audience members ask female presenters 0.2 more questions than they ask male presenters, but female audience members ask female presenters 1.0 more questions than they ask male presenters. Given than most audience members are male (and so they ask more questions overall), this is a substantial difference.

Seré then finds that there is a difference in the nature of interruptions between male and female audience members:

Being a female presenter is related with an increase in the number of questions from females in the audience and with a decrease in the ones made by males. Furthermore, while men in the audience make more comments on average when the presenter is female, the gender of the presenter has no effect on the number of comments made by women.

This is interesting, as it is closest to the results on tone from Handlan and Sheng. Questions require a response, whereas comments do not. Looking across both studies, if female audience members ask more questions of female presenters rather than making comments, perhaps those questions come across as more negative in tone?

Together, these two papers paint an interesting picture of some of the dynamics within economics seminars. Clearly, this is just the beginning of research in this area. In particular, Handlan and Sheng have made their algorithm available for others to use for coding and analysing other seminar series. Expect more research in this space in the future.

[HT: Marginal Revolution for the Handlan and Sheng paper, and Development Impact for the Seré paper]

Read more:

Sunday, 5 February 2023

Mendelian randomisation doesn't necessarily overturn the alcohol J-curve

The alcohol J-curve is the common empirical finding that moderate drinkers have better health than abstainers, and better health than heavy drinkers (see this post, for example). If you plot the relationship between the amount a person drinks and negative measures of health, the resulting curve is shaped like the letter J (as in that earlier post). However, few of the studies that establish a J-curve relationship demonstrate a causal relationship. That's because it is difficult to randomise people into a level of drinking.

However, a relatively new development in epidemiology is the idea of Mendelian randomisation. People are randomly assigned genes at birth. Some of those genes are associated with alcohol consumption. So, alcohol consumption is (partially) randomly assigned by the assignment of genes. Studies can then use instrumental variables regression (a relatively common technique in economics) to estimate the causal effect of alcohol consumption on a range of health (and other) outcomes.

What happens to the J-curve in these Mendelian randomisation studies? This editorial published in the journal Addiction in 2015 (open access), by Tanya Chikritzhs (Curtin University) and co-authors, summarises the state of knowledge up to that point. They note that there is no J-curve relationship observed in Mendelian randomisation studies, or at least that the results are much more equivocal about its existence, and conclude that:

The foundations of the hypothesis for protective effects of low-dose alcohol have now been so undermined that in our opinion the field is due for a major repositioning of the status of moderate alcohol consumption as protective.

I was recently referred to this editorial during a discussion on the J-curve. Having read a bit more about Mendelian randomisation though, I am not entirely convinced. The problem is that in these Mendelian randomisation studies, the two main assumptions of instrumental variables analysis must be met. First, the instrument (having the gene, or not) must be associated with the endogenous variable (alcohol consumption). That assumption should be relatively easy to meet. A researcher simply searches the literature on genome-wide association studies for some gene that is associated with alcohol consumption. Second, the instrument must only affect the dependent variable (health outcomes) through its effect on the endogenous variable (alcohol consumption), and not directly or through any other variable. This is known in economics as the exclusion restriction.

The exclusion restriction is a difficult to satisfy, in part because it is impossible to test statistically. Instead, most researchers settle for being able to identify an instrument that is 'plausibly exogenous'. That is, they find an instrument that is extremely unlikely to affect the outcome variable directly, or indirectly through any other mechanism than through the exogenous variable. In this case, that would mean identifying a gene that could only possibly affect health outcomes through its effect on alcohol consumption.

And that is the problem here. There is no gene for alcohol consumption. All that genome-wide association studies will identify is genes that are associated with alcohol consumption. Those genes all have some other purpose, and that other purpose may be linked to health outcomes in a way that doesn't involve alcohol consumption. As far as I am aware, the Mendelian randomisation studies to date haven't been able to establish that the gene in question has no other effects on health outcomes. That makes their claims to causality no better than those of the correlational studies that they are supposed to improve on.

Mendelian randomisation is good in theory, but not always in practice. For now, the J-curve lives.

Saturday, 4 February 2023

Academic achievement and class scheduling

It's only a few weeks until classes start for 2023. For the first time in many years, I will have a lecture scheduled in the dreaded Monday 9am timeslot. Will any students show up so early in the week? And if they do, will they be so bleary-eyed from a weekend of working and partying that it negatively impacts their learning?

I'll find out an answer to the first question on the first day of classes. Given my experience last trimester, I'm not holding out hope for high attendance (although that will seriously be to the detriment of the students - more on that in a future post, as I am currently looking into how students' engagement in my classes affected their performance). As for the second question, this 2018 article by Kevin Williams (Occidental College) and Teny Shapiro (Slack, Research & Analytics), published in the journal Economics of Education Review (sorry I don't see an ungated version online), provides some answer.

Williams and Shapiro use data from the US Air Force Academy, where students:

...alternate daily between two class schedules within the same semester. Students have a similar academic course load, but the alternating schedule creates variation in how much time students spend in class on a given schedule-day. This allows us to assess how a student performs with one schedule relative to their own performance with a different schedule.

Also, since students are randomly assigned to instructors and schedules, this provides an opportunity to estimate the causal effect of different classroom schedules on students' academic performance. This random assignment has made the USAFA data a popular choice in the economics of education (see here, for example). In terms of scheduling:

USAFA runs on an M/T schedule. On M days, students have one set of classes and on T days they have a different set of classes. The M/T schedules alternate days of the week...

Williams and Shapiro observe that there are two effects of class scheduling on student academic performance:

The first is the cognitive load a student has experienced before the start of a class. We refer to this as the student fatigue effect or cognitive fatigue. The second is the timing of a class: students may perform less well if classes are scheduled when they’re naturally less alert. We refer to this as the time-of-day effect. We expect student fatigue to unambiguously hinder academic performance. The time-of-day effect may vary throughout the day.

Williams and Shapiro use data from 6981 students (of whom 4788 are freshmen), and in total have over 230,000 course-level observations (of which over 180,000 are core courses). The key outcome is students' normalised grade, while the explanatory variables that Williams and Shapiro are most interested in are the number of consecutive classes and the number of cumulative classes on that day (to measure the student fatigue effect), and the time of day (to measure the time-of-day effect). They find that:

Consecutive classes have a consistently negative impact on performance... A student sitting in their second consecutive class is expected to perform 0.031 standard deviations worse than if she took the same course after a break. We take this as solid evidence of cognitive fatigue. When student’s schedules require them to sit in multiple classes in a row, they perform significantly worse in the latter classes, likely because of a decreased ability to absorb material.

The effect of cumulative classes, the total number of prior courses a student has taken on a given day, varies more across our models, but is significantly negative in our preferred specification... This suggests that students suffer both from the immediate effect of consecutive classes and the cumulative effect of heavy course loads in a single day...

The penalty for students taking a 7 am hour is consistent and robust. All time-of-day coefficients are positive and significant, which suggests holding class every hour of the day after the first benefits students. These effects are large in magnitude. Students taking a 9 am or later are expected to perform 0.16 standard deviations better than students taking the same class at 7am.

So, there is evidence for both negative student fatigue effects and time-of-day effects, with students performing better when they have classes later in the day. However, given my own predicament, the results are not all bad. Williams and Shapiro compare class times with a 7am class. When you compare with a 9am class, student performance is not substantially worse than later in the day, at least up to 1pm (phew!).

However, not all students are affected the same. When Williams and Shapiro separate students into terciles of academic ability, they find that:

Fatigue has the largest impact on the bottom tercile of USAFA students. A consecutive class reduces a bottom-tercile student’s expected grade by 0.042 standard deviations, compared with 0.030 or 0.019 for top-tercile and middle-tercile students, respectively. Two or more consecutive classes only have a significantly negative impact on bottom tercile students.

As with so many things, the students at the bottom of the ability distribution are the worst affected (although, it is worth noting that the bottom of the USAFA distribution is still students who are in the top 15 percent of all high school graduates).

So, class scheduling does appear to make a difference to student grades. Now, we can't schedule all classes in the late afternoon. But, when there are classes that students actually attend in person (rather than those they mostly watch recorded or online), perhaps those in-person classes could be scheduled later in the day? I might have to make some enquiries before the scheduling of my classes next year.