Monday, 20 December 2021

Three papers on the gender gap in economics seminars

The culture of the economics profession has attracted a lot of attention of late (see this post, or the list of posts at the end of this one for more on the gender gap in economics more generally). Through all of this attention, there have been some interesting and important changes underway (see this post, for example). A reasonable question, then, is whether things are improving.

Three recent papers may provide some early indications, in relation to economics seminars. The first is this NBER Working Paper by Pascaline Dupas (Stanford University) and co-authors. They systematically collected data from 460 economics seminars and job market talks across 32 top universities from January to May of 2019, as well as presentations at the NBER Summer Institute. So, these data don't tell us much about changes over time, but may provide a useful baseline against which to compare progress. Importantly, their data focuses on the questions and interruptions that speakers face, which is a key aspect of the climate of seminars and presentations (see this post, for example). Overall, Dupas et al. find from the seminar and job market talks data that:

On average, roughly 26 questions are asked during a regular economics seminar and 35 questions are asked during a job market talk. For a 90-minute seminar, this represents one interruption every 3.5 and 2.5 minutes, respectively - although interruptions are not uniformly distributed during the time allotted. Moreover, there is considerable heterogeneity with the number of questions ranging from a low of 5 to a high of 69 for any given seminar. There are 3.6 times as many questions from men as from women during regular seminars - and 7.6 times during job market talks - despite men only outnumbering women roughly 2 to 1 in attendance (and 3 to 1 in job talks)...

In terms of the type of questions, roughly 35 percent of all questions in regular seminars (37 percent in job market talks) are classified as clarifications, followed by another 17 percent (13 percent) that are classified as comments. Suggestions, follow-ups, and criticisms each account for 10 percent or less in both regular seminars and job market talks - perhaps countering the reputation that economics is as an overly critical discipline.

Comparing male and female presenters, they find that:

...women presenters are asked 3.8 additional questions (p<0.01) relative to men (a 12 percent increase). Accounting for the influence of a range of other factors about the audience, the presenter, the topic, and the coders, reduces the differential to 2.4 questions (p<0.05). This disparity appears most pronounced during recruitment (“job market”) talks (3.7 extra questions, p<0.05) and regular seminars with an external (rather than in-house) speaker (2.6 extra questions, p<0.05)...

Although we find that women receive a greater number of suggestions and clarifying questions, we also find that they are more likely to be asked questions that are rated as patronizing or hostile.

The latter finding should be particularly of concern. Also:

Aggregating across negative tones (questions which are patronizing, disruptive, demeaning or hostile), women receive 0.5 more such questions, three quarters of which seem to be coming from men.

Dupas et al. also analyse a much more limited dataset from the NBER Summer Institute, but find broadly similar results. However, this dataset allows them to investigate whether the format of the seminars makes a difference. On that point, they find that:

...having a discussant and/or Q&A at the end does not mitigate the differential treatment of women presenters. Indeed, women receive more questions than men even in those presentations that had formal discussants... The only mitigating factor appears to be the “moratorium” on questions in the first 5 or 10 minutes of the talk: with the caveat that this represents a very small sample of presentations (N=45), we find that the moratorium completely undoes (if anything, reverses) the gender gap. And this appears to be the result of fewer “clarifying” questions that end up being deferred anyway or followed up on later when asked too early.

So, there is some suggestive evidence that the format of the seminar may make a difference. In my experience, people interrupting the speaker to ask clarifying questions are often not adding value to the talk for anyone but themselves. The takeaway from this paper, though, should be that economics still has a lot of work to do to improve the seminar climate.

The second paper is this article by Jennifer Doleac (Texas A&M University), Erin Hengel (University of Liverpool), and Elizabeth Pancotti (Employ America), published in the papers and proceedings issue of the American Economic Review (ungated version here). Doleac et al. report on the characteristics of invited seminar speakers in a panel of 66 economics departments (60 from the US, and 6 from outside the US) from 2014 to 2019. So, here we get a longitudinal dimension, although Doleac et al. don't have data on the dynamics within each seminar. They focus on the trends over time, distinguishing between genders and between under-represented minority (URM) speakers and non-URM speakers. Here's the key trends, from their Figure 1:

If you squint your eyes, you may pick up a slight upward trend in the proportion of speakers who are non-URM women, and a slight downward trend for non-URM men (the medians are less noisy, so may better represent the trends over time). However, there has not been much change for URM speakers (men or women). Indeed:

Forty-three of the 66 departments in our sample did not invite a single URM woman to speak during this period; 39 did not invite a single URM man to speak.

Maybe there's some small sign of improvement for non-URM women in terms of their proportion of invited speaking opportunities. However, Doleac et al.'s data only goes up to 2019, and it would be interesting to see how more recent trends have played out, especially given the pandemic. Fortunately, that's pretty much what the third paper does, which is this discussion paper by Marcus Biermann (UC Louvain). Biermann looks at how the coronavirus pandemic has affected economics seminars, using data from the seminar series of 270 institutions worldwide (including 243 universities, 14 central banks, 11 research institutes, and 2 international organisations). The dataset includes over 12,000 seminars over the period from 2018 to 2020. Overall:

At the seminar level, 21.8 percent of the seminars are held by female speakers. The average speaker has about 12.2 years of experience after PhD award. The top 1 percent of researchers in terms of their overall output and in terms of their publication record in the last 10 years account for 7 and 12.5 percent of seminars, respectively. The 200 top young economists held 3.3 percent of the seminars...

Looking at the impact of the pandemic, Biermann finds:

...a 7.47 percentage point increase in the relative likelihood that the seminar speaker after the technology shock is female, which is about 34.3 percent in terms of the pre-technology shock mean.

That is quite a substantial effect. Also, he finds that the increase in likelihood of a speaker being female is larger for distances between the home and host institutions of between 1475 and 5000 kilometres. Biermann suggests that:

This implies that parts of the increase in the share of female speakers are driven by a supply side response for medium length distances. The requirement to travel to medium length distant places and to stay overnight may have hindered women to accept seminar invitations before the technology shock.

 Biermann's paper also shows some other interesting trends separate from the change in gender composition of seminar speakers, including:

...that the overall number of seminars declined and that the decline was not driven by the short-run supply of speakers... The distribution of seminars speakers shifted toward researchers of better quality... The geography of knowledge dissemination changed significantly as the average distance between host and speakers’ institutions increased by 32 percent and the share of seminars across borders also increased. Finally... the inequality in presentation opportunities manifested itself in inequality in citations.

Overall, these three papers provide a lot of food for thought on the state of and trends in economics seminars and the gender gap. Clearly, the seminar climate needs to improve from where it was in 2019. The trends over time suggest only small change in the gender distribution of seminar speakers over time (and it is possible that the difficult climate for female seminar speakers may contribute to that), but the recent shift to online seminar series may be a factor in reducing the gender gap in seminar speakers. However, these three papers also leave a lot of questions unanswered, including whether the change in gender composition of speakers as a result of the pandemic has been sustained as most seminar series moved to online or hybrid formats, whether the change in gender composition was matched by a change in composition in favour of URM speakers, and importantly, if the seminar climate is a problem, how it can best be improved. Hopefully, the increasing research attention in this area will provide us with some answers to these questions soon.

[HT: Marginal Revolution for the Dupas et al. paper, and David McKenzie at the Development Impact Blog for the Biermann paper]

Read more:

Saturday, 18 December 2021

Women are more competitive when they can be more pro-social

There is an array of research that demonstrates that men are more competitive than women (and boys are more competitive than girls; e.g. see here). This effect has been amply demonstrated in laboratory experiments for example (e.g. see here). In a recent article published in the journal PLoS ONE (open access, with a non-technical summary on The Conversation), Alessandra Cassar (University of San Francisco) and Mary Rigdon (University of Arizona) investigated whether the gender gap in competitiveness in laboratory contexts arises from the payoff structure of those experiments. Cassar and Rigdon argue that:

...women may be just as competitive as men, if the incentives involved reflect the social environment.

In their experiment, which was undertaken with 238 undergraduate students at Chapman University; University of California, Santa Cruz; and Simon Fraser University, Cassar and Rigdon:

...focus on one such scenario in which the prize for winning a tournament includes a social dimension: Winners have the option to share some of the prize with one of the losers. This prosocial option, known to the participants ahead of the competition, may appeal to women who are motivated to gain control of the distribution of resources or to repair social connections post-competition.

Research participants were:

...randomly assigned to one of two treatments: Baseline or Dictator. Each treatment consists of three rounds of a real effort task, the matrix search, under varying payment schemes: a piece rate per correct answer (round 1), a mandatory tournament where all subjects experience the competitive environment (round 2), and a choice between being paid according the piece-rate scheme or the tournament scheme (round 3).

Cassar and Rigdon then look at differences in the choices made in Round 3 of the experiment, between those who won were able to share some of the proceeds of the task with others (the 'Dictator' treatment, named after the 'Dictator Game'), and those who didn't get the option to share (the 'Baseline' treatment). The outcome is well-summarised by Figure 1 in the paper (which shows the proportion of male and female participants who chose the tournament rather than the piece rate in Round 3 of the experiment):

Women are clearly more competitive (more willing to choose the tournament rather than the piece rate) in the Dictator treatment than in the Baseline treatment. This difference is highly statistically significant, and in further analysis:

...controlling for these differences in individual abilities (using performance in the mandatory round 2 tournament), risk preferences, and beliefs explains some of the gender gap in competitiveness, but leaves the interaction effect of gender and treatment reported in model 1 largely unchanged and equally significant...

So, changing the context of the reward scheme can change the incentives such that women are just as competitive as men. The question now is, how can this be translated to real-world contexts (such as tournaments in the labour market or in education)?

[HT: The Conversation]

Read more:

Friday, 10 December 2021

Gender differences in answering multiple choice in a high-stakes test

Back in October, I posted about gender differences in multiple choice answering, which was based on research that used data from the PISA survey of high school children. The results demonstrated that male students perform better than female students, which is a common feature of multiple-choice tests (see the links at the bottom of this post for more). The research also provided some suggestive evidence that confidence in answering was important, along with stereotype threat, as I noted in this 2019 post. Related to confidence, part of the difference in performance between male and female students depends on differences in the propensity to leave some questions blank. Students that are less confident (who are relatively more likely to be female students than male students) are less likely to leave an answer blank.

Now, PISA is a 'low-stakes' test, in the sense that the results don't matter for the students at all. So, there is no negative consequence to leaving a question unanswered. That isn't the case in a high stakes examination, where guessing may come with a positive net payoff (if there is no penalty for a wrong answer), or may offer no net advantage (if there is a penalty). In that case, students may differ in whether they leave questions blank, depending on their confidence and their degree of risk aversion.

The propensity of male and female students to leave questions blank is investigated in this recent article by Perihan Saygin and Ann Atwater (both University of Florida), published in the journal Economics of Education Review (sorry, I don't see an ungated version online). They use data from the Turkish OSS, which is the main college admissions examination. This examination has an interesting feature that it has several sections, and those sections have different weights for different students. As Saygin and Atwater explain:

The Turkish high school curriculum features high school students being split into tracks of their choosing in their second year. The tracks are Science-Mathematics (Quantitative), Literature-Mathematics (Equally Weighted), Social Science-Literature (Qualitative), Foreign Languages, and the Arts. These fields line up with the topics covered on the ¨OSS. It has a total of four core sections: social science (which covers history, geography, and philosophy), science (biology, chemistry, and physics), mathematics, and literature... For each subject, the exam has a lesser difficulty section and a higher difficulty one. Thus, there are 8 sections in total on the ¨OSS...

The section weights then depend on which track the students were in during high school. So, a student in the quantitative track will have a higher weighting on the maths sections, but a lower weight on the social science section, whereas a student in the social science-literature track has the opposite. This allows Saygin and Atwater to investigate how differences in the stakes associated with particular sections affects the difference in propensity to leave questions blank between male and female students. Based on a sample of 1792 randomly selected OSS participants taking the OSS examination for the first time, they find that:

...not only is the gender difference in tendency to leave questions blank largest on sections that cover mathematics, but that this difference is only significant for the test takers on the track that places the most emphasis on these sections... We also provide evidence that this gap is larger on more difficult sections of a given subject despite these sections being weighted equally to the lower difficulty sections in score calculations.

Saygin and Atwater argue in a number of places in their paper that their results demonstrate that risk aversion isn't playing a role. However, when you see the results summarised as they are above, it is hard to draw any other conclusion. If a student is worried about the risk associated with guessing, then that risk is highest on the sections of the examination where the stakes are highest. Unfortunately, Saygin and Atwater don't have any measure of risk aversion. They do have measures of confidence, and looking at those they find that:

...a positive self-assessment on a given subject is related to skipping behavior on that subject and explains part of the gender differences in tendency to leave questions blank. In addition to this, we provide evidence for gender differences in self-assessment in a given subject conditional on the performance on the corresponding test section. Male test-takers are more likely to report a positive self-assessment than their test performance would suggest in math, science, and social science while this gender difference is inverted in literature. This variation in reported self-assessment across subjects matches the pattern of the observed gender differences in tendency to leave questions blank across subjects.

So, overconfidence explains at least some of the difference in leaving questions blank. The inability to convincingly eliminate risk aversion, or the interaction between risk aversion and confidence, leaves the question of the mechanisms underlying this behaviour still open.

Also, the sample that Saygin and Atwater use is somewhat idiosyncratic. The 1792 final sample is drawn from a larger random sample of nearly 10,000 OSS students. As far as I can tell, the smaller sample arises when they eliminate students that are attempting the OSS examination for the second or subsequent time. That suggests that around 80 percent of students in the OSS make multiple attempts before they get an examination ranking that they are satisfied will gain them entry into a good Turkish university. Clearly, there is a risk of selection bias in the sample in this research that is not accounted for.

Anyway, this paper modestly contributes to the idea that confidence contributes to the gender difference in multiple choice examination performance. However, we still need more research to better understand this topic.

Read more:

Thursday, 9 December 2021

Is this increasing gender equity, or gender inequity?

A post at the Dangerous Economist pointed me to this 2020 Medium article by Koen Smets (which is worth reading in its entirety):

Motor insurance in Europe forms a very interesting case study. Traditionally, insurers charged women less, because they tend to be safer drivers, and hence make fewer and smaller claims. Unlike life expectancy, the factors determining the risk here are much more linked to individual choice and behaviour. Since 2012, an EU directive forbids insurers to use gender as an element in the calculation of the premium. So, in Q4 2011, men paid on average 17% more than the overall average premium, while women paid 20% less. In 2018, that difference with the overall average premium had shrunk to a 5% uplift for male drivers, and a 6% discount for female drivers. (The residual difference stems from the fact that men tend to drive more miles per year, in more powerful cars.) So, relatively speaking, the ‘gender equity’ intervention has increased the average premium for the lower-risk women by more than 17%, while it has cut it by just under 10% for the higher-risk men. Is this an improvement? It’s not so sure. [sic]

Insurers charge higher motor vehicle insurance premiums to male drivers, because male drivers cost the insurers more. Male drivers tend to drive more kilometres, and have more severe accidents (see also here). It is reasonable for insurers to charge a higher premium to a more costly segment of the population, especially where those drivers are more costly because of their own behaviour (men could be lower cost to insurers, if they drove differently).

Insurance is subject to an asymmetric information problem that economists refer to as adverse selection. The uninformed party (the insurer) cannot easily tell drivers with 'good' attributes (low-risk drivers) apart from drivers with 'bad' attributes (high-risk drivers). To minimise the risk to themselves of engaging in an unfavourable market transaction, it makes sense for the insurer to assume that every driver is high risk. This leads to a pooling equilibrium - low-risk drivers are grouped together with the high-risk drivers and all drivers pay the same premium, because they can't easily differentiate themselves.

Since the premium is based on drivers of all risks on average, many low-risk drivers will find the cost of insurance to be too high. They will drop out of the market. The average risk of the remaining pool of insured drivers will increase, so the insurer will need to increase the insurance premium. Medium-risk drivers may then find the premiums too high, and drop out of the market. So, insurers raise premiums again. And so on, until the market fails because only the riskiest drivers would be left, and the insurer surely doesn't want to insure them!

One way of avoiding this adverse selection problem is for insurers to try to reveal how risky a driver each insurance applicant is. That way, we would have a separating equilibrium, where high-risk drivers pay higher premiums, and low-risk driver pay lower premiums. When the uninformed party tries to reveal private information (like how risky a driver an insurance applicant is), we refer to this as screening. In this case, the insurer uses various characteristics of the insurance applicant to estimate how risky they are likely to be. These characteristics might include their insurance history, past driving behaviour, the type of car they are insuring, and their demographic characteristics (including age and gender).

By imposing a law that equalises insurance premiums for men and women, the government is essentially telling insurers that they can no longer use gender as a screening tool for determining which drivers are higher risk. In the absence of that information, the insurer is a bit less informed than before. It moves things back towards the pooling equilibrium (but not all the way, because insurers still know the other details of the applicants). At the margin, insurers will now tend to over-estimate the risk of female drivers, and under-estimate the risk of male drivers. The consequence of this is that the premium for riskier male drivers becomes lower than it would have been without the law, and the premium for female drivers becomes higher than it would have been without the law. Essentially, relatively safe female drivers are cross-subsidising relatively riskier male drivers.

Wait! Wouldn't the intention of the EU directive have been to increase equity between female and male drivers? If you have a policy that equalises insurance premiums for male and female drivers, but in so doing makes male drivers better off and female drivers worse off, is that actually increasing gender equity, or decreasing gender equity? It would be interesting to know whether the policy makers had thought about this at all.

Read more: