Showing posts with label Gender bias. Show all posts
Showing posts with label Gender bias. Show all posts

Thursday, 11 December 2025

Do men and women pitch science proposals differently, and does it matter for funding outcomes?

If male academics and female academics write academic papers and grant proposals differently, does that lead to different outcomes by gender? Past studies have worried about whether grant funding decisions are affected by gender bias (see here, for example), and differences in writing style may contribute to that. However, the article I discussed in this post from earlier this year concluded that there was little evidence of bias in grant funding, at least since 2000 in the US.

Nevertheless, I thought it would be interesting to read this 2020 article by Julian Kolev (Southern Methodist University), Yuly Fuentes-Medel, and Fiona Murray (both MIT), published in the journal AEA Papers and Proceedings (ungated here), because it not only looks at the grant funding decisions, but also at writing style. Kolev et al. focus on grant applications submitted to the Bill and Melinda Gates Foundation and the National Institutes of Health (NIH) over the period from 2008 to 2016. The sample includes 6931 Gates Foundation applications and 12,589 NIH applications.

Kolev et al. first subject the applications to textual analysis to the abstract of each application, evaluating the positivity of the text (and the extent to which the word "novel" is used), the readability (using the Flesch reading ease score), concreteness of the language (as opposed to abstractness), and three measures of how narrow or broad the abstract is. In this textual analysis, they find that:

...female applicants are less likely to present their research using positive vocabulary, they are more likely to write with high readability, and they prefer concrete language. Moving to our final three measures, we find an interesting dichotomy: even as female applicants use fewer broad words and more narrow words in their abstracts, we find that their research is characterized by lower MeSH concentrations, meaning that they cover a wider range of medical subjects in their work, at least within the NIH sample. Effect sizes are relatively small: the impact of gender ranges from approximately 0.04 to 0.08 standard deviations for our significant effects.

So, there are small but statistically significant differences in writing style between male academics and female academics in these funding proposal abstracts. Does that translate into differences in outcome? Kolev et al. test for whether the measures of writing style correlate with funding outcome, while controlling for:

...calendar time and application topic fixed effects, controls for total word count and the count of relevant words for dictionary-based metrics, and applicant publication history and gender.

In this analysis, Kolev et al. find that:

For Gates applicants, high levels of concreteness tend to improve the odds of funding; by contrast, at the NIH, we find a strong positive impact for MeSH concentration and marginal effects for both broad and narrow words.

So, the evidence is weak that writing style matters, or that writing style differences between the genders affect the success of funding applications. However, not so fast. There is a key problem with this analysis. If you read the quote above about the control variables in this second analysis, you may note that they control for gender. That might sound sensible, but if you're wanting to evaluate whether writing style differences between the genders affect funding outcomes, you don't want to control for both writing style and gender. What Kolev et al. have actually tested is whether writing style differences within each gender affect funding application outcomes, finding that they don't. As an example, their analysis doesn't answer the question of whether readability differences between male and female academics affect funding outcomes, it answers the question of whether readability differences matter overall (which it appears they don't), controlling for the average difference in funding outcomes between men and women. Those are quite different questions.

In other words, we probably want to know whether style mediates the effect of gender on funding outcomes, but this analysis doesn't do that. Instead, they should either run the analysis with gender as the main explanatory variable, then add the style variables and see if the coefficient on the gender variable shrinks, or run the analysis with interactions between gender and the style variables.

The results are, on the one hand, surprising. Quality of writing should matter. Stylistic differences should matter less. However, the quality of the proposed research should matter even more than the quality or style of writing. And this study wasn't even evaluating the quality or style of all of the writing (or the quality of the proposed research), only the style of writing in the abstract for the proposal. That, along with the issue with the second analysis above, make this paper of limited use for understanding whether there is a gender difference in funding outcomes (and, if there is, whether writing style differences contribute to the difference). The difference in writing style is an interesting result in itself, but we need to know more.

Read more:

Sunday, 22 June 2025

What the persuasiveness of economics experts and signalling tell us about the gender gap in economics

Are the opinions of expert economists persuasive? The answer to that question will no doubt depend on who you ask. And at least some would answer 'yes, but I wish they weren't'. Does the gender of the economist matter for how persuasive their opinion is? That is a much more difficult question to answer. However, this recent article by Hans Sievertsen and Sarah Smith (both University of Bristol), published in the Journal of Economic Behavior and Organization (open access, but see this paywalled FT article as well), provides us with a starting point. Sievertsen and Smith:

...use an information provision experiment... to test whether the opinion about a topical policy issue expressed by a senior female economist is more, or less, persuasive than the same opinion expressed by a senior male economist. We run the same experiment twice – the first time, members of the public are shown credentials of expertise, the second time they are not.

Specifically, Sievertsen and Smith survey people in the US, asking their opinion on a range of issues (rated on a five-point scale from 'strongly disagree' to 'strongly agree'). Alongside a statement of the issue, each research participant was provided with the opinion of an economist, drawn from the US Economic Experts Panel run out of the University of Chicago. By comparing the results from the survey with those of an earlier survey of people who didn't get to see an economist's opinion, the persuasiveness of the economist whose opinion is provided (on each issue) can be evaluated.

The ten issues are based on the following statements:

1. Use of artificial intelligence over the next ten years will lead to a substantial increase in the growth rates of real per capita income in the US and Western Europe over the subsequent two decades.

2. There needs to be more government regulation around Twitter’s content moderation and personal data protection.

3. It would serve the US economy well to make it unlawful for companies with revenues over $1 billion to offer goods or services for sale at an excessive price during an exceptional market shock. (Price Gouging)

4. Efforts to achieve the goal of reaching net-zero emissions of greenhouse gases by 2050 will be a major drag on global economic growth.

5. Given the centrality of semiconductors to the manufacturing of many products, securing reliable supplies should be a key strategic objective of national policy.

6. A significant factor behind today’s higher US inflation is dominant corporations in uncompetitive markets taking advantage of their market power to raise prices. (Greedflation)

7. Financial regulators in the US and Europe lack the tools and authority to deter runs on banks by uninsured depositors.

8. When economic policy-makers are unable to commit credibly in advance to a specific decision rule, they will often follow a poor policy trajectory.

9. A windfall tax on the profits of large oil companies ‚ with the revenue rebated to households‚ would provide an efficient means to protect the average US household.

10. A ban on advertising junk foods (those that are high in sugar, salt, and fat) would be an effective policy to reduce child obesity.

The economists' actual views on each of the statements is not important for the research question (although the article does tell you, and you can probably guess for some of them what the average opinion of the economists is). Sievertsen and Smith first evaluate how persuasive economists are in general, finding that:

On average, a one-point change on the Likert scale in expert opinion is associated with a 0.17 point change in public opinion... Expert opinions have no effect on public opinions about Greedflation, while there are stronger effects for Price Gouging, Financial Regulation and Economic Policy. There is some support for the argument... that persuaders are more effective when receivers are less certain: The degree of persuasiveness is weaker on issues where baseline public opinion is more certain... The degree of persuasiveness is also stronger on issues where there is less distance between sub-panel expert opinion and baseline public opinion... this suggests that experts may be perceived as less credible when their views are further out of line with those of the general public.

So, the opinions of economic experts are convincing (somewhat). The last point is interesting though - people are most convinced when the economist's views are similar to those of the general public. This suggests some confirmation bias - people believe the experts more when the experts agree with them! It is also interesting who is persuaded most:

The degree of persuasiveness is greater for men [p = 0.002] and for non-whites [p = 0.000]. It is also greater for those with a degree [p = 0.000] and for those with higher self-reported economics knowledge [p = 0.000]. Those who identify as Republicans are also more persuaded by economists’ opinions than Democrats/Independents [p = 0.030].

Those who are more educated, and who claim to have more economics knowledge, are more persuaded by expert economists. More educated people likely give more credence to the views of other educated people, while those who claim to know more economics are more likely to modify their views to fit with those of economics experts.

What about gender differences in persuasion? Sievertsen and Smith switch the analysis to evaluating whether the opinion of the research participants exactly matches that of the expert whose opinion they are provided, and find that:

...members of the public are 1.1 percentage points more likely to match with the opinion of a female expert than with the same opinion expressed by a male economist.

Sievertsen and Smith find similar results when evaluating the distance between the opinion of the research participant and the expert (the distance is smaller for female experts). The effect (1.1 percentage points, against a match rate for male experts of 33.5 percent) seems quite small, but Sievertsen and Smith note that some matches would happen purely by chance, and after accounting for that:

...the degree of persuasiveness of female expert opinions is around 20 per cent higher {than male experts]...

Why are female economics experts so persuasive? Here's where things get interesting. Sievertsen and Smith run their survey a second time. The first time around, research participants were told the name and institutional affiliation of the expert economist (as well as shown their profile photo). In the second survey, research participants were only told the name of the expert economist (and shown the photo). In that second survey:

The overall effect of

female expert on the probability of matching opinion drops from 0.011 in the main experiment to 0.0002 in the follow-up.

The extra persuasiveness of female experts (over and above male experts) disappears! Sievertsen and Smith conclude that:

...in the first experiment, credentials provided an information signal that favoured senior female experts. Removing that signal in the follow-up experiment eliminates the gender difference.

It is worth explaining that last result in a bit more detail, because what it really shows is that the general public recognises the gender bias in top economics institutions.

The quality of a purported economics expert is private information. it is known to the expert themselves, but not known to the public. This is a case of asymmetric information. The expert is the informed party, and the general public is the uninformed party. Since the general public cannot tell high-quality and low-quality experts apart, they might assume that all experts are low quality. How can a high-quality expert instead convince the general public of their high quality? They must credibly reveal their quality to the public - this is called signalling.

For a signal to be effective, it must meet two conditions: (1) it must be costly; and (2) it must be costly in such a way that those with low quality attributes would not want to attempt the signal. Getting a tenured position is signal of high quality for experts. It is costly to get such a position, and it is costly in a way that low-quality experts wouldn't want to attempt it (because they would be unsuccessful in getting tenure anyway). So, a position at a top university is a signal of quality.

Now, why would this signal be even more effective for female economists? If getting a tenured position at a top institution is even more costly for female economists than for male economists, then the quality of the signal is higher for women than for men. And therefore, the general public would be even more believing of the signal for female economists than for male economists. The gender gap in economics is pervasive (see this post, and the posts linked at the bottom of that post). What is interesting is that this gender gap is so well established that the general public is acting on it!

Sievertsen and Smith finish by pointing out a bit of a paradox:

...if senior female economists have greater credibility in the eyes of the public, then why are they less confident in giving their opinion. This remains an open question.

Indeed. Hopefully, at the margin, this research will convince senior female economists to use their persuasiveness for the good of all (where 'good' is defined as persuaded more people to believe in economists' opinions).

Tuesday, 22 April 2025

The gender gaps in academia may not arise entirely from gender biases

I've written a lot of posts about the gender gap in academia, in economics and in other (mostly STEM) disciplines (see the list at the end of this post). However, this 2023 review article by Stephen Ceci (Cornell University), Shulamit Kahn (Boston University), and Wendy Williams (Cornell University), published in the journal Psychological Science in the Public Interest (open access), suggests that the gender gap may not be a substantial as previously believed (including by me). This article is quite credible, having arisen as an 'adversarial collaboration', meaning a collaboration between researchers who previously disagreed on the key conclusions from the literature. As Ceci et al. explain:

This article represents more than 4.5 years of effort by its three authors. By the time readers finish it, some may assume that the authors were in agreement about the nature and prevalence of gender bias from the start. However, this is definitely not the case. Rather, we are collegial adversaries who, during the 4.5 years that we worked on this article, continually challenged each other, modified or deleted text that we disagreed with, and often pushed the article in different directions...

Kahn has a long history of revealing gender inequities in her field of economics, and her work runs counter to Ceci and Williams’s claims of gender fairness.

Ceci et al. focus on seven questions of relevance to understanding the gender gap in academia:

In this article, we comprehensively examine evidence in six key evaluation contexts: (a) Are similarly accomplished women and men treated differently by academic hiring committees? (b) Are grant reviewers biased against female PIs? (c) Are journal reviewers biased against female authors? (d) Are recommendation-letter writers biased against female applicants for tenure-track positions? (e) Are faculty salaries biased against women? And, (f) are student teaching evaluations biased against female instructors? Claims of gender bias are omnipresent in all six of these domains... (We also review the literature in a seventh context, gender differences in publication rates, because publishing productivity can moderate evaluation in most of these six contexts.)

Ceci et al. focus on research published since 2000, which is more likely to represent the 'current' state of academia (although, arguably, they should weight more heavily more recent studies, which they don't). They also distinguish between:

...the most mathematically intensive fields—geosciences, engineering, economics, mathematics/computer science, and physical science (GEMP)—and less math-intensive fields—life sciences, psychology, and social sciences (LPS).

It is generally claimed that gender gaps are more prevalent in the GEMP fields than in the LPS fields (and a quick read through the links at the end of this post would suggest that is definitely true of economics, as one example).

The article is very thorough and has a lot of detail related to each of those contexts, so I'm just going to hit the headlines in relation to each of the seven research questions. If you are interested in any particular finding, the article is open access, so you can easily look at their review and unpack the details there. In relation to the first research question (are similarly accomplished women and men treated differently by academic hiring committees?), Ceci et al. conclude that:

The vast majority of findings—from (a) synthetic cohort analysis, (b) institutional hiring records, and (c) experiments—indicate that women are less likely than men to apply for tenure-track jobs, but when they do apply, they receive offers at an equal or higher rate than men do.

So, the news is both good and bad. On the positive side, there is little evidence of bias. However, where a gender gap in hiring persists, and is not because of bias in hiring, the gender gap must arise from differences in the rates of applying for academic positions between men and women. Indeed, Ceci et al. note that:

...women are more likely than men to give up their initial aspirations to become tenure-track professors while in graduate school, a finding primarily true of women with children or contemplating children. Undoubtedly, broad systemic factors are partly responsible, along with biological factors, for these women not applying for tenure-track positions.

That still suggests that there is further work to do (both in terms of research and in terms of addressing the problem), but this particular article skirts around that issue because it focuses on the gender biases for academics, and not the graduate student experience. In relation to the second research question (are grant reviewers biased against female principal investigators (PIs)?), Ceci et al. conclude that:

...pre-2006 evidence suggests that although some agencies evaluated men and women differently, on average they did not.

Ceci et al. then conduct their own meta-analysis of the literature since 2000 (including 39 studies), and conclude that:

Taken together, both the analytic dissection and our meta-analyses appear not to support the claim that the grant peer-review process has been rigged against women PIs during the past 20 years in the United States. This is particularly true when analyses controlled for PIs’ research productivity...

On the third research question (are journal reviewers biased against female authors?), Ceci et al. conduct both a review and meta-analysis and conclude that:

...overall, our meta-analyses and our dissection of key studies revealed no evidence of systematic bias against female authors, notwithstanding claims to the contrary.

For the fourth research question (are recommendation-letter writers biased against female applicants for tenure-track positions?), Ceci et al. conclude that:

On the basis of our analysis of the nine studies in this domain, we conclude that no persuasive evidence exists for the claim of antifemale bias in academic letters of recommendation.

In relation to the fifth research question (are faculty salaries biased against women?), Ceci et al. conclude that:

...the evidence supports the claim that women are paid less than men in tenure-track academia, although the magnitude of the gap is much smaller (60%–80% smaller) than often claimed in executive summaries and headlines, and in some situations has disappeared.

Ceci et al. also dig a bit deeper on the salary differences, noting that:

Some of the unexplained gender salary gap may be due to implicit bias (although this seems unlikely in biology, where starting salaries are higher for women), and some of it may be due to differences in willingness to negotiate and solicit outside offers... some of the remaining pay gap may be due to women’s work discontinuities for family leave... or to a desire to keep jobs flexible... Finally, some of the relatively small remaining pay gap may be due to women’s lower likelihood of negotiating higher salaries or their lower likelihood of pursuing more lucrative job offers. The lower likelihood of negotiating higher salaries may itself be due to bias... Without specific data on family leaves, past employment, and job pursuit, it is impossible to know how much, if any, of the less than 4% unexplained pay gap is attributable to bias.

The salary gap of 4% is small, but it is not zero. The fact that much of the gap can be explained by the factors outlined above would accord with research by 2023 Nobel Prize winner Claudia Goldin (whose work they cite, among others). However, that leaves open the question of how large the salary gap would be, after accounting for work discontinuities, preferences for flexibility, and negotiation? Again, a research question to be addressed in the future.

On the sixth research question (are student teaching evaluations biased against female instructors?), Ceci et al. conclude that:

...the evidence supports the claim that female instructors are penalized for being women, independent of the content and delivery of their lectures and independent of students’ actual learning. The effect sizes we calculated indicate penalties for women that ranged between small and moderately large (ds = 0.10–0.50). So, unlike the domains in which we were able to unequivocally reject claims of widespread gender bias, in this domain, we conclude that there is gender bias.

This is consistent with many research findings on student evaluations of teaching, including studies I have blogged about before (most recently here, and see the other links at the end of that post for more). Unfortunately, this seems to be a pervasive finding across all teaching contexts, and it doesn't appear to be getting any better.

Finally, in relation to the additional research context (gender differences in research productivity), Ceci et al. conclude that male researchers do have higher research productivity (more publications), and that:

...gender productivity differences are smallest in GEMP fields (with the exception of economics) and are largest (and possibly growing) in biology, psychology, and economics.

The overall takeaway from this work is that there are still gender gaps in academia, but that many of the gaps, or claimed gaps, don't seem to arise from gender bias. At least, that's what we should conclude from the research to date. That is true of all of the domains except salaries and teaching evaluations. None of this means that we should conclude that all is rosy for female academics (unlike the title of this post), and that is certainly not the case across all fields. Indeed, some of the changes in recent years might actually make things worse for female academics before they get better (as noted in yesterday's post). We still have some way to go, especially in the more technical fields, including STEM and economics.

[HT: Marginal Revolution, back in 2023]

Read more:

Monday, 21 April 2025

Co-authorship in economics in the aftermath of #MeToo

The #MeToo movement was a necessary corrective action recognising decades of toxic behaviour across many occupations. Economics was not immune (for example, see here or here). However, could the #MeToo movement have had an unintended consequence on the careers of female economists? If having a female co-author increases the chances of a male economist being called out for even minor indiscretions, does this meaningfully raise the cost of having female co-authors? And if the cost of having female co-authors meaningfully increases, we would expect to see fewer male-female collaborations (especially where the male economist is more senior).

That is the topic addressed in this new article by Noriko Amano-Patiño, Elisa Faraglia, and Chryssi Giannitsarou (all Cambridge University), published in the journal European Economic Review (open access). The use data on co-authorships in the nearly 27,000 working papers published in the NBER and CEPR working paper series between January 2004 and December 2020. They first note that:

The MeToo movement’s impact on the economics profession may have fostered a more respectful research environment, increased scrutiny of existing practices, and promoted greater diversity and inclusivity within the community. Conversely, it could have induced a chilling effect on collaborations, potentially causing researchers to become more hesitant in forming partnerships outside their established networks due to heightened concerns about trust and reputational risk.

Although the #MeToo movement started in 2017, Amano-Patiño et al. use the second quarter of 2018 as the effective date for the onset within the economics profession (that dates to the fallout arising from a series of studies, including a particularly notable study by Alice Wu, which I blogged about here). However, Amano-Patiño et al. vary the effective start date and find little difference in their results. So, comparing papers written before and after 2018, and controlling for a variety of author characteristics, Amano-Patiño et al. find that there was:

...a rise in the proportion of women coauthors for men, both overall and within junior and senior subgroups. Conversely, we find a decrease in the proportion of women coauthors for women, both overall and within corresponding seniority levels. Using a back-of-the-envelope calculation, these increases in mixed-gender collaborations, translate to an estimated 12.3% increase in women coauthors per 100 men-authored papers.

That seems to go against what we might expect, if the cost of having female co-authors has increased for male economists after #MeToo. However:

...we estimate decreases in the proportion of senior coauthors (especially senior women) for juniors, and symmetrically, in the proportion of junior coauthors (particularly junior women) for seniors. The decreases in collaborations between senior and junior economists we quantify, suggest a 3.0% decrease in the share of senior authors collaborating with junior coauthors.

 Amano-Patiño et al. interpret this as showing that:

...post-MeToo, authors have increasingly sorted their collaborations by seniority rather than by gender.

What might explain these findings? Researchers who are worried that they might get called out by female co-authors might respond by reducing their collaborations with female co-authors generally, as I noted at the start of this post. Or, they might reduce their collaborations with new co-authors, who they have not developed trust with, while continuing to collaborate with more senior authors that they trust. This is also consistent with Amano-Patiño et al.'s further results, where they note that:

...we find evidence of a general chilling effect on the expansion of economists’ professional networks. We estimate decreases in the share of new coauthors across all seniorities and genders, the share of new senior coauthors for juniors, and the share of new junior coauthors for seniors. Our estimates translate into 5.4% fewer new coauthorships per 100 papers. This trend is primarily driven by a substantial decrease in new coauthorships between senior and junior authors: for seniors, the share of new junior coauthors has dropped by 18.4%, with a particularly sharp 48% decrease in their share of new junior women coauthors.

Amano-Patiño et al. interpret their results as bad news, noting that if the results can be interpreted as causal:

First, authors may have prioritised increasing gender diversity in their collaborations. Second, senior authors have increasingly relied on their existing collaboration networks rather than forming new coauthorships. The latter trend, if persistent, could have long-lasting consequences for the career development of women economists and potentially exacerbate the already ‘leaky’ pipeline in the profession.

 Amano-Patiño et al. stop short of noting that this is a substantial negative unintended consequence of the #MeToo movement in economics. Although the environment for female economists may be improving, at least one aspect, being the opportunities for collaboration and mentoring from senior economists, appears to be declining. And that will be a difficult problem to address. Indeed, Amano-Patiño et al. aren't able to offer any concrete steps that could be implemented to solve this issue, concluding with some more general statements:

These results underscore the urgent need for sustained efforts to cultivate a supportive ecosystem and dismantle systemic barriers hindering the advancement of women and junior economists in the field. The economics profession must proactively continue to foster a safe, inclusive environment by evaluating, monitoring, and educating on relevant issues.

Sadly, I also can't offer anything concrete, and only hope that the current desire for change within the profession will ultimately lead to greater opportunities for female economists overall.

Sunday, 9 June 2024

The 'mighty girl effect' may only kick in when daughters reach school age

You may have heard of the 'mighty girl effect' (also called the 'eldest daughter effect') - the idea that fathers whose eldest child is a daughter are less likely to support traditional gender norms, and have more progressive views. There are several studies that support the existence of this effect (see here or here for examples). However, less studied is when this effect occurs. Does the birth of a daughter have an immediate impact, or does it take time for fathers to change their views? And if it takes time for father's views to change, how long does it take?

This 2019 article by Mireia Borrell-Porta, Joan Costa-Font, and Julia Philipp (all London School of Economics and Political Science), published in the journal Oxford Economic Papers (open access), provides an initial answer. They used data from the British Household Panel Survey waves between 1991 and 2012, a sample of nearly 28,000 observations of over 11,000 parents (about 44 percent men). Their key measure was agreement with the statement "a husband’s job is to earn money; a wife’s job is to look after the home and family", initially measured on a five-point scale ranging from "strongly agree" to "strongly disagree", but in most analyses they use a binary version of the variable, set equal to one where the respondent strongly agreed, agreed, or neither agreed nor disagreed with the statement (thereby demonstrating some level of support for traditional gender roles).

Borrell-Porta et al. take advantage of the fact that the gender of a child is essentially random, and compare parents with at least one daughter in the household with those with no daughters. They then extend that analysis (which is similar to previous research) to consider the age of the oldest daughter (in three categories: 0 to 5 years; 6 to 10 years; and 11 to 18 years). The first set of results are well summarised in the simplest analysis, presented in Figure 1 of the paper:

Fathers with daughters are less likely to support traditional gender roles, but the results are less clear for mothers. So far, nothing so different from earlier work in this area. In the full analysis, separating daughters by age, Borrell-Porta et al. find that:

...fathers’ probability to support traditional gender norms declines by approximately three percentage points (8% change) when parenting primary school-aged daughters and by four percentage points (11% change) when parenting secondary school-aged daughters. In contrast, the effect on mothers’ attitudes is smaller and generally not statistically significant.

All of that is based on self-reported responses to the question, so Borrell-Porta et al. look deeper for behavioural change, specifically looking at whether fathers of daughters are less likely to be in a couple that follows a 'male breadwinner norm' (with the father working, and the mother not working). They find that:

Parenting pre-school daughters is associated with a higher probability to behave traditionally. However, parenting primary and secondary school-age daughters is associated with a lower likelihood to follow a traditional male breadwinner norm in which the man works and the woman does not work, and this result holds both cross-sectionally and longitudinally. In terms of effect size, FEs estimates... indicate that parenting daughters aged six to 10 reduces the probability of a traditional gender division of work by seven percentage points, and parenting daughters aged 11 or older reduces that probability by five percentage points. Compared to the baseline probability of following a traditional norm for those without daughters of 20.3%, this is a sizeable reduction of 36% and 25%, respectively.

So, being father to a daughter not only appears to change attitudes, but also behaviour, but only when those daughters reach school-age. Borrell-Porta et al. note that these results are consistent with exposure theory (which says that men develop or change their understanding of women's place in society when exposed to situations that make them more sympathetic - something that mothers would have already experienced, but fathers experience through their daughters), as well as identity theory (which says that the child's wellbeing enters into the parent's utility function - that is, the parent feels better off when the child is better off). Borrell-Porta et al. aren't able to tease apart those possible mechanisms underlying the results.

I find these studies interesting, but I think they raise as many questions as they answer. Fathers have had daughters for millenia. If each generation of fathers became less likely to support 'traditional' gender norms than the previous generation, by the amounts that these studies find, then the 'traditional' gender norms would have disappeared long ago. What is really missing here is an answer to the question of, why now? Is there something about modern times that facilitates this change? Are there certain pre-conditions that need to be in place before daughters can have an appreciable impact on the attitudes and behaviours of their fathers? Do these results hold across cultures that are less progressive than the UK and the US (where the studies that I have seen have been based)? These are all questions that would be interesting to answer, and give us a better understanding of how daughters contribute to the breakdown of fathers' support for traditional gender norms.

Saturday, 20 April 2024

The gender of a doctor matters for medical evaulations

There is lots of evidence that there is gender bias in healthcare. This Medical News Today article summarises some examples and consequences. It seems plausible that at least some of the gender bias in healthcare arises when male doctors examine or treat female patients. A useful question to ask, then, is what would happen to bias if patients were examined by same-gender doctors?

That is essentially the research question underlying this recent article by Marika Cabral (University of Texas at Austin) and Marcus Dillender (Vanderbilt University), published in the journal American Economic Review (ungated earlier version here). Cabral and Dillender first outline the problem, being that:

...female patients, relative to male patients, receive less health care for similar medical conditions and are more likely to be told by providers that their symptoms are emotionally driven rather than arising from a physical impairment... Differences in doctors’ evaluations of medical issues for male and female patients may be a key factor contributing to observed differences in treatment. Beyond influencing the treatments patients receive, medical evaluations also impact benefit eligibility in social insurance programs. Recent evidence suggests there are large gender disparities in social insurance programs that rely on medical evaluations...

Cabral and Dillender make use of:

...comprehensive administrative data and random assignment of doctors to patients within the Texas workers’ compensation insurance system. Random assignment of doctors to patients occurs in this setting through the dispute resolution process. Insurers and injured workers may request independent medical evaluations to settle disputes over an injured worker’s impairment level... The random assignment of doctors to patients means that differences in assessments between male and female doctors stem from the doctors themselves rather than from differences in the types of patients assigned to doctors.

That last point is important. It is the random assignment of patients to doctors that means that the results from this study can be interpreted as causal evidence of the effect of doctor gender on patients' outcomes, and evaluate the difference in those outcomes between male and female patients. Essentially, this is a form of difference-in-differences analysis, looking at the difference in outcomes between male and female patients with a male doctor, and comparing that with the difference in outcomes between male and female patients with a female doctor.

The outcomes that Cabral and Dillender look at are whether the patient is evaluated as having a disability, and the amount of cash disability benefits they receive after the evaluation. Having controlled for patient characteristics such as the type of injury and the industry that the patient worked in, there should be no differences between male and female patients in either disability assessment or disability benefits, depending on whether they have a male or female doctor. Instead, Cabral and Dillender find that:

...patient-doctor gender match increases evaluated disability and subsequent cash disability benefits for female patients but has little impact on outcomes of male patients... Compared to differences among their male patient counterparts, female patients randomly assigned a female doctor rather than a male doctor are 3.1 percentage points more likely to be evaluated as having an ongoing disability and receive 8.6 percent more cash benefits on average, or $483 evaluated at the mean of $5,622. There is no analogous gender-match effect for male patients. We note the magnitude of these effects is sizable. The estimated 3.1 percentage point increase in the likelihood of being evaluated as disabled is nearly large enough to offset the entire observed gender gap in this outcome when male doctors evaluate claimants.

Cabral and Dillender then turn to explaining why this gender bias exists, and find that:

Controlling for available baseline patient information, the estimates indicate that female doctors evaluate female and male patients as similarly disabled while male doctors evaluate female patients as less disabled than male patients. While only suggestive, this evidence is consistent with male doctors evaluating female patients against a stricter standard than male patients and female doctors applying similar standards to male and female patients.

On that last point though, as Cabral and Dillender note in one of the footnotes in the paper, these results alone can't distinguish between whether it is male doctors who evaluate female patients to a higher standard, or female doctors who evaluate male patients to a lower standard. However, Cabral and Dillender report a range of survey evidence from a sample of over 1500 people that is consistent with the former, including:

...that women—relative to men—more often report having a negative experience where a doctor didn’t understand their concerns, had assumed something without asking, talked down to them, made them feel uncomfortable, or didn’t believe them. When asked about how a doctor’s gender influences the likelihood of having a positive interaction, women were much more likely than men to report an own-gender doctor would be more likely to treat them with respect, understand their concerns, believe them, provide needed testing and treatments, make them feel comfortable, and ask appropriate questions instead of making assumptions.

Cabral and Dillender also report on the intensity of preferences over doctor gender, showing that:

...48.5 percent of women are willing to pay an additional $5 copay to see an own-gender provider compared to only 29.3 percent of men—a 19.2 percentage point difference.

It would have been interesting if they had extended that analysis to an estimate of the female patients' average willingness-to-pay for having a female (rather than a male) doctor, but they didn't. Finally, Cabral and Dillender looked at the policy implications, noting that based on their results:

...increasing the share of independent medical evaluations performed by female doctors from 17 percent to 50 percent would cause a 0.88 percentage point increase in the share of female patients evaluated as disabled, closing approximately 41 percent of the gender gap conditional on observables among disputed claims.

Given that still less than half of medical school graduates in the US are female, there is a long way to go before we get to that point. For comparison, in New Zealand in 2019, over 58 percent of medical school graduates were female. I guess that is good news for New Zealand, in terms of reducing the gender bias in medical evaluations here.

Monday, 8 August 2022

Online teaching and gender bias in teaching evaluations

The gender bias in student evaluations of teaching is well established (see my most recent post on the topic here, or the links at the end of this post). However, during the pandemic teaching underwent a sudden and unexpected change to an online mode. It is reasonable to ask, has online teaching reduced gender bias in evaluations? On the one hand, we know that women faced additional difficulties in managing the transition to online work, and especially juggling work with unequal home and family responsibilities. This may have impacted on female teachers' ability to deliver teaching in the online mode effectively. On the other hand, at the risk of gross generalisation, women tend to have a different teaching style that favours connections over content, which may have better helped students with the transition to online learning. It is therefore unclear whether female teachers would be helped, or hurt, in terms of student evaluations of teaching, by the move to online teaching.

This new article by Sara Ayllón, published in the journal Economics of Education Review (ungated earlier version here), provides us with an initial answer. Ayllón uses data from the University of Girona in Spain. Using a difference-in-differences approach, she compares the difference in teaching evaluations between the first and second semester of the 2018/19 academic year, with the difference in evaluations between the first and second semester of the 2019/20 academic year. Since online teaching was enforced in Spain for much of the second semester of 2019/20, this comparison is intended to pick up the impact of online teaching on evaluations. In her baseline analysis, she finds that:

...when I use controls (student’s age, its square, gender, whether the student is repeating the course, and field of study) and robust standard errors clustered at the student level... teaching evaluations during the online semester were, on average, no different from those in previous semesters. Interestingly, though, separate regressions by gender of the lecturer indicate a different story... the online semester had, on average, no impact on the evaluation of male lecturers; but for female lecturers, the average evaluation score decreased by 0.063 points in the online semester compared to previous semesters (about 5.4% of a standard deviation). Thus, while the new teaching environment had, on average, no effect on men’s scores, it did negatively impact the scores received by women...

Ayllón then attempts to tease out the reasons underlying the negative impact of online teaching on female lecturers' student evaluations. She finds no difference in how students felt about how well the course materials were adapted to online learning between male and female lecturers. She also finds no robust difference in grades between students with female lecturers and students with male lecturers. And there is no difference in students' opinions on various aspects of lecturer performance. Ayllón notes that:

...the gendered difference in the teaching evaluation result of the online environment does not appear to be driven by (potentially more objective) aspects of the teacher’s performance. The bias creeps in when students evaluate overall performance...

Students couldn't have easily sorted themselves to have different lecturers, because they chose their programme of study at the start of the academic year. Nevertheless, Ayllón shows that the results hold when the sample is limited to compulsory courses. Finally, looking more deeply at the characteristics of lecturers and students, she finds that:

...the results are particularly negative for young female instructors without a permanent contract, and are strongly driven by male students and low achievers who - even before they know their final grade - retaliate against female instructors, but not against male teachers. The findings are most apparent in Social Sciences. Online teaching did not lead to any positive bias on the part of female students towards female instructors. Yet a considerable degree of discrimination in favour of male instructors is found among high-achieving students.

These results are not dissimilar to other results in the literature. However, rather than showing the underlying gender bias in teaching evaluations, this study shows that the gender bias is larger during online teaching. Unfortunately, as with most studies like this, the solution to the problem is unclear. With online teaching, female lecturers can't give their student evaluations a bump by giving students chocolate. However, these results, especially if confirmed in other studies, should make us even more cautious about using data from student evaluations of teaching in promotion or tenure or appointment decisions, unless we want to perpetuate gender bias.

Read more:

Wednesday, 30 March 2022

The bias against research on gender bias?

There is a large (and growing) literature on gender bias and the gender gap. Regular readers of this blog will no doubt have noted that it is a recurring theme. In fact, there is so much research on gender bias that it is hard to believe that there could be a bias against such research. Nevertheless, that is the conclusion of this 2018 article by Aleksandra Cislak (Nicolaus Copernicus University), Magdalena Formanowicz (University of Bern), and Tamar Saguy (Interdisciplinary Center Herzliya), published in the journal Scientometrics (open access).

Cislak et al. collated data on publications listed in PsycINFO and PsycARTICLES over the period from 2008 to 2015, on gender bias or racial bias. After removing irrelevant articles and duplicates, their analysis was based on slightly more than 1000 articles over that time. They then compared articles on gender bias with articles on racial bias, in terms of prestige. Prestige was measured in two ways: (1) by the impact factor of the journal they were published in, for the year of publication); and (2) by whether the research had been funded. They found that:

...research on gender bias was funded less often (B = - .20; SE = .09; p = .02) and published in lower Impact Factor journals (B = - .67; SE = .20; p = .001).

So, research on gender bias appears from this study to have attracted less prestige than research on racial bias. Cislak et al. take that as evidence that there is bias against research on gender bias. However, there is good reason to doubt their conclusion. It relies on an assumption that, in the absence of any bias, there would be no difference in the impact factor or funding between research on gender bias and research on racial bias. That assumption strikes me as difficult to support.

In order for there to be bias, the ideal experiment would be two otherwise identical groups of articles, one group on gender bias, and one group on racial bias, submitted to the same journals. That would hold constant the quality of the articles, the quality of the journals, general editorial policies and practices, authorship, authors' incentives (for writing long comprehensive articles, or shorter articles on sub-topics), and article context (other than gender or racial bias). Differences in acceptance rates between these two groups of articles might be taken as evidence of bias.

Instead, we have an observational study based on different articles published by different journals. It tells us almost nothing about the editorial process from submission to publication. We really have no idea if there observed difference arises because of bias, or because of differences in article quality or something else. In fact, the data could even be consistent with bias in favour of articles on gender bias, if the acceptance rate of submitted gender bias articles was higher than the acceptance rate of submitted racial bias articles of otherwise similar quality and other attributes. Maybe the bar for acceptance of an article on racial bias is set higher than the bar for acceptance of an article on gender bias? Cislak et al. engage in some hand-waving about the quality of the research being the same because of the use of "similar methods and paradigms". However, that's not a very convincing argument, and they don't actually control for research quality (noting whether articles use quantitative or qualitative methods is not a control for research quality).

Similarly, if the source of funding is not held constant between gender bias and racial bias, it tells us little about bias (unless we are considering bias in the availability of funding, which Cislak et al. probably argue they are getting at). Nevertheless, without knowing the rate of acceptance of funding applications for gender bias and racial bias, there is no reason to believe that more articles being funded is evidence of bias in either direction. 

In short, this research is unconvincing. Show me an audit study on this topic, and I'd likely give it more weight. But this observational research simply doesn't cut it.

Tuesday, 15 February 2022

Hardly the final words on student evaluations of teaching

I've written a few posts this week about student evaluations of teaching (see here and here), a few others in previous years. One of those earlier posts asked whether student evaluations of teaching are even measuring teaching quality. The research I cited there was a meta-analysis that suggested that there was no correlation between teaching quality (as measured by final grade or final exam mark or similar) and student evaluations of teaching. However, measuring teaching quality objectively using grades could be problematic if there is reverse causation (for example, teachers give higher grades in hopes of receiving better teaching evaluations). A better approach may be to use some measure of teacher value-added, such as the grade in subsequent classes (with different teachers), or grades in standardised tests (that the teacher doesn't grade themselves).

The former approach, based on teacher value-added, is the one adopted in this 2014 article by Michela Braga (Bocconi University), Marco Paccagnella (Bank of Italy), and Michele Pellizzari (University of Geneva), published in the journal Economics of Education Review (ungated earlier version here). They use data from students in the 1998/99 incoming cohort at Bocconi University, where the students were randomly allocated to teaching classes in all of their compulsory courses (which eliminates problems of selection bias). Looking at the effect of future student performance on current teaching evaluations, Braga et al. find that:

Our benchmark class effects are negatively associated with all the items that we consider, suggesting that teachers who are more effective in promoting future performance receive worse evaluations from their students. This relationship is statistically significant for all items (but logistics), and is of sizable magnitude. For example, a one-standard deviation increase in teacher effectiveness reduces the students’ evaluations of overall teaching quality by about 50% of a standard deviation. Such an effect could move a teacher who would otherwise receive a median evaluation down to the 31st percentile of the distribution.

Those results are consistent with the meta-analysis results, that teachers who do a better job of preparing students for their future studies receive worse teaching evaluations. However, when looking at exam performance in the current class, Braga et al. find that:

...the estimated coefficients turn positive and highly significant for all items (but workload). In other words, the teachers of classes that are associated with higher grades in their own exam receive better evaluations from their students. The magnitudes of these effects is smaller than those estimated for our benchmark measures: one standard deviation change in the contemporaneous teacher effect increases the evaluation of overall teaching quality by 24% of a standard deviation and the evaluation of lecturing clarity by 11%.

They interpret those results as showing that teachers who 'teach to the test' for the current semester receive better teaching evaluations. Braga et al. conclude, unsurprisingly, that:

Overall, our results cast serious doubts on the validity of students’ evaluations of professors as measures of teaching quality or effort.

Aside from being a measure of teaching quality or effort, perhaps student evaluations of teaching provide useful information that teachers use to improve? This 2020 article by Margaretha Buurman (Free University Amsterdam) and co-authors, published in the journal Labour Economics (ungated earlier version here), addresses that question using a field experiment. Specifically, from 2011 to 2013 Buurman et al.:

...set up a field experiment at a large Dutch school for intermediate vocational education. Student evaluations were introduced for all teachers in the form of an electronic questionnaire consisting of 19 items. We implemented a feedback treatment where a randomly chosen group of teachers received the outcomes of their students’ evaluations. The other group of teachers was evaluated as well but did not receive any personal feedback. We examine the effect of receiving feedback on student evaluations a year later...

They find that:

...receiving feedback has on average no effect on feedback scores of teachers a year later. We find a precisely estimated zero average treatment effect of 0.04 on a 5-point scale with a standard error of 0.05...

Buurman et al. suggest that this may be because they estimate the effect a year later, and that the evaluations feedback may have shorter run effects. I don't find that convincing. However, there were differences by gender:

Whereas male teachers hardly respond to feedback independent of the content, we find that female teachers’ student evaluation scores increase significantly after learning that their student evaluation score falls below their self-assessment score as well as when they learn their score is worse than that of their team. Moreover, in contrast to male teachers, female teachers adjust their self-assessment downwards after learning that students rate them less favorably than they rated themselves.

That should perhaps worry us, given the gender bias in evaluations. If it causes female teachers to expend additional effort in trying to improve their teaching evaluations to match those of male teachers, then they will be expending more effort on teaching than the male teachers will for the same outcome. In a university context, that would likely have a negative impact on female teachers' research productivity, with negative consequences for their career. This might be an intervention that is best avoided, unless the gender bias in student evaluations of teaching is first addressed (see yesterday's post for one idea).

As the title of this post suggests, this is hardly the final words on student evaluations of teaching. However, we need to understand what works best, what avoids (or minimises) biases against female or minority teachers, and how teachers can best use the outcomes of evaluations to improve their teaching.

Read more:

Monday, 14 February 2022

Can a simple intervention reduce gender bias in student evaluations of teaching?

Following on from yesterday's post, which discussed research demonstrating bias against female teachers in student evaluations of teaching [SETs] (see also this post, although the meta-analysis in yesterday's post suggested that the only social science to have such a bias is economics), it is reasonable to wonder how the bias can be addressed. Would simply drawing students' attention to the bias be enough, or do we need to moderate student evaluations in some way?

The question of whether a simple intervention would work is addressed in this recent article by Anne Boring (Erasmus School of Economics) and Arnaud Philippe (University of Bristol), published in the Journal of Public Economics (ungated earlier version here). Boring and Philippe conducted an experiment in the 2015-16 academic year across seven campuses of Sciences Po in France, where each campus was assigned to one of three groups: (1) a control group; (2) a "purely normative" treatment; or (3) an "informational" treatment. As Boring and Philippe explain:

The administration sent two different emails to students during the evaluation period. One email—the ‘‘purely normative” treatment—encouraged students to be careful not to discriminate in SETs. The other email—the ‘‘informational” treatment—added information to trigger bias consciousness. It included the same statement as the purely normative treatment, plus information from the study on gender biases in SETs. The message contained precise information on the presence of gender biases in SET scores in previous years at that university, including the fact that male students were particularly biased in favor of male teachers.

Of students at the treated campuses, half received the email and half did not. No students at the control campuses received an email. In addition, the emails were sent after the period for students to complete evaluations had already started, so some students completed their evaluations before the treatment, and some after. This design allows Boring and Philippe to adopt a difference-in-differences analysis, comparing the difference in evaluations before and after the email for students at treatment campuses who were assigned to receive the email, with the difference in evaluations before and after the email for students at control campuses. The difference in those two differences is the effect of the intervention. Conducting this analysis, they find that:

...the purely normative treatment had no significant impact on reducing biases in SET scores. However, the informational treatment significantly reduced the gender gap in SET scores, by increasing the scores of female teachers. Overall satisfaction scores for female teachers increased by about 0.30 points (between 0.08 and 0.52 for the confidence interval at 5%), which represents around 30% of a standard error. The informational treatment did not have a significant impact on the scores of male teachers...

The reduction in the gender gap following the informational email seems to be driven by male students increasing their scores for female teachers. On the informational treatment campuses, male students’ mean ratings of female teachers increased from 2.89 to 3.20 after the emails were sent... Furthermore, the scores of the higher quality female teachers (those who generated more learning) seem to have been more positively impacted by the informational email.

That all seems very positive. Also, comparing evaluations from students at control campuses with evaluations from students in the control group at treated campuses before and after the email was sent allows Boring and Philippe to investigate whether the interventions had a spillover effect on those who did not receive the email. They find that:

...the informational treatment had important spillover effects. On informational treatment campuses, we find an impact on students who received the email and on students who did not receive the email. Anecdotal evidence suggests that this email sparked conversations between students within campuses, de facto treating other students.

The anecdotal evidence (based on responses to an email asking students about whether they had discussed the informational email with others) both provides a plausible mechanism to explain the spillover effects, and suggests that the emails may have been effective in spurring important conversations on gender bias. Also, importantly, the informational emails had an enduring effect. Looking at evaluations one semester later, Boring and Philippe find that:

The effect of the informational treatment remains significant during the spring semester: female teachers improved their scores. The normative treatment remained ineffective.

So, it appears that it is possible to reduce gender bias in student evaluations of teaching with a simple intervention.

Read more:

Sunday, 13 February 2022

More on gender bias in student evaluations of teaching

Back in 2020, I wrote a post on gender biases in student evaluations of teaching, highlighting five research papers that showed pretty clearly that student evaluations of teaching (SET) are biased against female teachers. I've recently read some further research on this topic that I thought I would share, some of which supports my original conclusion, and some of which should make us pause, or at least draw a more nuanced conclusion.

The first article is this 2020 one by Shao-Hsun Keng (National University of Kaohsiung), published in the journal Labour Economics (sorry, I don't see an ungated version online). Keng uses data from 2002 to 2015 from the National University of Kaohsiung, covering all departments. They have data on student evaluations, and on student grades, which they use to measure teacher value-added. In the simplest analysis, they find that:

...both male and female students give higher teaching evaluations to male instructors. Female students rate male instructors 11% of a standard deviation higher than female instructors. The effect is even stronger for male students. Male students evaluate male instructors 15% (0.109 + 0.041) of a standard deviation higher than female instructors.

 Interestingly, Keng also finds that:

Students who spend more time studying give higher scores to instructors, while those cutting more classes give lower ratings to instructors.

It is difficult to know which way the causality would run there though. Do students who are doing better in the class recognise the higher-quality teaching with better evaluations? Or do students who are enjoying the teaching more spend more time studying? Also:

Instructors who have higher MOST [Ministry of Sciences and Technology] grants receive 1.2% standard deviation lower teaching evaluations, suggesting that there might be a trade-off between research and teaching.

That suggests that research and teaching are substitutes (see my earlier post on this topic). Keng then goes on the separately analyse STEM and non-STEM departments, and finds that:

Gender bias in favor of male instructors is more prominent among male students, especially in STEM departments. Female students in non-STEM departments, however, show a greater gender bias against female instructors, compared to their counterparts in STEM departments.

In other words, both male and female students are biased against female teachers, but male students are more biased. Male STEM students are more biased than male non-STEM students, but female non-STEM students are more biased than female STEM students. Interesting. Keng then goes on to show that:

...the gender gap in SET grows as the departments become more gender imbalanced.

This effect is greater for female students than for male students, so female students appear to be more sensitive to gender imbalances. This is not as good as it may sound - it means that female students are more biased against female teachers in departments that have a greater proportion of male teachers (such as STEM departments). Finally, Keng uses their measure of value-added to make an argument that the bias against female teachers is related to statistical discrimination. However, I don't find those results persuasive, as they seem to rely on an assumption that as teachers remain at the institution longer, students learn about their quality. However, students are only at the institution for three or four years, and don't typically see the same teachers across multiple years, so it is hard to see that this is a learning effect on the students' side. I'd attribute it more to the teachers better understanding what it take to get good teaching evaluations.

Moving on, the second article is this 2021 article by Amanda Felkey and Cassondra Batz-Barbarich (both Lake Forest College), published in the AEA Papers and Proceedings (sorry, I don't see an ungated version online). Felkey and Batz-Barbarich conduct a meta-analysis of gender bias in student evaluations of teaching. A meta-analysis combines the results across many studies, allowing us to (hopefully) overcome statistical biases arising from looking at a single study. Felkey and Batz-Barbarich base their meta-analysis on US studies covering the period from 1987 to 2017, which includes 15 studies and 39 estimated effect sizes. They also compare economics with other social sciences. They find that:

In the 30 years spanned by our metadata, there was significant gender difference in SETs that favored men for economics courses... Gender difference in the rest of the social sciences favored women on average but was statistically insignificant...

The p-value for other social sciences is 0.734, so is clearly statistically insignificant. The p-value for economics is 0.051, which many would argue is also statistically insignificant (although barely so). However, in a footnote, Felkey and Batz-Barbarich note that:

We found evidence that our results for economics were impacted by publication bias such that the gender difference is actually greater...than our included studies and analyses suggest.

They don't present an analysis that accounts for the publication bias, which might have shown a more statistically significant gender bias. This is bad news for economics, but might it be good news for other disciplines? It's not consistent with the results of other analyses of gender bias in SETs, where it appears across all disciplines (see the Keng study above, or studies in this earlier post). Usually, I would strongly favour the evidence in a meta-analysis over individual studies, but it is difficult when the meta-analysis seems to show something different from the studies I have read. Moreover, Felkey and Batz-Barbarich don't find any evidence of publication bias in disciplines other than economics, which suggests the null finding for those disciplines is robust. Perhaps gender bias in teaching evaluations is really just a feature of economics and STEM disciplines? I'd want to see a more detailed analysis (the AEA Papers and Proceedings don't offer the opportunity for authors to include a lot of detail), before drawing a strong conclusion, but this should make us more carefully evaluate the evidence on gender bias, especially outside of the STEM disciplines.

Read more:

Friday, 11 February 2022

Gender bias in principals' evaluations of primary teachers in Ghana

I've written before about gender bias in student evaluations of teaching (see here, with more to come soon in a future post). There is good reason to worry that student evaluations don't even measure teaching quality (see here, with more on that to come too). However, it turns out that it isn't just students that are biased in evaluating teachers. This article by Sabrin Beg (University of Delaware), Anne Fitzpatrick (University of Massachusetts, Boston), and Adrienne Lucas (University of Delaware), published in the AEA Papers and Proceedings last year (ungated version here), shows that primary school principals, at least in Ghana, are biased as well.

Their data come from the Strengthening Teacher Accountability to Reach All Students (STARS) project, a randomised trial that collected data from 210 schools in 20 districts in Ghana. They asked fourth and fifth-grade teachers to rate their own performance (by comparing themselves to teachers at similar schools), and asked principals to rate their teachers. They also presented teachers and principals with vignettes, where the gender of the teacher described was randomised, and asked the principals to rate the teacher described in the vignette. Finally, they measured the 'teacher value-added' using standardised tests administered at the beginning and end of the year. Comparing ratings between male and female teachers, Beg et al. find that:

Female and male teachers were equally likely to assess themselves as at least more effective than other teachers at similar schools... In contrast, principals were about 11 percentage points less likely to assess female teachers this highly relative to male teachers...

The gender bias of principals was not statistically significant though (a point that Beg et al. do not note in the paper, preferring to concentrate on the magnitude of the coefficient). They also:

...test for gender differences in the objective measure of effectiveness based on student test scores and find that female teachers had on average 0.28 standard deviations higher effectiveness than their male peers...

This puts the statistical insignificance of the principals' gender bias into more context. Using an objective measure of teacher value-added, female teachers are better teachers than male teachers, and yet female teachers are not statistically significantly rated any better than male teachers by principals. No difference in subjective assessments, when objective assessments say that female teachers are better than male teachers, provides evidence of bias.

Coming to the vignettes though, Beg et al. find that:

Principals further demonstrated evidence of bias against women in their hypothetical assessments. Principals rated individuals 0.12 standard deviations less effective when they had a female name instead of a male one...

Again, the difference is not statistically significant (and doesn't provide strong evidence of bias, because there was by construction no difference in teacher quality between male and female teachers in the vignettes). Overall, I was a bit surprised by this study, because when Beg et al. graph principals' subjective assessments of male and female teachers against the teacher value-added, you get this (their Figure 1):

At every level of objective teacher value-added, male teachers (the black line) are subjectively rated better than female teachers (the blue line) by principals. And yet, the difference is statistically insignificant in Beg et al.'s regression model. Perhaps Beg et al. should have included the objective measure in their models (or included confidence intervals in their Figure 1). Overall, this provides some weak additional support for gender bias in the evaluation of teachers.

Read more: