The gender gap in STEM: what the research shows
Girls and boys now post nearly identical maths scores, yet physics and computing classrooms stay stubbornly male. Between those two facts sits one of psychology’s most famous ideas — and its hardest replication fight. Here is what the research actually supports, and what schools can do with it.
The finding: Gender gap STEM research tells two stories at once. In maths performance, the gap has essentially closed: across roughly seven million US students, the difference in average scores is trivial at every grade. In participation, the gap is real but field-specific — the life sciences are majority-female while physics, engineering and computing are not.
The mechanism: The famous explanation, stereotype threat, is genuinely contested. The founding lab result was real, but the school literature carries publication bias, and the best-controlled classroom studies find little or nothing. The sturdier drivers sit elsewhere: the image of each field, early experience of it, belonging, and course choices steered by relative strengths.
The product: Future Proof Education™ is built for the second story. Adaptive practice holds every learner to the same standard, the Adaptive Diagnostic measures skill rather than confidence, and teacher dashboards surface participation gaps at the moment subjects are chosen — before a quiet opt-out becomes a career path.
In this article
- 01Two gaps, one label
- 02The performance gap that closed
- 03The similarities hypothesis
- 04The famous experiment
- 05The replication debate
- 06Where the gap actually lives
- 07What helps: belonging and representation
- 08What the evidence doesn’t show
- 09Closing the gap by the evidence
The phrase “gender gap in STEM” smuggles two different claims under one label. The first is about ability: that girls cannot do maths and science as well as boys. The second is about participation: that women are scarce in some scientific fields. The first claim is measurable, and it has been measured on an enormous scale. The second is visible in any physics classroom in the world. The central finding of the research is that these two claims have surprisingly little to do with each other.
A famous experiment sits at the centre of the story. In 1999, psychologists showed that a one-line reminder of a stereotype could depress able women’s scores on a hard maths test (Spencer, Steele & Quinn, 1999). “Stereotype threat” became one of the most cited ideas in modern psychology. It reached teacher training, education policy and a shelf of popular books. Then the replication era arrived, and the effect began to shrink under scrutiny. This article treats that fight as a feature, not an embarrassment. Watching a field audit its own favourite idea is the best science lesson on offer.
What follows moves in layers. First the performance data, which is unusually clean. Then the stereotype-threat literature and its replication debate, which is not. Then the question the gap actually poses — why able girls walk away from particular fields — and what schools, parents and ministries can do about it. Eight anchor studies carry the argument.
Two gaps, one label
Start by splitting the label. The performance gap is a difference in what boys and girls can do: test scores, grades, examination marks. The participation gap is a difference in where they go: subjects chosen, degrees taken, careers entered. The two gaps have different sizes, different causes and different fixes. Treating them as one problem called “STEM” is how most policy in this area goes wrong.
The participation gap is also lumpier than the acronym suggests. Women earn roughly half of US science and engineering bachelor’s degrees overall. But the average hides the pattern. Women are a solid majority in psychology and the life sciences, near parity in chemistry and maths — and under a fifth of graduates in physics, engineering and computer science (Cheryan et al., 2017). There is no single STEM gap. There is a cluster of specific fields that lose girls, and a research question about why.
Hold that shape in mind through everything that follows. The ability story and the field story are usually told together, each borrowed as evidence for the other. The data lets us pull them apart — and once apart, they point to very different actions.
The performance gap that closed
Older readers were taught, implicitly or otherwise, that boys are better at maths. For the average scores of their generation there was something to it: through the middle of the last century, boys held a modest advantage on maths tests, especially in older students. That advantage then shrank decade by decade as girls’ course-taking caught up — a decline documented across successive research syntheses (Hyde, 2005).
The landmark test of where things now stand came in 2008. Hyde and colleagues used US state assessment data gathered under the No Child Left Behind act: roughly seven million students, grades 2 to 11 (Hyde et al., 2008). The result was stark. At every grade, the gender difference in average maths performance was trivial — effect sizes between about d ≈ −0.02 and d ≈ 0.06. (Cohen’s d expresses a difference in standard-deviation units; 0.20 is conventionally called “small”.) In plain terms, the two score distributions sit almost exactly on top of each other.
≈7 million US students in the state-assessment analysis. At every grade from 2 to 11, the gender difference in mean maths score fell between d ≈ −0.02 and d ≈ 0.06 — a rounding error by the standards of educational effects (Hyde et al., 2008).
Two caveats belong beside the headline. The state tests were thin on genuinely hard problem-solving items, so the study says least about the very top of the curve. And boys’ scores were slightly more spread out than girls’ — somewhat more boys at both extremes — a variance question the averages cannot settle (Hyde et al., 2008). Neither caveat rescues the folk belief. On what schools actually test, the ability gap is gone.
The similarities hypothesis
Maths is one instance of a broader pattern. In 2005, Hyde reviewed 46 meta-analyses of psychological gender differences — pooled studies covering cognition, communication, personality and motor skills (Hyde, 2005). Her tally became famous: roughly 78% of the pooled differences were small or near zero. She named the conclusion the gender similarities hypothesis. On most psychological variables, males and females are simply more alike than different.
The hypothesis is not a claim that no differences exist. Some are large and robust: throwing velocity is the textbook case, with a difference above two full standard deviations. A few cognitive measures, such as mental rotation, show moderate gaps (Hyde, 2005). The point is proportion. The differences the culture obsesses over — maths ability chief among them — sit in the crowded near-zero end of the distribution.
It helps to feel what a small d means. At d ≈ 0.1, the two bell curves overlap almost completely, and knowing a child’s gender tells you nearly nothing about their maths score. Averages this close cannot explain university classrooms that are 80% male. Something other than ability is doing that work.
The famous experiment
None of this parity was obvious in the 1990s, when the test gaps were closing but not yet closed. Into that moment landed the experiment that defined a field. Spencer, Steele and Quinn recruited university men and women with strong, roughly equal maths records and gave them a very hard test built from graduate-entrance items (Spencer, Steele & Quinn, 1999). Half were told the test had shown gender differences in the past. Half were told it had not. The framing changed everything. Under the gender-difference framing, women scored well below men. Under the no-difference framing, the gap disappeared.
The proposed mechanism was elegant. A negative stereotype about your group adds a second task to the test: the fear of confirming it. That fear occupies working memory — the limited mental workspace hard maths depends on — and performance drops. The idea was hopeful, too. If part of the gap was created by the testing situation, then cheap changes to the situation should shrink it. No remediation, no decade-long reform: just remove the threat.
The finding travelled fast and far. It reached teacher training, test-administration guidance and policy documents. For years the published record seemed to keep agreeing, as dozens of small conceptual replications reported effects in many groups and settings (Flore & Wicherts, 2015). The theory hardened into fact. Then researchers started checking the fact’s foundations.
The replication debate
Doubt arrived through the field’s standard instruments: funnel plots and pre-registration. In 2015, Flore and Wicherts meta-analysed the studies that matter most for education — stereotype-threat experiments on school-aged girls (Flore & Wicherts, 2015). The raw pooled effect was modest: under threat, girls scored lower by roughly d ≈ 0.2. Then the authors tested the literature itself. Small studies with big effects were plentiful; small studies with null results were strangely absent. That asymmetry is the signature of publication bias. Corrected for it, the best estimate fell close to zero.
Direct tests in ordinary school samples pointed the same way. Ganley and colleagues ran threat manipulations across three studies and more than nine hundred students, from primary age to high school (Ganley et al., 2013). No condition produced the predicted drop in girls’ scores. Adequately powered, honestly analysed, published anyway — exactly the kind of study the older literature lacked.
A literature can look unanimous and still be wrong. Small threat studies that found effects were published; comparable studies that found nothing mostly were not. Correct for that bias and the school-age estimate lands near zero (Flore & Wicherts, 2015). Any program sold to schools on stereotype-threat grounds should be priced with that fact on the table.
What should a careful reader conclude? Not that the original experiment was fake, and not that social pressure never touches performance. The honest verdict is narrower. Stereotype threat, as measured in real school settings, is not the robust general force the 2000s took it to be. It cannot carry the weight of explaining the participation gap. Policies that lean on it alone are leaning on the weakest plank in the platform.
Gender similarities characterize math performance.Hyde, Lindberg, Linn, Ellis & Williams, Science, 2008
Where the gap actually lives
If ability is equal and threat is shaky, why are physics and computing still male? The most useful framework comes from Cheryan and colleagues’ review of field-level differences (Cheryan et al., 2017). Their move is to change the question. Not “why do women avoid STEM?” — women dominate several STEM fields — but “why do some STEM fields recruit women while others repel them?”
Three factors separate the balanced fields from the unbalanced ones in their account (Cheryan et al., 2017). First, masculine field cultures: the stereotype of who belongs in computing or physics — the obsessive lone genius — fits fewer girls’ sense of themselves, and signals it early. Second, insufficient early experience: pupils meet biology and chemistry properly at school, but many choose or reject physics and computing degrees having barely met the subjects. Third, self-efficacy gaps: girls’ confidence in these specific fields runs below their measured skill, and confidence feeds choices even when ability does not differ.
A stranger finding complicates the story. Stoet and Geary examined achievement and degree choices across dozens of countries and found what they called the gender-equality paradox: women’s share of STEM degrees is often lower in the most gender-equal, affluent countries (Stoet & Geary, 2018). One reading is that where economic pressure is light, choices track interests and relative strengths more freely. And relative strengths differ on average: girls who are strong in maths are, more often than boys, even stronger verbally — so comparative advantage quietly steers some of them elsewhere (Stoet & Geary, 2018). The paradox’s interpretation is contested, and this article leans on it only as a caution: participation numbers are not a simple gauge of fairness, and closing them is not purely a matter of removing barriers.
What helps: belonging and representation
The intervention evidence is younger and thinner than the descriptive evidence, but it points somewhere useful. The best-known classroom result comes from university physics. Miyake and colleagues randomized a values-affirmation exercise — two short writing tasks in which students wrote about values that mattered to them — into an introductory course (Miyake et al., 2010). The gender gap in exam scores shrank substantially. On a standardized concept test, it essentially disappeared in that cohort. Fifteen minutes of writing, twice, moved a gap that lectures had not.
The honest footnote: later attempts to reproduce affirmation effects have not always succeeded, and the conditions under which they work are still being mapped. This article returns to that limit below. The fair summary of belonging-style interventions is promising, cheap and unreliable — worth trying, not worth promising. Readers who followed the growth-mindset scale-up story on our corporate research site will recognize the shape: see the growth-mindset evidence. Light-touch psychology travels badly.
The structural levers look sturdier in the review evidence (Cheryan et al., 2017). Give girls real early experience of computing and physical science before the choice points arrive. Redesign introductory courses and their imagery so the field’s face is broader than the lone genius. Put counter-stereotypical role models in front of pupils routinely, not as an event. None of these is a psychological trick. They change what the field appears to be, which is what the choice evidence says matters.
What the evidence doesn’t show
An honest reading of this literature has to mark its own boundaries. Six matter most.
- Threat is contested, not debunked. The bias-corrected school estimate sits near zero on average, but averages permit moderators; specific tasks, ages or settings may yet produce real effects (Flore & Wicherts, 2015).
- The tails are still argued. Mean parity coexists with slightly greater male score variance in some datasets, and what that means at the 99th percentile remains disputed (Hyde et al., 2008).
- One-session fixes may not scale. The physics affirmation result is real, but the wider affirmation literature is heterogeneous and school-age evidence is thin (Miyake et al., 2010).
- No single cause explains field choice. Culture, early experience, confidence, interests and discrimination are entangled; the review evidence apportions between them, it does not exonerate any (Cheryan et al., 2017).
- The paradox has rival readings. The cross-national equality pattern is real in the data used but its measurement and meaning are both contested (Stoet & Geary, 2018).
- Averages do not diagnose a child. Every result here is a group statistic; none licenses a prediction about the girl or boy in front of you.
Where the evidence stops
- 1Threat is contested, not debunked
- 2The tails are still argued
- 3One-session fixes may not scale
- 4No single cause explains field choice
- 5The paradox has rival readings
- 6Averages do not diagnose a child
Closing the gap by the evidence
Read together, the eight studies turn into a short operating manual for schools and school systems — one that spends effort where the mechanisms actually are.
Teach the parity data, plainly. Pupils, teachers and parents should know that the maths score gap is roughly zero across seven million students (Hyde et al., 2008). Expectations are an instructional input like any other, and this one is free.
Audit the choice points, not just the marks. The gap opens where subjects become optional, not on exam day. Track who drops physics and computing, and exactly when. A school that measures attainment but not participation is watching the wrong dial (Cheryan et al., 2017).
Fix the field’s image before the field. Early, real experience of programming and physical science; introductory courses whose examples and imagery admit more kinds of person; role models as routine, not spectacle. These are the levers the field-difference evidence supports (Cheryan et al., 2017).
Handle the threat story with care. Stripping stereotype cues from tests and classrooms is cheap and harmless, so do it. Building funded programs on the lab effect is another matter — the school evidence will not hold the weight (Flore & Wicherts, 2015).
Counsel strengths without closing doors. A girl who is better at English than maths may still be better at maths than most of her class. Comparative advantage should inform guidance, never decide it silently (Stoet & Geary, 2018). Make the trade-off visible to the pupil, and the choice stays hers.
How Future Proof Education™ applies this.
The evidence says the ability gap is closed and the participation gap is made of choices, confidence and course images. So the platform watches choices, not stereotypes. Adaptive practice holds every learner to the same standard by construction. The Adaptive Diagnostic measures skill directly, so a confident guesser and a hesitant expert are both seen accurately. The Knowledge Map shows each pupil the full route into physics, computing and engineering — not just the next test. Teacher dashboards flag participation gaps at subject-choice moments, parents see the same signal at home, and ministries get field-level participation reporting at system scale.
See Future Proof for schools →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above.
The evidence, by year
- 1999Spencer
- 2005Hyde
- 2008Hyde
- 2010Miyake
- 2013Ganley
- 2015Flore
- 2017Cheryan
- 2018Stoet
- Spencer, S.J., Steele, C.M., & Quinn, D.M. (1999). Stereotype threat and women’s math performance. Journal of Experimental Social Psychology 35(1): 4–28. PDF
- Cheryan, S., Ziegler, S.A., Montoya, A.K., & Jiang, L. (2017). Why are some STEM fields more gender balanced than others? Psychological Bulletin 143(1): 1–35. PDF
- Hyde, J.S. (2005). The gender similarities hypothesis. American Psychologist 60(6): 581–592. PDF
- Hyde, J.S., Lindberg, S.M., Linn, M.C., Ellis, A.B., & Williams, C.C. (2008). Gender similarities characterize math performance. Science 321(5888): 494–495. DOI
- Flore, P.C., & Wicherts, J.M. (2015). Does stereotype threat influence performance of girls in stereotyped domains? A meta-analysis. Journal of School Psychology 53(1): 25–44. DOI
- Ganley, C.M., Mingle, L.A., Ryan, A.M., Ryan, K., Vasilyeva, M., & Perry, M. (2013). An examination of stereotype threat effects on girls’ mathematics performance. Developmental Psychology 49(10): 1886–1897. PDF
- Stoet, G., & Geary, D.C. (2018). The gender-equality paradox in science, technology, engineering, and mathematics education. Psychological Science 29(4): 581–593. DOI
- Miyake, A., Kost-Smith, L.E., Finkelstein, N.D., Pollock, S.J., Cohen, G.L., & Ito, T.A. (2010). Reducing the gender achievement gap in college science: A classroom study of values affirmation. Science 330(6008): 1234–1237. PDF
Catch the gap where it opens.
Book a 20-minute demo. We’ll show you adaptive maths practice that holds every learner to the same standard — and dashboards that surface participation gaps at the moment subjects are chosen.