© 2026 FUTURE PROOF™
Inside the Classroom · Teacher expectations

The teacher expectations effect, retested.

In 1968, a fake test and a random label appeared to raise children’s IQ — and Pygmalion became the most famous experiment in education. The real story took fifty years to finish: a fierce critique, a shrinking effect, and modern designs showing that a small, real influence can still reach a long way.

TL;DR

The finding: The teacher expectations effect is real — and far smaller than its legend. Pygmalion reported first graders gaining roughly 15 IQ points from a random label, but the pooled experimental estimate sits nearer 0.1 standard deviations, and the typical classroom effect around 0.1 to 0.2. The effect appears mainly when teachers do not yet know their pupils. Modern causal designs confirm it is small, real, biased for some groups — and able to reach outcomes as distant as finishing college.

The mechanism: Teacher expectations are mostly accurate — they follow the evidence teachers already have, which is why expectation and outcome correlate so strongly. The causal sliver travels through behaviour: who gets the harder question, the longer patience, the richer feedback. Labels do their worst when the evidence underneath them is stale — new classes, transferred pupils, children whose improvement nobody has re-measured.

The product: Future Proof Education™ attacks the stale-evidence problem directly: the Adaptive Diagnostic re-measures every child’s actual level continuously, dashboards show teachers where measurement and expectation disagree, and the AI Tutor pitches work to the data — it has no reputations to remember.

In this article

  1. 01The experiment that named an effect
  2. 02The backlash
  3. 03What hundreds of studies settled
  4. 04The accuracy turn
  5. 05Where expectations bite hardest
  6. 06The modern causal designs
  7. 07What the evidence doesn’t show
  8. 08Expectations by the evidence
© 2026 FUTURE PROOF™
The route. 8 sections, from “The experiment that named an effect” to “Expectations by the evidence”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

In the spring of 1964, every child in an ordinary California elementary school sat a test. Teachers were told it was the “Harvard Test of Inflected Acquisition”, an instrument that could spot children about to bloom intellectually. It could do no such thing — it was a standard IQ test, and the “bloomers” whose names teachers received, about one child in five, had been picked at random (Rosenthal & Jacobson, 1968). The only thing that differed about those children was what their teachers believed about them.

A year later the children were tested again, and the belief appeared to have become biology. In the youngest classes the gap was startling: first-grade bloomers were reported gaining roughly 27 IQ points against roughly 12 for their classmates — an advantage of about 15 points from a label drawn out of a hat. Second graders showed a similar, smaller pattern, and across the school the average advantage was modest, near 4 points (Rosenthal & Jacobson, 1968).

The number

≈+15 IQ pts The reported first-grade advantage of randomly labelled “bloomers” over their classmates after one year — the number that made Pygmalion famous, and the one the following fifty years of research could never reproduce at that size (Rosenthal & Jacobson, 1968).

Rosenthal and Jacobson called it Pygmalion in the Classroom, and the name stuck to the phenomenon itself: the self-fulfilling prophecy, expectations creating the reality they predict. Few results in social science travelled faster. Within years it was policy language, teacher-training doctrine and courtroom evidence. Which makes what happened next — inside the research literature — one of education’s most important corrections.

Labelled ‘bloomers’ Control classmates First graders +27.4 +12.0 Second graders +16.5 +7.0 All six grades +12.2 +8.4 0 +10 +20 +30 IQ-point gain over one school year — Rosenthal & Jacobson (1968), schematic © 2026 FUTURE PROOF™
Figure 1. Pygmalion’s reported IQ gains after one year: the dramatic advantage sat almost entirely in the first two grades, while the whole-school gap was a modest 3–4 points (Rosenthal & Jacobson, 1968). Values as reported in the original study; Section 2 explains why the youngest grades — tested with the least reliable instrument — are exactly where a sceptic would expect inflated numbers. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The backlash

The most influential review of Pygmalion was also the least kind. Robert Thorndike, one of the era’s leading measurement experts, went through the study’s numbers and found the foundations soft. The IQ instrument was used on children younger than it was built for, and administered to whole groups of six-year-olds. In places it produced classroom averages so low that, read literally, entire classes sat near the test’s floor (Thorndike, 1968). Gains measured from a floor that unreliable can be manufactured by measurement error alone.

Note what the critique targeted: not the idea that beliefs can matter, but the claim that this study had shown IQ moving because of them. The gaudy first-grade numbers came from precisely the classrooms where the measurement was weakest — the pattern a sceptic would predict (Thorndike, 1968). Replication attempts multiplied through the 1970s, and most found little or nothing; the effect, where it appeared, was a shadow of the original.

The dispute hardened into camps because the stakes were real. If a planted belief could add points to a child’s IQ, then teacher attitudes were a policy lever of enormous cheapness. If it could not, a generation of blame had been aimed at teachers on the strength of a broken ruler. Settling that required more than one school and one spring. It required counting the whole literature.

By the late 1970s the field had split into believers and debunkers, each with studies to wave. What it needed was arithmetic — pooling the whole pile, testing what moderates the result. Two syntheses supplied it, and between them they drew the picture that still stands.

What hundreds of studies settled

The first came from Rosenthal himself, with Rubin: a meta-analysis of 345 interpersonal-expectancy studies across laboratories, workplaces and classrooms. Across that whole literature the effect was real and unmistakably nonzero — people do behave differently toward those they expect more from, and the targets’ performance moves (Rosenthal & Rubin, 1978). The existence question was closed. The size question, for classrooms specifically, was not.

Raudenbush closed it. He gathered the 18 experiments that had tried to move pupil IQ by giving teachers induced expectations, and pooled them with one crucial moderator: how long the teacher had known the pupils before the label arrived. The pooled effect was small — near a tenth of a standard deviation. And the moderator swallowed almost all of it. When labels arrived in the first week or two of acquaintance, effects appeared. Once teachers had a few weeks of their own evidence, the induced label did essentially nothing (Raudenbush, 1984).

That timing gradient is the single most practical fact in this literature. A false belief could not beat live data. Expectancy effects are, at root, a property of information-poor moments: the new class, the new school, the fresh transfer file. Where teachers hold current, personal evidence about a child, labels lose (Raudenbush, 1984).

≈+0.3 when the label came first ≈0 once teachers knew their pupils +0.3 +0.2 +0.1 0 0 1 2 4 8 12 weeks the teacher had known the class before the label (schematic) © 2026 FUTURE PROOF™
Figure 2. The credibility gradient across 18 expectancy-induction experiments: sizeable effects when the label arrived before teachers had their own evidence, falling to roughly zero once they had known the class for more than about two weeks (Raudenbush, 1984). Schematic: the curve renders the synthesis’ reported gradient; individual points are approximate readings, not published estimates. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The accuracy turn

A parallel line of work asked a more basic question: in ordinary classrooms — no fake tests, no planted labels — how do teacher expectations relate to what pupils go on to do? The correlations are strong. But Brophy’s synthesis of the naturalistic research put the causal share low — around 5 percent of achievement variance. The reason is mundane: most teacher expectations are simply accurate, formed from real performance and updated as evidence arrives (Brophy, 1983).

Jussim and Harber’s landmark review made the accounting explicit. Self-fulfilling prophecies in classrooms are real, replicated — and typically small, with effects around 0.1 to 0.2. Expectations predict outcomes strongly mostly because teachers predict well, not because their beliefs manufacture the result. Claims of effects accumulating year over year into large gaps found little support. And the review flagged one sharp exception worth its own section: the effects are not evenly distributed across children (Jussim & Harber, 2005).

The misread

An expectation that predicts is not an expectation that causes. Teacher forecasts track pupil outcomes mainly because they are accurate (Brophy, 1983) (Jussim & Harber, 2005). Reading every correlation as prophecy overstates teacher power — and misdirects the fix, which is calibration, not optimism on command.

Where expectations bite hardest

The small average hides a patterned spread. Jussim and Harber’s review found self-fulfilling effects two to three times larger for pupils from stigmatised or low-status groups — children from poor families, Black students in the American studies, low prior achievers generally (Jussim & Harber, 2005). Accuracy is the norm; where it frays, it frays along social lines. The children with the least margin absorb the most prophecy.

How does a belief travel from a teacher’s head into a child’s results? The observational work mapped the routes early. High-expectation pupils get more chances to answer, longer pauses before help arrives, and more precise feedback. Low-expectation pupils get easier questions, faster rescues, and more praise for less work (Brophy, 1983). Each behaviour is small on its own. Together they change the daily diet of challenge a child receives — and challenge is what drives learning.

The bias is measurable before its effects are. Researchers can define expectation bias precisely: the gap between what a teacher predicts and what the child’s measured achievement and aptitude would predict. Those gaps are not random noise. They tilt against pupils from lower-income homes and some minority groups, and they persist (de Boer, Bosker & van der Werf, 2010). In the American data, teachers’ degree expectations for the same Black student differ by the teacher’s own background — a disagreement that cannot be explained by the student (Papageorge, Gershenson & Kim, 2020).

The modern causal designs

The old experiments planted false labels; the modern designs measure real expectations and chase their consequences with better statistics. De Boer and colleagues followed thousands of Dutch pupils from the end of primary school through secondary school. Pupils whose teachers under-expected them, relative to their measured ability, sat in lower school positions five years on. The bias effect did not wash out with time. Over-expected pupils drifted upward in parallel (de Boer, Bosker & van der Werf, 2010). Small pushes, delivered at a sorting point, held.

Papageorge, Gershenson and Kim closed the causal loop on the longest outcome yet. In a national American cohort, two teachers each rated how far a tenth grader would go in school; their disagreements gave the analysis leverage to separate the expectation from the student. The finding: teacher expectations causally move the probability of finishing college — modestly per teacher, meaningfully across a school career. And because expectations differ systematically across teachers for the same child, the bias documented above is not cosmetic; it compounds into credentials (Papageorge, Gershenson & Kim, 2020).

Fifty years on, the ledger closes neatly. Pygmalion’s headline number does not survive scrutiny (Thorndike, 1968). Its central idea does — at perhaps a tenth of the advertised size (Raudenbush, 1984). The effect concentrates where evidence is stale and status is low (Jussim & Harber, 2005). And it matters because school is a chain of sorting decisions, where small pushes settle into records (Papageorge, Gershenson & Kim, 2020).

First graders (1968) ≈1.0 SD All grades (1968) ≈0.25 18 experiments pooled (1984) ≈0.1 Typical modern estimate (2005) ≈0.1–0.2 0 0.2 0.4 0.6 0.8 1.0 Expectancy effect on IQ / achievement (SD), approximate, by source © 2026 FUTURE PROOF™
Figure 3. The estimate shrinks as the evidence hardens. Pygmalion’s first-grade advantage of roughly 15 IQ points is about one standard deviation; the study’s whole-school advantage is nearer 0.25 (Rosenthal & Jacobson, 1968); the pooled 18-experiment estimate sits near 0.1 (Raudenbush, 1984); and the typical naturalistic classroom effect lands at roughly 0.1–0.2 (Jussim & Harber, 2005). Approximate conversions on one scale for comparison — small, real, and a long way from the legend. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Teacher expectations and self-fulfilling prophecies: knowns and unknowns, resolved and unresolved controversies. Jussim & Harber, Personality and Social Psychology Review, 2005

What the evidence doesn’t show

The expectations literature is unusually rich in things people believe it shows and it does not. Six boundaries keep the reading honest.

  • No IQ transformations. The 15-point first-grade gain sits exactly where the measurement was weakest, and nothing near it has replicated; the defensible pooled effect is an order of magnitude smaller (Thorndike, 1968) (Raudenbush, 1984).
  • Expectations are mostly accurate. The strong expectation–outcome correlations mainly reflect teachers predicting well; treating every correlation as bias misreads the evidence (Brophy, 1983) (Jussim & Harber, 2005).
  • Accumulation is unproven. The claim that small yearly effects compound into life-defining gaps found little direct support in the classic review; the college-completion result shows reach, not unbounded growth (Jussim & Harber, 2005) (Papageorge, Gershenson & Kim, 2020).
  • Group moderators rest on thinner pools. Larger effects for stigmatised groups recur, but in fewer studies with smaller samples than the headline literature (Jussim & Harber, 2005).
  • Deception designs travel badly. Planted false labels tell us what a false belief can do in an information vacuum — not what banning judgment would do in a live classroom (Raudenbush, 1984).
  • The lever is behaviour, not belief. Expectations act through challenge, patience, grouping and feedback; no study shows that commanding optimism, without changing those behaviours, changes children (Brophy, 1983).

Where the evidence stops

  1. 1No IQ transformations
  2. 2Expectations are mostly accurate
  3. 3Accumulation is unproven
  4. 4Group moderators rest on thinner pools
  5. 5Deception designs travel badly
  6. 6The lever is behaviour, not belief
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Expectations by the evidence

For schools, the fifty-year correction converts into a short and slightly surprising manual: the target is not teachers’ hearts but their information.

Refresh the evidence expectations run on. Labels beat teachers only in information-poor moments; live data beats labels (Raudenbush, 1984). The dangerous weeks are the first ones — new classes, new arrivals, handover files. Fast, skill-level measurement at every transition shrinks the window where prophecy can operate.

Audit calibration, not attitude. The measurable problem is the gap between what a teacher expects and what the child’s data predicts — and its tilt by income and ethnicity (de Boer, Bosker & van der Werf, 2010). Schools can compute that gap. Optimism training without measurement is the intervention the evidence declined to support (Brophy, 1983).

Keep the challenge channel open. Expectation effects travel through who gets asked the hard question, given the demanding text, allowed the second attempt (Brophy, 1983). Holding task difficulty open for every child blocks the mechanism without requiring anyone to believe anything on command.

Let improvement register. A sustaining expectation — holding last term’s picture of a child who has since moved — is the everyday version of the effect (Brophy, 1983). Re-measure often enough that a child’s growth forces an update, and route the update to whoever holds the next sorting decision (Papageorge, Gershenson & Kim, 2020).

Guard the sorting points hardest. The modern results show small pushes persisting when they land at transitions — track placement, subject choice, the college conversation (de Boer, Bosker & van der Werf, 2010). Those are the moments to require data beside judgment, and to double-check the children the literature says carry the most prophecy (Jussim & Harber, 2005).

Applied at Future Proof Education

Expectations, re-measured.

The evidence says expectancy effects live where information is stale — so Future Proof Education™ keeps it fresh. The Adaptive Diagnostic re-measures each child’s actual level, skill by skill, all year: every new class and every transfer arrives with current evidence instead of a reputation. Teacher dashboards show exactly where expectation and measurement disagree — the calibration gap the research warns about — and flag the child whose data has outrun their placement. The AI Tutor pitches every task to the measured level; it has no reputations to remember and no priors to sustain. Parents see levels and growth rather than labels, and school systems can watch the expectation gaps that pattern by background, at scale, while they close.

See it in the classroom
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 1968Rosenthal
  • 1968Thorndike
  • 1978Rosenthal
  • 1983Brophy
  • 1984Raudenbush
  • 2005Jussim
  • 2010de Boer
  • 2020Papageorge
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 1968–2020, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Rosenthal, R., & Jacobson, L. (1968). Pygmalion in the Classroom: Teacher Expectation and Pupils’ Intellectual Development. Holt, Rinehart & Winston. PDF
  2. Thorndike, R.L. (1968). Review of Pygmalion in the Classroom. American Educational Research Journal 5(4): 708–711. PDF
  3. Rosenthal, R., & Rubin, D.B. (1978). Interpersonal expectancy effects: The first 345 studies. Behavioral and Brain Sciences 1(3): 377–386. PDF
  4. Raudenbush, S.W. (1984). Magnitude of teacher expectancy effects on pupil IQ as a function of the credibility of expectancy induction: A synthesis of findings from 18 experiments. Journal of Educational Psychology 76(1): 85–97. PDF
  5. Brophy, J.E. (1983). Research on the self-fulfilling prophecy and teacher expectations. Journal of Educational Psychology 75(5): 631–661. PDF
  6. Jussim, L., & Harber, K.D. (2005). Teacher expectations and self-fulfilling prophecies: Knowns and unknowns, resolved and unresolved controversies. Personality and Social Psychology Review 9(2): 131–155. PDF
  7. de Boer, H., Bosker, R.J., & van der Werf, M.P.C. (2010). Sustainability of teacher expectation bias effects on long-term student performance. Journal of Educational Psychology 102(1): 168–179. PDF
  8. Papageorge, N.W., Gershenson, S., & Kim, K.M. (2020). Teacher expectations matter. Review of Economics and Statistics 102(2): 234–251. PDF
Try Future Proof Education

Fresh evidence beats old labels.

Book a 20-minute demo. We’ll show the Adaptive Diagnostic re-measuring what each child can actually do, dashboards that surface calibration gaps, and tutoring that pitches to the data — not the reputation.

8 citations Reviewed August 2026 Open peer review welcomed