© 2026 FUTURE PROOF™
Inside the Classroom · Class size

Class size: what the research evidence shows

Smaller classes are the most popular idea in education, and one of the most expensive. For once the evidence is unusually complete: a true randomized experiment, a famous null result, and a statewide scale-up that went wrong. Here is what it says — and what it costs.

TL;DR

The finding: The best class size research evidence comes from one true experiment. In Tennessee’s Project STAR, cutting kindergarten-to-grade-3 classes from about 23 pupils to about 15 raised achievement by roughly 0.2 standard deviations, with effects about twice as large for Black students. Around it sit a famous null — modest changes in ordinary class-size ranges buy nothing measurable — and a California scale-up that delivered a fraction of the pilot.

The mechanism: The leading candidate is attention per pupil in the earliest grades. Test-score gains fade once children return to regular classes, but the experiment leaves a long echo: more college-test taking, higher college attendance, better adult outcomes. The gains survive only where teacher quality survives the hiring wave that class-size cuts trigger.

The product: Future Proof Education™ works the same mechanism from the software side. The AI Tutor gives each pupil individual explanation and feedback at any class size, dashboards make a class of 28 legible the way a class of 15 is, and ministries weighing class-size budgets get measurement that shows what the spending actually changed.

In this article

  1. 01The experiment Tennessee ran
  2. 02What STAR actually found
  3. 03The fade-out and the long echo
  4. 04Hoxby’s null
  5. 05Maimonides’ rule and the wider evidence
  6. 06California: the scale-up warning
  7. 07The cost-effectiveness question
  8. 08What the evidence doesn’t show
  9. 09Class size by the evidence
© 2026 FUTURE PROOF™
The route. 9 sections, from “The experiment Tennessee ran” to “Class size by the evidence”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Ask parents what would improve their child’s school and smaller classes usually top the list. Ask a finance ministry what smaller classes cost and the enthusiasm cools: teachers are the most expensive input in schooling, and cutting class size means hiring more of them, plus building the rooms to put them in. Few education policies combine this much public affection with this much money. Which makes it strange how rarely anyone cites the evidence — because for once, the evidence is excellent.

Class size is one of a handful of education questions with a true randomized experiment behind it. In the 1980s, the state of Tennessee assigned thousands of young children, at random, to smaller or larger classes, and then followed them (Mosteller, 1995). Around that experiment sits an unusually complete supporting cast: careful re-analyses, a famous and precise null result, natural experiments from other countries, and a statewide rollout that shows what happens when a pilot meets reality.

This article walks the whole set. First the Tennessee experiment and what it really found, including the parts advocates skip. Then the fade-out, and the strange long echo in adulthood. Then the null results and the scale-up failure, which are not contradictions but boundary markers. Last, the question every budget holder actually faces: is a standard deviation bought this way worth its price?

The experiment Tennessee ran

Project STAR — Student/Teacher Achievement Ratio — ran from 1985 to 1989. About 6,500 kindergarten children in roughly 80 Tennessee schools were randomly assigned to one of three kinds of class: small (13 to 17 pupils), regular (22 to 25), or regular with a full-time teaching aide (Mosteller, 1995). Teachers were randomized too. Children kept their assignment through grade 3, new entrants were randomized as they arrived, and around 11,600 students passed through the experiment in total. Then everyone returned to ordinary classes, and Tennessee kept following them.

The design matters as much as the result. Because assignment was random within schools, small and regular classes had the same mix of families, incomes and prior abilities. Any later difference could be read as caused by class size. Education has very few levers tested this way at this scale, which is why a statistician of Frederick Mosteller’s standing treated STAR less as a study than as an event (Mosteller, 1995).

The first reports landed in 1990. Finn and Achilles found clear advantages for small classes in both reading and maths by the end of the early grades, and essentially nothing for the teaching-aide classes (Finn & Achilles, 1990). That second result deserves more fame than it has. An extra adult in the room, at real cost, moved achievement by approximately zero. Whatever small classes do, a cheaper adult-shaped substitute did not do it.

What STAR actually found

Big experiments attract big scrutiny, and STAR had real flaws: children left the study, some switched class types, and schools varied in how faithfully they held sizes. The definitive accounting came from the economist Alan Krueger, who rebuilt the data and stress-tested the result against every one of those problems (Krueger, 1999). The effect survived. Entering a small class raised test performance by roughly four percentile points in the first year, and the advantage grew by about one further point for each additional year spent there.

In effect-size terms the small-class advantage came to roughly 0.2 standard deviations by the end of the experiment — a modest, real, useful effect. (A standard deviation, SD, is the yardstick of educational effects; 0.2 SD moves a middle-of-the-pack child noticeably, not dramatically.) Two details carry most of the policy weight. The gains were concentrated in the earliest years. And they were not evenly spread: effects for Black students ran roughly twice the average, with larger-than-average gains for low-income students too (Krueger, 1999). Class size, tested honestly, is partly an equity instrument.

The number

≈11,600 children passed through the STAR experiment — randomly assigned, with their teachers, to small or regular classes across some 80 Tennessee schools. It remains the largest true randomized experiment on class size ever run (Mosteller, 1995).

Kindergarten ≈4 pts Grade 1 ≈5 pts Grade 2 ≈6 pts Grade 3 ≈7 pts 0 2 4 6 8 Small-class advantage on standardized tests (percentile points) © 2026 FUTURE PROOF™
Figure 1. The STAR advantage compounds while children stay in small classes: roughly four percentile points in the first year, growing by about one point per further year, to around 0.2 SD by grade 3 — and roughly twice as large for Black students (Krueger, 1999). Schematic: values are approximate averages across entry cohorts, not per-grade point estimates. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The fade-out and the long echo

After grade 3 the experiment ended and everyone moved to regular classes. The test-score advantage then behaved the way early-education gains usually behave: it faded. By middle school, the former small-class students’ edge had shrunk to a few percentile points — detectable, but a fraction of its grade-3 peak (Krueger & Whitmore, 2001). For years, that fade-out was quoted as the punchline. Buy an effect, watch it evaporate.

Then the follow-ups reached adulthood, and the punchline reversed. Krueger and Whitmore found that students from small classes were more likely to take a college entrance exam — roughly 44% versus 40% (Krueger & Whitmore, 2001). The effect sat exactly where the test gains had been. Among Black students, exam-taking rose from about 32% to about 40%, sharply cutting the gap with white students (Krueger & Whitmore, 2001). A test-score effect that had mostly faded was still steering lives a decade later.

The deepest follow-up came from Chetty and colleagues, who linked STAR’s children to their adult tax records (Chetty et al., 2011). Students assigned to small classes were about two percentage points more likely to attend college. Stranger still, the overall quality of a child’s kindergarten class — measured by classmates’ score gains — predicted earnings at age 27, home ownership and retirement savings (Chetty et al., 2011). Yet that same class quality had stopped moving test scores by grade 8. The authors’ reading: early classrooms build skills that tests stop measuring but employers do not. Fade-out, in this literature, is not the last word.

experiment ends 0 0.1 0.2 test-score effect (SD) K 3 5 8 grade (STAR cohort) regular small K–3 40.0 43.7 31.8 40.2all students Black students took a college entrance exam (%) © 2026 FUTURE PROOF™
Figure 2. The fade-out and the long echo. Left: the small-class test advantage fades after children return to regular classes in grade 4, from roughly 0.2 SD toward 0.05 by grade 8 — the curve is schematic, drawn through approximate cohort values. Right: years later the same students were more likely to take a college entrance exam, with the largest jump among Black students (Krueger & Whitmore, 2001). Adult tax-record data extend the echo to college attendance and earnings (Chetty et al., 2011). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
One of the most important educational investigations ever carried out. Mosteller on Project STAR, The Future of Children, 1995

Hoxby’s null

If the story ended there, every budget debate would be short. It does not end there. In 2000, Caroline Hoxby published the study that advocates least like to cite (Hoxby, 2000). She used natural variation in class size across 649 Connecticut elementary schools — the accidental ups and downs that enrolment numbers produce from year to year, including the sharp jumps that occur when a cohort crosses a class-splitting threshold. These accidents move class size without anyone selecting who sits where, which makes them a fair test.

The result: no detectable effect of class size on achievement. And not a vague null — the estimates were precise enough to rule out even quite modest effects across the ranges Connecticut actually experienced (Hoxby, 2000). In ordinary schools, drifting between, say, 19 and 25 pupils, achievement simply did not move with the roster.

The temptation is to stage Hoxby against STAR as a contradiction. Read carefully, they are answering different questions. STAR tested a large, sustained cut — about eight pupils, held for four years, in the earliest grades. Hoxby tested the small fluctuations that ordinary demography deals out, across a wider range of grades and mostly around larger baseline sizes (Hoxby, 2000). Both results can be true at once, and probably are: big early cuts help, small trims do nothing. For policy, the null is as valuable as the experiment. It marks the price floor below which the purchase is worthless.

Maimonides’ rule and the wider evidence

Between the experiment and the null sits a third kind of evidence: natural experiments abroad. The classic comes from Israel, where a rule dating to the medieval scholar Maimonides caps classes at 40. When a grade cohort passes 40 pupils, the school must split it — so a cohort of 41 suddenly sits in classes of about 20, while a cohort of 39 stays at 39. Angrist and Lavy used those sharp accidents of arithmetic to estimate class-size effects in Israeli primary schools (Angrist & Lavy, 1999).

The results landed between the poles. Reading scores in grades 4 and 5 rose as classes shrank (Angrist & Lavy, 1999). In grade 5 the gain was on the order of 0.2 SD for a reduction of around eight pupils, with smaller effects in maths and in the younger grade (Angrist & Lavy, 1999). The study became one of the founding designs of modern education economics, and its message matches STAR’s shape: meaningful reductions, in the right conditions, move achievement — though not everywhere and not uniformly.

Read as a set, the pattern is coherent. True experiment: a large early-grade cut works. Natural experiment: large cuts driven by a threshold work, moderately. Natural variation: everyday wobble in class size does nothing detectable. The dose matters, the grade matters, and the baseline matters.

California: the scale-up warning

In 1996, California decided to buy the STAR result for everyone. The state pushed kindergarten-to-grade-3 classes down to 20 or fewer, almost overnight, at a cost that settled above a billion dollars a year. Jepsen and Rivkin’s evaluation is the essential study of what happened next (Jepsen & Rivkin, 2009). Achievement did rise where classes shrank — by roughly 0.06 to 0.10 SD for a ten-student reduction in grade 3 — real, but well short of what the Tennessee arithmetic promised.

The shortfall had a mechanism, and it was hiring. Shrinking every class in a state means staffing tens of thousands of new classrooms at once. Schools filled them with novice and not-yet-certified teachers, and the least experienced arrivals concentrated in the poorest schools — the very places the policy was meant to help most. The gains from smaller classes were partly offset by the dip in average teacher experience, especially in high-poverty schools (Jepsen & Rivkin, 2009).

The catch

Class-size cuts are a labour-market event, not just a pedagogy decision. Cut every class at once and you must hire at once. California’s evidence is that the resulting wave of novice teachers eats a large share of the gain, worst in the schools with least hiring power (Jepsen & Rivkin, 2009). The pilot is not the policy.

None of this says STAR was wrong. It says STAR was a pilot: scarce, monitored, staffed from the existing teacher pool. Statewide policy competes with itself for teachers. Any government planning a class-size cut inherits that constraint on day one, and the honest plan phases the cut slowly enough for teacher supply to keep up.

The cost-effectiveness question

Now the uncomfortable arithmetic. Moving a class of 23 to 15 means roughly half again as many classrooms, teachers and salaries, forever. Against that price, the experiment bought about 0.2 SD, concentrated in the early grades and in disadvantaged groups (Krueger, 1999). Whether that is a good buy depends entirely on what the same money could buy instead — and on which children receive it.

The evidence in this article sets the terms of that comparison honestly. Small trims are money burned: the null covers them (Hoxby, 2000). Universal, rapid cuts dilute themselves through the hiring channel (Jepsen & Rivkin, 2009). The strongest case for spending sits where STAR’s own effects sat: large reductions, in the earliest grades, aimed at disadvantaged cohorts, phased to protect teacher quality. And the adult follow-ups strengthen that targeted case considerably — an intervention judged “faded” on grade-8 tests was still raising college entry a decade later (Chetty et al., 2011). Budget models that stop at test scores quietly discard the best part of the return.

Tennessee STAR (experiment) ≈0.20 Israel · Maimonides’ rule ≈0.18 California CSR (at scale) ≈0.06–0.10 Connecticut (Hoxby) ≈0.00 precise enough to rule out even modest effects 0 0.1 0.2 0.3 Approximate effect of smaller classes on achievement (SD units) © 2026 FUTURE PROOF™
Figure 3. Four landmark estimates on one axis, each for a reduction of roughly 7 to 10 pupils, so magnitudes are broadly comparable: the STAR experiment (Krueger, 1999), the Israeli natural experiment (grade-5 reading) (Angrist & Lavy, 1999), California’s statewide rollout (Jepsen & Rivkin, 2009) and Connecticut’s natural variation (Hoxby, 2000). Values are approximate central estimates; contexts, grades and outcome tests differ across studies. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What the evidence doesn’t show

The class-size literature is unusually strong, which makes its boundaries easy to state and important to keep.

  • Fade-out is not the last word. Test effects shrink after grade 3, yet college-test taking, college entry and adult outcomes still moved — so grade-8 scores understate the return (Chetty et al., 2011).
  • The null covers modest trims. Hoxby’s precise zero applies to ordinary fluctuations in ordinary ranges; it neither tests nor refutes a large early-grade cut (Hoxby, 2000).
  • Scale broke the pilot. STAR’s magnitude did not survive a rushed statewide rollout, and no honest projection should assume it will (Jepsen & Rivkin, 2009).
  • This is K–3 evidence, mostly. The experimental result is early-grade; older-grade evidence is quasi-experimental, thinner and mixed (Angrist & Lavy, 1999).
  • Mechanisms remain argued. More attention per pupil, fewer disruptions, earlier detection of struggle — the experiment measures the bundle, not the ingredients (Mosteller, 1995).
  • Aides were not a substitute. The cheaper adult-in-the-room arm delivered approximately nothing, a null with its own budget lesson (Finn & Achilles, 1990).

Where the evidence stops

  1. 1Fade-out is not the last word
  2. 2The null covers modest trims
  3. 3Scale broke the pilot
  4. 4This is K–3 evidence, mostly
  5. 5Mechanisms remain argued
  6. 6Aides were not a substitute
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Class size by the evidence

Held together, the studies compress into rules a school board or ministry can actually use.

Buy big cuts or none. The measured gains come from reductions of about seven to ten pupils; the precise null covers small trims (Hoxby, 2000). Shaving two pupils from every class spends real money inside the range where nothing has ever been detected.

Aim at the early grades and at disadvantage. STAR’s effects were largest in kindergarten and grade 1, and roughly double for Black students (Krueger, 1999). A targeted program puts the smallest classes where the evidence says the mechanism bites hardest.

Phase the rollout to protect teacher quality. California’s warning is quantitative: hire in a rush and novice-teacher dilution consumes much of the gain, worst in poor schools (Jepsen & Rivkin, 2009). Slower is cheaper per unit of learning.

Judge the spend on long dials, not just next spring’s scores. The follow-ups moved college-test taking and college entry years after the test effects faded (Krueger & Whitmore, 2001). Evaluation windows shorter than five years bias every class-size decision toward no.

Do not buy the discount substitute. The aide arm was the cheap version of the same theory, and it returned nothing measurable (Finn & Achilles, 1990). If the budget cannot fund real reductions where they count, spend it on interventions with their own evidence — not on a diluted imitation of this one.

Applied at Future Proof

How Future Proof Education™ applies this.

The likeliest ingredient in the small-class effect is attention per pupil — and attention is what software can multiply without a hiring wave. The AI Tutor gives every pupil individual explanation, worked feedback and unlimited patience, whatever the roster says. The Adaptive Diagnostic and teacher dashboards make a class of 28 legible the way a class of 15 is: who is stuck, who is coasting, who needs the teacher today. The Memory Coach carries practice follow-through that large classes drop. And for ministries weighing a class-size budget, the platform measures what the spend changes — so the decision runs on evidence, not on affection.

See Future Proof for schools
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 1990Finn
  • 1995Mosteller
  • 1999Krueger
  • 1999Angrist
  • 2000Hoxby
  • 2001Krueger
  • 2009Jepsen
  • 2011Chetty
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 1990–2011, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Mosteller, F. (1995). The Tennessee study of class size in the early school grades. The Future of Children 5(2): 113–127. PDF
  2. Finn, J.D., & Achilles, C.M. (1990). Answers and questions about class size: A statewide experiment. American Educational Research Journal 27(3): 557–577. PDF
  3. Krueger, A.B. (1999). Experimental estimates of education production functions. Quarterly Journal of Economics 114(2): 497–532. PDF
  4. Krueger, A.B., & Whitmore, D.M. (2001). The effect of attending a small class in the early grades on college-test taking and middle school test results: Evidence from Project STAR. Economic Journal 111(468): 1–28. PDF
  5. Chetty, R., Friedman, J.N., Hilger, N., Saez, E., Schanzenbach, D.W., & Yagan, D. (2011). How does your kindergarten classroom affect your earnings? Evidence from Project STAR. Quarterly Journal of Economics 126(4): 1593–1660. DOI
  6. Hoxby, C.M. (2000). The effects of class size on student achievement: New evidence from population variation. Quarterly Journal of Economics 115(4): 1239–1285. PDF
  7. Angrist, J.D., & Lavy, V. (1999). Using Maimonides’ rule to estimate the effect of class size on scholastic achievement. Quarterly Journal of Economics 114(2): 533–575. PDF
  8. Jepsen, C., & Rivkin, S. (2009). Class size reduction and student achievement: The potential tradeoff between teacher quality and class size. Journal of Human Resources 44(1): 223–250. PDF
For schools & ministries

Attention per pupil, without the hiring wave.

Book a 20-minute demo. We’ll show you AI tutoring that gives every pupil individual feedback, and dashboards that make a class of 28 as legible as a class of 15.

8 citations Reviewed August 2026 Open peer review welcomed