© 2026 FUTURE PROOF™
Systems & Policy · Early childhood

Preschool programs’ long-term effects, on trial.

The two most famous experiments in education promise decades of returns from two preschool years. The most rigorous modern trial found a pre-K advantage that vanished by kindergarten and turned negative by sixth grade. Both results are real — and the space between them is the entire policy question.

TL;DR

The finding: Preschool programs’ long-term effects are real and can be enormous. Perry and Abecedarian moved graduation, employment, college entry and arrest rates decades later, with estimated returns of roughly 7–10% a year. But scale is not destiny: in the strongest modern randomised trial, Tennessee’s state pre-K produced a boost that vanished by kindergarten and turned negative by grade 6.

The mechanism: Test-score boosts fade almost everywhere within a few years — even in the programs that paid off for life. What persists travels through skills tests miss: self-regulation, persistence, staying on track. Payoffs hinge on program quality, on what the alternative was, and on whether the schools that follow build on the head start or flatten it.

The product: Future Proof Education™ is built around that mechanism. The Adaptive Diagnostic measures each child’s actual skills at school entry — not just letters and numbers — and the AI Tutor carries early gains into the primary years with practice that starts where each child is, while teachers and parents watch the trajectory on live dashboards.

In this article

  1. 01Two experiments from another era
  2. 02The fade-out that fooled everyone
  3. 03The adult ledger
  4. 04Head Start, the national middle case
  5. 05Tennessee, the modern trial
  6. 06Both are real: reconciling the record
  7. 07What buys a payoff
  8. 08What the evidence doesn’t show
  9. 09Reading the evidence like a ministry
© 2026 FUTURE PROOF™
The route. 9 sections, from “Two experiments from another era” to “Reading the evidence like a ministry”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

No education policy argument leans harder on two small studies. When a minister says preschool pays for itself many times over, the claim traces back to a few dozen children in 1960s Michigan and 1970s North Carolina, followed for decades. The returns estimated from those cohorts — on the order of 7 to 10 percent a year — rival the stock market (Heckman et al., 2010).

Then there is the other result. Tennessee ran its state pre-K program through a genuine lottery, the strongest modern randomised test of scaled public preschool. The pre-K children left the program ahead. They lost the lead by the end of kindergarten. By sixth grade they scored below the control group, with more discipline referrals besides (Durkin et al., 2022).

It is tempting to pick one result and dismiss the other. That would be a mistake: both come from strong designs, honestly analysed. This article lays out the whole record: the famous model programs, the fade-out pattern, the national middle case of Head Start, and the sobering Tennessee trial. Then it asks what separates early investments that compound from early investments that evaporate.

Two experiments from another era

The HighScope Perry Preschool Project randomised 123 children in Ypsilanti, Michigan, between 1962 and 1967. All were Black, from low-income families, with low measured IQ at entry. The program itself was two school years of half-day preschool, run by certified teachers with small groups, plus a weekly home visit. The control group got nothing — in 1962, nothing was the normal alternative (Schweinhart et al., 2005).

The Carolina Abecedarian Project, begun in 1972, went further. It enrolled 111 infants from low-income families and provided full-day, year-round educational childcare from a few months old until age five — an order of magnitude more contact time than Perry (Campbell et al., 2002).

Both were tiny, expensive and obsessively documented. That last property is why they matter: the researchers kept following the children for decades. Perry’s cohort was interviewed at 19, 27 and 40. Abecedarian’s was assessed at 12, 15, 21 and beyond. Almost nothing else in education has an evidence trail that long.

The fade-out that fooled everyone

The early Perry results made headlines for the wrong reason. IQ scores jumped — a gap of roughly 11 points at the end of the program. Then the gap shrank every year, and by around age 10 it was essentially gone (Schweinhart et al., 2005). Critics reasonably called the program a failure. The gains it was designed to produce had washed out.

Fade-out turned out to be the rule, not a Perry quirk. Across roughly 84 program evaluations since the 1960s, the average effect at the end of a preschool program is about 0.35 standard deviations — and it shrinks steadily in the years that follow (Duncan & Magnuson, 2013). On achievement tests, the treated and untreated children converge. This is one of the most replicated patterns in education research.

If the story had ended there, preschool would be remembered as an expensive way to buy a temporary test bump. The story did not end there, because the Perry and Abecedarian teams kept counting — and switched from measuring test scores to measuring lives.

The adult ledger

By age 40, the Perry treatment group looked different where it counts. Roughly 77% had finished high school, against about 60% of the controls. Around 76% were employed, against about 62%. The arrest record tilted hardest: about 36% of the program group had been arrested five or more times, against roughly 55% of the controls (Schweinhart et al., 2005). Earnings ran higher too.

Abecedarian’s ledger, read at age 21, pointed the same way with different line items. About 36% of the treated group had enrolled in a four-year college, against roughly 14% of controls. Reading and maths advantages were still visible in early adulthood, and teenage parenthood ran lower (Campbell et al., 2002). Unlike Perry, some of Abecedarian’s cognitive advantage never fully faded — plausibly because it started in infancy.

Heckman and colleagues then put Perry through a proper cost-benefit audit, correcting earlier, looser claims. Counting earnings, crime, welfare and schooling costs, they estimated an annual social return of roughly 7 to 10 percent — benefit-cost ratios on the order of 7-to-1 to 12-to-1 under plausible assumptions (Heckman et al., 2010). Most of the return came from crime reduction and earnings, not test scores. The skills that persisted — self-control, persistence, social behaviour — were precisely the ones the achievement tests stopped seeing.

The number

7–10% The estimated annual social rate of return to the Perry Preschool Program — driven by crime, earnings and schooling outcomes decades later, not by the IQ gains that faded by age 10 (Heckman et al., 2010).

≈+11 ≈0 adult outcomes diverge decades later 0 +4 +8 +12 3 4 5 6 7 8 10 Perry Preschool IQ gap by age in years (points, approximate) © 2026 FUTURE PROOF™
Figure 1. The famous fade-out, schematic. Perry’s treatment-control IQ gap peaked at roughly 11–12 points at the end of the program and was essentially gone by around age 10 (Schweinhart et al., 2005) — the pattern that convinced early critics the program had failed, decades before the adult outcomes reported in Figure 2. Curve interpolates approximate published test points; the same fade shows up across the wider preschool literature (Duncan & Magnuson, 2013). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Head Start, the national middle case

Perry and Abecedarian were boutique experiments. Head Start is a nation-scale program that has served millions of low-income American children since 1965. It sits between the model programs and a modern state rollout — cheaper and more variable than Perry, older and better studied than any state pre-K. Its evidence splits cleanly along the same line as Perry’s.

On test scores, the verdict is sobering. The Head Start Impact Study randomised access for roughly 4,700 children — the program’s only national experiment. Offers of a Head Start place produced modest gains in the preschool year, mostly in literacy. By the end of third grade, the treatment and control groups were essentially indistinguishable on achievement (Puma et al., 2012). Fade-out, again, at national scale.

On life outcomes, the verdict flips. Deming compared siblings where one attended Head Start and the other did not, following them into adulthood. Test gains faded by adolescence, exactly on script. Yet by young adulthood the Head Start siblings scored about 0.23 standard deviations higher on a composite of graduation, college attendance, health and staying out of trouble. That is roughly 80% of the benefit of the model programs, at a fraction of the cost (Deming, 2009).

Hold the two Head Start results together and a rule emerges: the achievement test at age 8 is a poor forecaster of the life ledger at 25. Whatever the lasting asset is, third-grade scores do not price it.

Tennessee, the modern trial

Tennessee’s Voluntary Pre-K program is what scaled public preschool actually looks like: a full school-day program, thousands of classrooms, ordinary district staffing, real budgets. Oversubscribed sites ran admission lotteries, and Lipsey, Farran and Durkin followed about a thousand randomised children — the closest thing to a Perry-quality design ever run on a modern state program (Lipsey, Farran & Durkin, 2018).

The children left pre-K clearly ahead — roughly a third of a standard deviation on early achievement. The lead did not survive contact with school. By the end of kindergarten the control children had caught up. By second and third grade the lines had crossed on several measures, with the pre-K group slightly behind (Lipsey, Farran & Durkin, 2018).

The sixth-grade follow-up hardened the reversal. The randomised pre-K group scored lower on state achievement tests and logged more discipline referrals and more special-education placements than the controls (Durkin et al., 2022). The effects are small. They are also negative, in the strongest design available, at the scale ministers actually buy. No honest reading of the early-childhood literature gets to skip this study.

The catch

A randomised lottery, a scaled state program, and a negative sign by grade 6 (Durkin et al., 2022). Tennessee did not show that preschool cannot work. It showed that “preschool” is not one thing — and that scale, quality and what follows the program can flip the outcome.

program group control group Perry: finished high school 77% 60% Perry: employed at 40 76% 62% Perry: 5+ arrests by 40 36% 55% Abecedarian: 4-year college 36% 14% 0% 25% 50% 75% 100% Adult outcomes, program vs control (percent, approximate) © 2026 FUTURE PROOF™
Figure 2. The adult ledger the fade-out hid. Perry’s treatment group at age 40: more likely to have finished high school and to be employed, and far less likely to be repeatedly arrested — note that for the arrests row, shorter is better (Schweinhart et al., 2005). Abecedarian’s group at 21: two and a half times more likely to have entered a four-year college (Campbell et al., 2002). Approximate percentages from small samples (123 and 111 children); confidence intervals are wide. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Both are real: reconciling the record

So the two headline results stand: 7–10% annual returns from 1960s Michigan, and a negative sign from 2010s Tennessee. Four differences carry most of the reconciliation.

The counterfactual moved. Perry’s control children stayed home in an era with almost no alternatives. A modern control group attends Head Start, subsidised childcare or an informal arrangement. Preschool evaluations increasingly measure “this program versus other care”, not “preschool versus nothing” — one reason recent programs show smaller initial effects than older ones (Duncan & Magnuson, 2013).

Quality and dosage differ by an order of magnitude. Perry ran tiny groups with certified teachers and weekly home visits; Abecedarian ran five years of full-day enrichment from infancy (Campbell et al., 2002). A scaled state program spends far less per child and varies enormously classroom to classroom. Averaging over that variation can average away the payoff.

What follows the program differs. Fade-out is fastest when the receiving school re-teaches what pre-K children already know, at one pace, from the same starting line. The Tennessee team themselves point in this direction: a head start only compounds if primary school builds on it (Lipsey, Farran & Durkin, 2018).

The outcomes measured differ. Perry looks transformative on lives and unremarkable on tests; Tennessee has so far been judged mostly on tests. That is not a defence of Tennessee — discipline referrals moved the wrong way too (Durkin et al., 2022) — but it cautions against reading any program’s obituary at grade 3 (Deming, 2009).

What buys a payoff

Strip the record to what correlates with lasting benefit and a short list survives. Intensity and duration: Abecedarian’s five full-time years left marks that Perry’s two half-days-plus-visits years did not, including cognitive gains that never fully faded (Campbell et al., 2002). Skilled adults and small groups, consistently. Content aimed at how children learn at that age — structured play, language-rich interaction, self-regulation practice — rather than a compressed first grade.

Targeting matters too. Every large payoff on record comes from children with the least at home; benefits shrink as the alternative improves (Deming, 2009). And the handoff decides retention: gains persist where the next classroom starts from each child’s actual level instead of the grade’s assumed one (Lipsey, Farran & Durkin, 2018).

None of this is exotic. It is expensive, per child, and it resists dilution — which is exactly why the model programs and the scaled ones keep producing different answers.

≈+0.32 ≈0 ≈−0.06 ≈−0.10 the boost reverses +0.3 +0.2 +0.1 0 −0.1 end of pre-K end of K grade 3 grade 6 Tennessee pre-K: treatment effect on achievement (SD, approximate) © 2026 FUTURE PROOF™
Figure 3. The Tennessee trajectory, schematic. Randomised pre-K attendees left the program roughly 0.3 SD ahead, were caught by the end of kindergarten (Lipsey, Farran & Durkin, 2018), and by grade 6 sat slightly below the control group on state achievement tests (Durkin et al., 2022). Values are approximate composites across measures; the grade-6 deficit is small but statistically credible in the randomised sample. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
The rate of return to the HighScope Perry Preschool Program. Heckman, Moon, Pinto, Savelyev & Yavitz, Journal of Public Economics, 2010

What the evidence doesn’t show

The early-childhood literature is long on conviction in both directions. The record itself supports neither triumph nor dismissal without conditions.

  • Two cohorts carry the famous returns. Perry and Abecedarian together enrolled 234 children, in specific communities, half a century ago. The return estimates are careful, but they are estimates from tiny, era-bound samples (Heckman et al., 2010).
  • Fade-out on tests is near-universal. No one should promise lasting achievement gaps from one pre-K year; the pooled record shows initial effects shrinking steadily after exit (Duncan & Magnuson, 2013).
  • Sleeper effects are not guaranteed. Head Start’s adult payoffs emerged from quasi-experimental designs, not the national RCT — and Tennessee’s RCT shows follow-up can also reveal harm, not hidden benefit (Durkin et al., 2022).
  • The mechanism is inferred, not proven. Self-regulation and related skills are the leading explanation for persistence, but no trial has randomised the mechanism itself (Deming, 2009).
  • Universal programs are thinner evidence. The strongest payoffs come from targeted programs for the least-resourced children; extrapolating those returns to universal rollouts outruns the data (Duncan & Magnuson, 2013).
  • Tennessee is one state’s implementation. The reversal indicts that program as run, and warns about scale — it does not prove every scaled pre-K fails (Lipsey, Farran & Durkin, 2018).

Where the evidence stops

  1. 1Two cohorts carry the famous returns
  2. 2Fade-out on tests is near-universal
  3. 3Sleeper effects are not guaranteed
  4. 4The mechanism is inferred, not proven
  5. 5Universal programs are thinner evidence
  6. 6Tennessee is one state’s implementation
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Reading the evidence like a ministry

For a school system deciding what to fund, the record reads as five working rules.

Buy quality before coverage. The programs that paid off were intense, skilled and small-group; the one that backfired was scaled and variable. Expanding seats faster than quality is the one move the whole record argues against (Duncan & Magnuson, 2013).

Target first. The children with the least at home gain the most, in every strong study — Perry, Abecedarian and the Head Start sibling evidence agree on this even where they agree on little else (Deming, 2009). Universal ambitions should grow outward from the children the evidence is actually about.

Design the handoff, not just the program. Fade-out lives in the transition. A primary school that measures where each child actually is, and teaches from there, is the difference between compounding a head start and re-teaching it away (Lipsey, Farran & Durkin, 2018).

Judge programs on more than grade-3 tests — but do judge them. Adult ledgers redeemed Perry and Head Start after their test effects died (Schweinhart et al., 2005). Tennessee shows the opposite ending exists too (Durkin et al., 2022). Track behaviour, self-regulation and progression alongside achievement, and keep tracking after exit.

Embed the lottery. Every oversubscribed program is a free experiment. Tennessee’s willingness to randomise is why anyone knows the truth about it; that honesty should be the norm, not the scandal (Lipsey, Farran & Durkin, 2018). Fund the follow-up for decades — the entire Perry dividend was invisible at year five (Heckman et al., 2010).

Applied at Future Proof

How Future Proof Education applies this.

The fade-out literature has one operational message: early gains survive only when the next classroom builds on them. Future Proof Education™ is built for that handoff. The Adaptive Diagnostic measures each child’s actual skills at school entry — early literacy, numeracy and the self-regulation signals the long-run studies price — so no head start gets re-taught into the average. The AI Tutor then keeps every child on an individual gradient through the primary years, while the Knowledge Map shows teachers exactly which foundations each cohort still owes. Parent and ministry dashboards keep the trajectory visible long after program exit — because the Perry lesson is that the payoff shows up in the follow-through, not the launch year.

See the platform
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 2002Campbell
  • 2005Schweinhart
  • 2009Deming
  • 2010Heckman
  • 2012Puma
  • 2013Duncan
  • 2018Lipsey
  • 2022Durkin
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 2002–2022, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Heckman, J.J., Moon, S.H., Pinto, R., Savelyev, P.A., & Yavitz, A. (2010). The rate of return to the HighScope Perry Preschool Program. Journal of Public Economics 94(1–2): 114–128. PDF
  2. Durkin, K., Lipsey, M.W., Farran, D.C., & Wiesen, S.E. (2022). Effects of a statewide pre-kindergarten program on children’s achievement and behavior through sixth grade. Developmental Psychology 58(3): 470–484. PDF
  3. Schweinhart, L.J., Montie, J., Xiang, Z., Barnett, W.S., Belfield, C.R., & Nores, M. (2005). Lifetime effects: The High/Scope Perry Preschool study through age 40. Monographs of the HighScope Educational Research Foundation, 14. Ypsilanti: HighScope Press. PDF
  4. Campbell, F.A., Ramey, C.T., Pungello, E., Sparling, J., & Miller-Johnson, S. (2002). Early childhood education: Young adult outcomes from the Abecedarian Project. Applied Developmental Science 6(1): 42–57. PDF
  5. Duncan, G.J., & Magnuson, K. (2013). Investing in preschool programs. Journal of Economic Perspectives 27(2): 109–132. DOI
  6. Deming, D. (2009). Early childhood intervention and life-cycle skill development: Evidence from Head Start. American Economic Journal: Applied Economics 1(3): 111–134. DOI
  7. Puma, M., Bell, S., Cook, R., Heid, C., et al. (2012). Third Grade Follow-up to the Head Start Impact Study: Final Report. OPRE Report 2012-45. Washington, DC: U.S. Department of Health and Human Services. PDF
  8. Lipsey, M.W., Farran, D.C., & Durkin, K. (2018). Effects of the Tennessee Prekindergarten Program on children’s achievement and behavior through third grade. Early Childhood Research Quarterly 45: 155–176. PDF
See it in a classroom

Make the head start stick.

Book a 20-minute demo. We’ll show you the Adaptive Diagnostic measuring school-entry skills, the AI Tutor building on each child’s actual level through the primary years, and dashboards that keep every cohort’s trajectory visible.

8 citations Reviewed August 2026 Open peer review welcomed