© 2026 FUTURE PROOF™
Systems & Policy · Accountability

High-stakes testing effects: what the research found

For two decades, school systems bet that testing every child and attaching consequences would raise achievement. The research verdict is in, and it refuses to pick a side: real maths gains, a reading line that never moved, and a long ledger of inflated scores, gamed thresholds and outright cheating. Both halves are true.

TL;DR

The finding: The research on high-stakes testing effects points both ways at once. NCLB-style accountability raised maths achievement by roughly a quarter of a standard deviation in fourth grade — and left reading essentially flat (Dee & Jacob, 2011). The gains are real. So are the distortions: score inflation, attention rationed to “bubble” students, and cheating in roughly 4 to 5 percent of classrooms (Jacob & Levitt, 2003).

The mechanism: Stakes make schools respond to whatever is measured — exactly as intended, and exactly as feared. Where the measure aligned with real skill, effort produced learning, mostly in maths. Where a shortcut existed — drilling test formats, focusing on children near the pass mark, reclassifying weak students — schools took it, and measured scores rose faster than actual learning (Koretz, 2008).

The product: Future Proof Education™ builds the measurement layer accountability always needed: an Adaptive Diagnostic that draws from large item banks and resists format drilling, growth tracking for every child rather than proficiency counts, and dashboards that let schools and ministries see learning — not test theatre.

In this article

  1. 01The deal behind the tests
  2. 02What NCLB did to maths
  3. 03Reading: the line that stayed flat
  4. 04The verdict across systems
  5. 05Score inflation: when the ruler bends
  6. 06The distortions: bubbles, exclusions, cheating
  7. 07Weighing the ledger
  8. 08What the evidence doesn’t show
  9. 09Accountability by the evidence
© 2026 FUTURE PROOF™
The route. 9 sections, from “The deal behind the tests” to “Accountability by the evidence”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

In January 2002, the United States signed the biggest experiment in school governance ever run. No Child Left Behind required every state to test every child in reading and maths, every year, from third grade to eighth. Results were published by school and by subgroup. Schools that missed rising proficiency targets faced escalating consequences, up to restructuring. Versions of the same bargain spread worldwide — league tables in England, minimum-competency regimes across US states, national assessments almost everywhere.

The bet rested on a simple theory. Schools know more about their own performance than anyone above them can see. Measure honestly, attach stakes, and effort will flow toward learning. The counter-theory was just as simple. Schools will respond to the measure, not the mission — and a test is a very narrow measure of a school.

Twenty years of research later, both theories turned out to be correct, often inside the same building. This article lays out that double verdict. First, what accountability did to achievement — the maths gains and the reading flatline. Then the part the test-score headlines hid: inflation, bubble students, strategic exclusion, and cheating. Finally, what a school system would build if it wanted the gains without the theatre.

The deal behind the tests

It helps to be precise about what “high stakes” means, because the stakes were mostly aimed at adults. Under NCLB and its state-level predecessors, consequences attached to schools and educators: public ratings, funding rules, reconstitution, in some districts pay and promotion. For students, stakes varied — some places tied grade promotion or graduation to scores, many did not.

Economists saw a classic incentive-design problem, and said so early. When you reward a measured proxy, you get more of the proxy — whether or not you get more of the thing it stands for. The research program that followed was therefore two-tracked from the start. Track one: did real learning rise? Track two: what else rose?

Answering track one is harder than it looks. Scores on the accountability test itself cannot settle it, for reasons that become obvious in a moment. The credible studies lean on an audit measure — above all NAEP, the low-stakes national assessment no school can prepare for — and on comparisons between states that adopted stakes early and late (Figlio & Loeb, 2011).

What NCLB did to maths

The cleanest national estimate comes from Dee and Jacob. Their design uses a fact of history: some states had built consequential accountability systems in the 1990s, before NCLB forced the rest to follow. If accountability works, the late adopters should gain ground after 2002 on the audit test, relative to the states already treated. That is what happened — in one subject (Dee & Jacob, 2011).

Fourth-grade maths rose by roughly 0.23 standard deviations by 2007 — a large gain by policy standards, worth well over half a school year of learning. A standard deviation is the researcher’s yardstick for achievement differences; effects near 0.25 are among the largest any national reform has produced. Eighth-grade maths showed smaller, less certain gains, concentrated among lower-achieving students. Reading, at both grades, showed no detectable effect at all (Dee & Jacob, 2011).

The number

≈0.23 SD The estimated effect of NCLB-style accountability on fourth-grade maths by 2007, measured on the low-stakes national audit test — with reading effects statistically indistinguishable from zero (Dee & Jacob, 2011).

Two details make the maths result hard to dismiss. It appears on NAEP, which teachers do not see, cannot drill, and whose results carry no consequences — so it is not inflation. And it is largest exactly where the policy aimed: younger students, disadvantaged groups, low-performing schools (Dee & Jacob, 2011). Pressure, in other words, produced some genuine teaching.

Reading: the line that stayed flat

The reading null is just as instructive, and it replicates across almost every serious study of test-based accountability (Figlio & Loeb, 2011). Why would pressure move one subject and not the other?

The most credible answer sits in how the two skills are built. Maths achievement, especially in primary school, is close to a pure school product — procedures and concepts taught in sequence, responsive to instructional time and focus. Reading comprehension is different. Past the decoding years, it depends heavily on vocabulary and background knowledge accumulated over years, at home as much as at school. A school squeezed for quick gains can add maths lessons and get maths scores. Adding “reading strategy” lessons does far less, because the binding constraint is knowledge, not strategy.

The asymmetry carries a hard lesson for every accountability designer. Test pressure is a lever on things schools can change quickly. It cannot conjure the slow inputs — and a system that punishes schools for slow-growing skills mostly teaches them to hunt shortcuts. Which brings us to what else the pressure produced.

The verdict across systems

Before the shortcuts, complete the achievement ledger. The NCLB estimate is one study of one policy. The wider record points the same way. Hanushek and Raymond compared states through the 1990s as accountability spread, and found that systems with real consequences raised achievement growth, while systems that merely published report cards did little (Hanushek & Raymond, 2005). Stakes, not information alone, moved behaviour.

Figlio and Loeb’s synthesis of the whole literature — dozens of state, national and international studies — lands on the same asymmetric verdict: modest positive effects on measured achievement, stronger and more consistent in maths than in reading (Figlio & Loeb, 2011). The pattern survives across designs and decades. Accountability is not a nothing. It is also not the transformation its architects promised, and the honest half of the gains sits almost entirely in one subject.

So far, the scoreboard reads: real but narrow gains, verified on audit tests. Now the other track — what happened on the tests that carried the stakes.

Maths, grade 4 ≈+0.23 Maths, grade 8 ≈+0.10 (less certain) Reading, grade 4 ≈0.00 (ns) Reading, grade 8 ≈0.00 (ns) 0 0.1 0.2 0.3 Approximate NCLB effects on the national audit test (SD) © 2026 FUTURE PROOF™
Figure 1. Accountability’s audited report card: a large fourth-grade maths gain, a smaller and less certain eighth-grade one, and reading flat at both grades. Approximate estimates after Dee & Jacob (2011), measured on NAEP by 2007; the reading bars are drawn at their near-zero point estimates, whose confidence intervals include zero. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Score inflation: when the ruler bends

Daniel Koretz spent a career documenting a phenomenon that should unsettle anyone who reads a school ranking: scores on a high-stakes test can rise for years while the learning they claim to measure barely moves (Koretz, 2008). The mechanism is mundane. Any test samples a small slice of a subject. Teach to the sample — its formats, its favourite topics, its predictable rubrics — and scores climb without the subject being learned. Koretz calls the result score inflation, and he documented a signature for it: gains that evaporate when the same students sit a test they were not prepared for.

The most famous case was the “Texas miracle”. Through the late 1990s, Texas’s own accountability test showed spectacular gains, celebrated nationally and used as the template for NCLB. Klein and colleagues at RAND compared those gains to the same state’s performance on the low-stakes national test. The miracle mostly vanished — Texas’s audited progress was several times smaller than its official numbers, and its racial achievement gaps were widening while the state test showed them closing (Klein, Hamilton, McCaffrey & Stecher, 2000).

Chicago provides the microscope version. Jacob studied the city’s 1996 accountability regime and found scores on the high-stakes test jumping sharply — while the same students’ results on a parallel low-stakes state test improved far less. The gap had fingerprints: maths gains clustered in the specific skills the high-stakes test sampled, and special-education placements rose as schools moved weak students out of the tested pool (Jacob, 2005).

The general rule, confirmed again and again: the tighter the stakes bind to one instrument, the less that instrument means. Any single number a school is judged by will, over time, describe the school’s test preparation more than its teaching. That is not cynicism. It is the best-replicated finding in the measurement literature (Koretz, 2008).

state test (stakes attached) audit test (no stakes) the inflation gap 0 0.2 0.4 0.6 cumulative gain (SD) 1994 1996 1998 2000 Schematic: high-stakes gains vs audited gains, Texas pattern © 2026 FUTURE PROOF™
Figure 2. Score inflation’s signature. On the test that carries stakes, scores soar; on the audit test the same students sit without preparation, gains are several times smaller. Schematic after the Texas comparison in Klein et al. (2000) and the Chicago high-stakes vs low-stakes divergence in Jacob (2005); curve magnitudes are illustrative of the documented pattern, not plotted from a single dataset. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The distortions: bubbles, exclusions, cheating

Inflation is the soft distortion. The harder ones involve choosing which children matter. NCLB judged schools on proficiency counts — the share of students clearing a fixed bar. Neal and Schanzenbach pointed out what any incentive designer would predict: under a threshold rule, the rational school concentrates on students near the bar. Children far below it cannot be rescued this year; children far above it are already counted. Both become, in the brutal shorthand teachers themselves adopted, non-bubble kids (Neal & Schanzenbach, 2010).

Chicago’s data confirmed the prediction. When the city moved to threshold-based accountability, test-score gains concentrated squarely in the middle of the achievement distribution. Students in the bottom deciles — the children the policy was named for — showed little or no improvement. The authors’ title is the finding: left behind by design (Neal & Schanzenbach, 2010).

Schools also edited the tested population itself. The accountability literature documents rises in special-education classification, strategic grade retention ahead of tested years, and suspensions that happened to fall during testing windows — even schools serving high-calorie lunches on test days (Figlio & Loeb, 2011). Each trick moves the number. None of them teaches anyone anything.

And at the far end of the spectrum sits fraud. Jacob and Levitt built a statistical detector for answer-sheet manipulation — improbable strings of identical answers, suspicious swings that vanish the next year — and ran it across Chicago’s classrooms. Their estimate: serious teacher or administrator cheating in roughly 4 to 5 percent of classrooms in any given year, with prevalence rising as accountability pressure rose (Jacob & Levitt, 2003). Years later, the Atlanta scandal — 178 educators implicated in coordinated answer-changing — showed the base rate was no Chicago quirk.

The catch

Cheating tracked incentives, not character. Classrooms most exposed to sanction thresholds cheated most, and prevalence moved when the stakes moved (Jacob & Levitt, 2003). A system that judges adults on a single manipulable number is not hiring worse people — it is manufacturing the temptation.

Weighing the ledger

So what is the honest bottom line? Start with what accountability defensibly bought. Real maths gains, largest for the disadvantaged students the policy targeted (Dee & Jacob, 2011). A norm of measuring every child and publishing every subgroup — which ended the era when struggling groups could disappear inside averages. And a research infrastructure that most school systems simply did not have before.

Against that: no detectable reading gains, official numbers that overstate true progress wherever stakes are attached, effort rationed away from the lowest achievers, and a measurable rate of adult fraud (Figlio & Loeb, 2011). The costs are not side effects of bad implementation. They are the predictable output of the design — one instrument, one threshold, high stakes, no audit.

The mature conclusion is not “testing bad” or “testing good”. It is that measurement changes what it measures, in proportion to the stakes — so the design of the measurement is everything. That principle, not nostalgia for either side of the testing wars, is what the next generation of systems should be built on.

the bubble little gain at the bottom 0 +0.1 score gain (SD) far below standard near the pass mark far above prior achievement (ordinal); gain curve schematic © 2026 FUTURE PROOF™
Figure 3. Who a proficiency threshold helps. Under pass-mark accountability, Chicago’s gains concentrated on students near the bar, with little improvement for those far below it — the group the policy was named after. Schematic after Neal & Schanzenbach (2010): the horizontal axis is ordinal and the curve illustrates the documented shape of the gains, not exact decile estimates. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Left behind by design: proficiency counts and test-based accountability. Neal & Schanzenbach, Review of Economics and Statistics, 2010

What the evidence doesn’t show

The accountability literature is large and unusually well identified. It still has edges, and honest policy needs them stated.

  • No randomized trial of accountability exists. Nations do not flip coins over school governance; the causal estimates come from staggered adoption and policy discontinuities, strong designs but not experiments (Dee & Jacob, 2011).
  • The maths-reading split is descriptive. The knowledge-based explanation for flat reading fits the data, but no study has directly tested why pressure moves one subject and not the other (Figlio & Loeb, 2011).
  • Long-run adult outcomes are thin. Whether accountability-era gains carried into graduation, earnings or wellbeing is far less studied than the test-score effects themselves (Hanushek & Raymond, 2005).
  • Inflation size varies and is hard to forecast. Audit gaps differ by state, test and decade; the Texas numbers travel as a warning, not a universal constant (Klein, Hamilton, McCaffrey & Stecher, 2000).
  • Cheating estimates are lower bounds. The detection algorithm catches crude answer-changing, not subtler coaching or exclusion — the true distortion rate is unknown (Jacob & Levitt, 2003).
  • Newer designs are less studied. Growth-based and multi-measure systems adopted after 2015 have not yet accumulated evidence of NCLB’s depth (Figlio & Loeb, 2011).

Where the evidence stops

  1. 1No randomized trial of accountability exists
  2. 2The maths-reading split is descriptive
  3. 3Long-run adult outcomes are thin
  4. 4Inflation size varies and is hard to forecast
  5. 5Cheating estimates are lower bounds
  6. 6Newer designs are less studied
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Accountability by the evidence

Twenty years of findings compress into a design manual. Every line of it follows from a study above.

Judge growth, not proficiency counts. A threshold creates bubble kids by arithmetic; measuring each child’s progress puts every child back in the incentive (Neal & Schanzenbach, 2010). If a pass bar must exist for reporting, it should never be the number consequences hang on.

Always run an audit measure. The only reason anyone knows about the Texas gap or the Chicago gap is that a low-stakes test existed alongside the official one (Klein, Hamilton, McCaffrey & Stecher, 2000). A sampled, no-stakes audit — NAEP-style — is the immune system of an accountability regime. Systems without one are choosing not to know.

Vary the instrument. Inflation feeds on predictability: fixed formats, recycled items, one test (Koretz, 2008). Broad item banks, rotating content and adaptive delivery shrink the payoff to drilling and push effort back toward the subject.

Treat reading differently. Pressure alone has never moved reading comprehension (Dee & Jacob, 2011). The evidence-consistent response is curriculum — building vocabulary and knowledge over years — with accountability tracking the inputs, not just demanding the score.

Police the pool, and lower the temperature. Exclusion tricks and answer-sheet fraud rose with stakes (Jacob & Levitt, 2003). Participation audits and forensic score checks should be routine — and consequences graduated, because the sharpest cliffs produced the worst behaviour, not the best teaching (Figlio & Loeb, 2011).

Applied at Future Proof

How Future Proof Education™ applies this.

The research says measurement improves schools only when it cannot be gamed and counts every child. That is the specification Future Proof Education is built to. The Adaptive Diagnostic assembles each assessment from large, rotating item banks — so there is no fixed format to drill and no answer sheet to doctor. The Knowledge Map tracks growth for every student against their own baseline, which means no bubble: progress at the bottom of the class counts exactly as much as progress at the bar. Teacher dashboards separate diagnosis from stakes, parents see real mastery rather than a pass label, and ministries get audit-grade cohort data without a test-prep industry growing around it.

See the measurement platform
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 2000Klein
  • 2003Jacob & Levitt
  • 2005Hanushek
  • 2005Jacob
  • 2008Koretz
  • 2010Neal
  • 2011Dee
  • 2011Figlio
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 2000–2011, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Dee, T.S., & Jacob, B. (2011). The impact of No Child Left Behind on student achievement. Journal of Policy Analysis and Management 30(3): 418–446. PDF
  2. Figlio, D., & Loeb, S. (2011). School accountability. Handbook of the Economics of Education, Vol. 3: 383–421. Elsevier. PDF
  3. Hanushek, E.A., & Raymond, M.E. (2005). Does school accountability lead to improved student performance? Journal of Policy Analysis and Management 24(2): 297–327. PDF
  4. Koretz, D. (2008). Measuring Up: What Educational Testing Really Tells Us. Harvard University Press, Cambridge MA. PDF
  5. Klein, S.P., Hamilton, L.S., McCaffrey, D.F., & Stecher, B.M. (2000). What do test scores in Texas tell us? Education Policy Analysis Archives 8(49). PDF
  6. Jacob, B.A. (2005). Accountability, incentives and behavior: The impact of high-stakes testing in the Chicago Public Schools. Journal of Public Economics 89(5–6): 761–796. PDF
  7. Neal, D., & Schanzenbach, D.W. (2010). Left behind by design: Proficiency counts and test-based accountability. Review of Economics and Statistics 92(2): 263–283. PDF
  8. Jacob, B.A., & Levitt, S.D. (2003). Rotten apples: An investigation of the prevalence and predictors of teacher cheating. Quarterly Journal of Economics 118(3): 843–877. PDF
Try the platform

Measurement without the theatre.

Book a 20-minute demo. We’ll show you adaptive assessment from rotating item banks, growth tracking for every child, and audit-grade reporting — the parts of accountability the evidence says to keep.

8 citations Reviewed August 2026 Open peer review welcomed