© 2026 FUTURE PROOF™
Inside the Classroom · Ability grouping

The ability grouping evidence, sorted.

Schools have sorted children by ability for a century, and argued about it for just as long. Sorted by form, the research record is unexpectedly tidy: separate classes do almost nothing, flexible small groups help, and moving a ready child ahead works better than nearly anything else a school can decide.

TL;DR

The finding: The ability grouping evidence splits cleanly by form. Sorting whole classes by general ability — tracking, streaming — moves achievement by roughly nothing: the best syntheses put it near zero at both elementary and secondary level. Regrouping children within a class, or across grades for one subject, helps — on the order of +0.2 to +0.45 standard deviations. And acceleration, the least used form, shows some of the largest effects in education research: roughly +0.7 against equally able age-mates.

The mechanism: A grouping label changes nothing by itself. Effects appear only when grouping changes what is actually taught — the pace, the level, the materials. Between-class tracks usually leave the lesson untouched and fix the label for years. Flexible, subject-specific groups that are re-formed as children move are the version the evidence pays.

The product: Future Proof Education™ builds the version that works: an Adaptive Diagnostic finds each child’s current level in each skill, the AI Tutor teaches at exactly that level, and teacher dashboards regroup children on live evidence — grouping without the labels, acceleration without the wait.

In this article

  1. 01A century of sorting children
  2. 02The forms, because the form decides
  3. 03The elementary synthesis
  4. 04The secondary verdict
  5. 05Small groups, moving groups
  6. 06Acceleration: the awkward winner
  7. 07The equity ledger
  8. 08What the evidence doesn’t show
  9. 09Grouping by the evidence
© 2026 FUTURE PROOF™
The route. 9 sections, from “A century of sorting children” to “Grouping by the evidence”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Every school sorts children, whether it admits it or not. A primary teacher builds reading groups in week one. A head teacher decides whether Year 7 will be streamed or mixed. A ministry sets the age at which pupils split into academic and vocational tracks. Parents, meanwhile, ask the only question that matters to them: will my child learn more, or less, because of where they were placed? Few school decisions carry more feeling. Fewer still are argued with less precision.

The research base is enormous and old. Studies of ability grouping date back to the 1910s, and by the 1980s there were hundreds of them, plus meta-analyses — studies that pool many studies into one estimate — and then syntheses of the meta-analyses themselves. The most recent of these second-order reviews pooled the results of thirteen meta-analyses of grouping and six of acceleration, covering a century of work (Steenbergen-Hu et al., 2016). Very few questions in education have been measured this many times.

Yet the public argument stays stubbornly binary: group or don’t. That framing is the main casualty of the evidence. The research does not return one answer for “grouping”. It returns different answers for different forms — and the differences are large enough to reverse the policy conclusion. Sorted by form, the century of data is surprisingly tidy. This article walks through it in that order.

The forms, because the form decides

Start with the vocabulary, because everything downstream depends on it. Between-class grouping — tracking in the United States, streaming or setting elsewhere — assigns children to separate classes by ability. In its comprehensive form, a test or a judgment sorts children into high, middle and low classes for the whole day, often for years. Researchers call the classic version XYZ grouping (Kulik & Kulik, 1992). This is the form most people mean by “ability grouping”, and the form most of the heat is about.

The other forms are smaller and quieter. Regrouping re-sorts children across classes for one subject only, keeping mixed classes the rest of the day. Cross-grade grouping — the Joplin plan is the famous example — regroups across age levels by reading level, so a strong eight-year-old and an average ten-year-old share a lesson pitched at both (Slavin, 1987). Within-class grouping forms small groups inside one classroom, usually for maths or reading. Gifted programs place high achievers in enriched or accelerated classes. And acceleration moves a ready child through the curriculum faster — a skipped grade, early entry, or a higher-grade class in a single subject (Kulik & Kulik, 1992).

Why belabour the taxonomy? Because the forms carry opposite verdicts, and any average across them is a statistical fiction. The null result that dominates headlines belongs almost entirely to the first form. The wins belong to the rest. Collapse them into one number and you get a small positive blur that supports nobody’s policy (Steenbergen-Hu et al., 2016).

The elementary synthesis

The modern reading of the literature begins with Robert Slavin’s best-evidence syntheses. The method combines the discipline of meta-analysis with old-fashioned standards: only studies with real comparison groups, reasonable duration and fair measures get pooled, and each study is also read on its own terms (Slavin, 1987). For elementary schools, the synthesis sorted the evidence by the same taxonomy as above. The results split immediately.

Comprehensive between-class plans — separate classes by general ability — produced a median effect close to zero. Not modestly positive, not clearly harmful: nothing, across decades of studies (Slavin, 1987). The two forms that did work were the modest ones. Within-class grouping in mathematics carried a median effect of roughly +0.34 standard deviations. Cross-grade regrouping by reading level, the Joplin plan, carried roughly +0.45 (Slavin, 1987). In plain terms: sorting classes did nothing, while re-sorting children for a specific subject, at a specific level, produced some of the better effects in elementary schooling.

Slavin drew the obvious conclusion and named the conditions. Grouping paid when it was organised around one subject rather than a global label, when it measurably narrowed the spread of what had to be taught, when placements were re-assessed and changed often, and when teachers actually varied the pace and materials to fit the group (Slavin, 1987). Remove those conditions and the effect went with them. The label alone taught nobody anything.

The secondary verdict

Three years later Slavin ran the same method over secondary schools, where between-class tracking is the default arrangement in much of the world. The synthesis covered 29 controlled studies of tracked versus untracked schools. The result was, again, a median effect of essentially zero — and it stayed near zero for high, middle and low achievers alike (Slavin, 1990). Tracking, as typically practised, neither lifted the top nor sank the bottom on achievement tests. It simply failed to matter.

Hold on to how strange that null is, because it cuts against both camps. Advocates promise that able children are freed to fly in selective classes; the data show no such lift. Critics warn that low tracks devastate the children in them; the achievement data, on average, show no such collapse either (Slavin, 1990). What the null really says is that relabelling a classroom does not change what happens inside it. Most tracked schools taught the same curriculum, from the same textbooks, at nearly the same pace — just to more uniform audiences.

The Kuliks, working with looser inclusion rules, read the same literature a shade more warmly: they found small positive effects when grouped classes actually adjusted the curriculum, especially for high achievers in enriched classes (Kulik & Kulik, 1992). The two camps argued for years, but their disagreement was narrow. Both found that grouping with an untouched curriculum produced approximately nothing. Both found that effects appeared when the teaching changed. The fight was over how often schools actually change it.

Between-class plans (elem.) ≈0.00 Between-class plans (sec.) ≈0.00 Within-class maths groups ≈+0.34 Cross-grade reading (Joplin) ≈+0.45 0 0.1 0.2 0.3 0.4 0.5 Median effect size (SD), Slavin best-evidence syntheses © 2026 FUTURE PROOF™
Figure 1. The best-evidence syntheses in one chart. Comprehensive between-class plans sit at a median of about zero in elementary (Slavin, 1987) and secondary schools (Slavin, 1990), while the two forms that change the day’s actual teaching — within-class maths groups and cross-grade reading regrouping — carry medians of roughly +0.34 and +0.45 (Slavin, 1987). Values are reported medians across pooled studies; treat magnitudes as approximate. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Small groups, moving groups

The within-class result did not stay Slavin’s alone. A decade later, Lou and colleagues pooled 66 studies of small-group teaching inside classrooms. Simply teaching in small groups, versus whole-class teaching with no grouping, was worth about +0.17 standard deviations. Where the comparison was between grouping by skill and grouping at random, skill-matched groups added roughly +0.12 more (Lou et al., 1996). These are modest numbers, but they attach to an arrangement that costs nothing and excludes nobody.

The moderators in that synthesis are the practical gold. Groups of three or four members outperformed larger ones. Effects grew sharply when teachers used materials actually adapted to the group, and shrank toward zero when the groups just sat together over the same worksheet (Lou et al., 1996). And the subgroup pattern complicates any simple sorting story: low achievers learned more in mixed-skill groups, average achievers in matched ones, while high achievers did about equally well in both (Lou et al., 1996). One arrangement does not fit all children — which is itself an argument for keeping arrangements loose.

Set the Kuliks’ program ladder next to this and the whole literature aligns. Re-labelled XYZ classes: about +0.03. Within-class groups: about +0.25. Cross-grade grouping: about +0.30. Enriched classes for gifted students: about +0.41. Accelerated classes: about +0.87 (Kulik & Kulik, 1992). The ladder climbs in exact step with how much the curriculum itself changes. At the bottom, new class lists and the same lesson. At the top, genuinely different teaching, delivered sooner.

XYZ / tracked classes ≈+0.03 Within-class groups ≈+0.25 Cross-grade grouping ≈+0.30 Enriched gifted classes ≈+0.41 Accelerated classes ≈+0.87 0 0.2 0.4 0.6 0.8 Mean effect size (SD) by program type — Kulik & Kulik (1992) © 2026 FUTURE PROOF™
Figure 2. The program ladder from the Kuliks’ meta-analytic summary: mean effects rise in step with how much the grouping changes the curriculum itself, from re-labelled XYZ classes near zero to accelerated classes at roughly +0.87 against same-age peers (Kulik & Kulik, 1992). Approximate pooled means; study pools differ by row, so read the ladder as an ordering with magnitudes attached, not five estimates of one quantity. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Acceleration: the awkward winner

The century-scale view confirms the ladder. In 2016, Steenbergen-Hu, Makel and Olszewski-Kubilius pooled the pooled: thirteen meta-analyses of grouping and six of acceleration, in two second-order meta-analyses. Between-class grouping landed at about +0.04 to +0.06 — negligible. Within-class grouping: roughly +0.19 to +0.30. Cross-grade subject grouping: about +0.26. Special grouping for gifted students: about +0.37 (Steenbergen-Hu et al., 2016). A hundred years of data, and the 1987 ordering stands almost untouched.

Then comes acceleration, and the scale changes. Accelerated students outperformed equally able students of their own age by roughly 0.70 standard deviations — among the largest pooled effects education research has produced. Compared instead with the older students they joined, the gap was about 0.09 and not statistically significant: the young entrants simply kept pace (Steenbergen-Hu et al., 2016). The Kuliks had found the same shape decades earlier — accelerates ran far ahead of age-mates and held level with their new, older classmates (Kulik & Kulik, 1992).

The number

≈0.70 SD The pooled advantage of accelerated students over equally able same-age peers, across a century of studies — while holding their own against the older students whose classes they joined (Steenbergen-Hu et al., 2016).

Acceleration is also the form schools use least, and the 2004 national report on the topic chose its title accordingly: A Nation Deceived. Reviewing the same evidence, it catalogued eighteen distinct forms of acceleration — grade-skipping, early entrance, single-subject acceleration and more — and found the common fears largely unsupported. Accelerated children were not, on the whole, socially or emotionally worse off; the practice was cheap; and refusal was usually a policy habit rather than a judgment about the child (Colangelo, Assouline & Gross, 2004). Two decades on, the gap between this evidence and everyday school practice remains one of the widest in education.

Steenbergen-Hu et al. (2016) — pooled achievement vs older classmates ≈0.09 vs same-age peers ≈0.70 Kulik & Kulik — accelerated classes vs older classmates ≈0.0 vs same-age peers ≈0.87 0 0.2 0.4 0.6 0.8 1.0 Effect size (SD): the same accelerated pupils, two comparison groups © 2026 FUTURE PROOF™
Figure 3. Acceleration’s double comparison. Measured against equally able age-mates left behind, accelerated students run roughly 0.7 to 0.9 standard deviations ahead; measured against the older students whose classes they join, the difference is near zero — they keep pace (Steenbergen-Hu et al., 2016) (Kulik & Kulik, 1992). Approximate pooled values; the two rows pool different study sets and eras. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
What one hundred years of research says about the effects of ability grouping and acceleration. Steenbergen-Hu, Makel & Olszewski-Kubilius, Review of Educational Research, 2016

The equity ledger

Achievement averages are not the whole account. The most influential critique of tracking was never mainly about test scores. Jeannette Oakes’ Keeping Track documented what life inside the tracks looked like across 25 schools: low-track classes moved slower, spent more time on drill and discipline, got less experienced teachers more often, and ran on visibly lower expectations. Placement correlated with race and class, and children knew exactly what their label meant (Oakes, 1985). None of that shows up in a median effect size, and all of it is part of the decision.

The survey evidence backs the concern with numbers. Gamoran and Mare, analysing a national longitudinal sample, asked whether tracking compensates for inequality, reinforces it, or leaves it alone. Their answer was reinforcement: high-track placement raised achievement and the chance of finishing school, low-track placement lowered both, and placement itself was tied to social origin beyond what prior achievement explained (Gamoran & Mare, 1989). Track assignment, once made, also proved sticky — movement between tracks was rare, and rarely upward.

Here is how the two literatures fit together. The experimental syntheses say the average child gains nothing from between-class sorting (Slavin, 1990). The sociological work says the arrangement still redistributes — opportunity, expectation, curriculum — and does so along social lines (Gamoran & Mare, 1989). A null average with an unequal spread is not a neutral policy. It is a policy that trades nothing, on average, for a quiet sorting of children by background.

The misread

A null average is not a clean bill of health. Between-class tracking barely moves mean achievement (Slavin, 1990), but it fixes labels early, assigns them unevenly by background, and hands the low track a thinner curriculum (Oakes, 1985). The cost lives in the distribution, not the mean.

What the evidence doesn’t show

The tidy ladder comes with honest limits. Six are worth holding on to.

  • Averages hide the deal. A near-zero median across schools does not mean each child breaks even. Within any tracked system, some children get a better year and others a worse one; the synthesis arithmetic nets them out (Slavin, 1987).
  • An ageing evidence base. Most controlled studies behind the syntheses were run before the 1990s, in US schools. Modern classrooms, curricula and setting practices are thinner ground, and recent randomised evidence is scarce (Steenbergen-Hu et al., 2016).
  • Achievement is not everything. Effects on self-concept, aspiration and friendship are measured far less often, and the measures are weaker. Small negative self-concept effects for grouped high achievers appear in some pools — a big fish, smaller pond cost (Kulik & Kulik, 1992).
  • Nulls describe typical practice. The between-class zero is a verdict on tracking as usually implemented — same curriculum, sticky labels. A setting system with genuinely adapted teaching and frequent re-assessment is closer to the regrouping literature than to the null (Slavin, 1987).
  • Selection clouds the surveys. The tracking studies with the starkest inequality findings are observational; children are not assigned to tracks at random, and statistical controls can never fully separate the track from the child (Gamoran & Mare, 1989).
  • Acceleration is for the ready. The +0.7 belongs to able, prepared, usually willing children — often volunteers. It is evidence for removing barriers when a child is ready, not for hurrying everyone (Steenbergen-Hu et al., 2016).

Where the evidence stops

  1. 1Averages hide the deal
  2. 2An ageing evidence base
  3. 3Achievement is not everything
  4. 4Nulls describe typical practice
  5. 5Selection clouds the surveys
  6. 6Acceleration is for the ready
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Grouping by the evidence

Read as one body of work, the century of research converts into a short operating manual for schools — and it is not the manual most systems follow.

Group by subject and current skill, never by general ability. The forms that pay are specific: a maths group, a reading level. The forms that fail are global: a “top set child”. Every effect in the winning half of the ladder is attached to a subject-specific placement (Slavin, 1987).

Let every label expire. Slavin’s working conditions include frequent re-assessment and easy movement (Slavin, 1987); the sociological record shows what happens without them — placements that harden into biography (Gamoran & Mare, 1989). A group a child cannot leave is a track, whatever the school calls it.

Change the teaching, not the seating. Grouping only ever works through instruction: adapted materials, adjusted pace, a narrower target (Lou et al., 1996). If two groups get the same lesson, the grouping is administrative decoration — the XYZ bar at the bottom of the ladder (Kulik & Kulik, 1992).

Prefer small and fluid over separate and permanent. Within-class groups of three or four, re-formed as evidence arrives, carry the reliable gains (Lou et al., 1996). Whole-day tracks carry the equity costs without the achievement return (Oakes, 1985).

Accelerate the ready, and audit the gates. The strongest effect in the literature belongs to the least used form (Steenbergen-Hu et al., 2016); the burden of proof now sits with refusal (Colangelo, Assouline & Gross, 2004). And wherever placement decisions are made — up or down — check who is landing where, against what evidence. That audit is the cheapest equity instrument a school owns.

Applied at Future Proof Education

Grouping without the labels.

The evidence favours grouping that is subject-specific, current and easy to leave — and punishes the fixed label. Future Proof Education™ is built on exactly that reading. The Adaptive Diagnostic locates each child’s present level in each skill, not one global “ability”. The AI Tutor then teaches at that level, so every child gets the within-class treatment with no between-class label. The Knowledge Map shows teachers which children need the same thing this week, so small groups form, do their work and dissolve. When a child is ready to move ahead in one subject, the platform moves them — acceleration without a committee. Parents see real levels instead of stream names, and school systems see the whole distribution, not just the average.

See it in the classroom
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 1985Oakes
  • 1987Slavin
  • 1989Gamoran
  • 1990Slavin
  • 1992Kulik
  • 1996Lou
  • 2004Colangelo
  • 2016Steenbergen-Hu
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 1985–2016, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Slavin, R.E. (1987). Ability grouping and student achievement in elementary schools: A best-evidence synthesis. Review of Educational Research 57(3): 293–336. PDF
  2. Slavin, R.E. (1990). Achievement effects of ability grouping in secondary schools: A best-evidence synthesis. Review of Educational Research 60(3): 471–499. PDF
  3. Kulik, J.A., & Kulik, C.-L.C. (1992). Meta-analytic findings on grouping programs. Gifted Child Quarterly 36(2): 73–77. PDF
  4. Lou, Y., Abrami, P.C., Spence, J.C., Poulsen, C., Chambers, B., & d’Apollonia, S. (1996). Within-class grouping: A meta-analysis. Review of Educational Research 66(4): 423–458. PDF
  5. Oakes, J. (1985). Keeping Track: How Schools Structure Inequality. Yale University Press. PDF
  6. Gamoran, A., & Mare, R.D. (1989). Secondary school tracking and educational inequality: Compensation, reinforcement, or neutrality? American Journal of Sociology 94(5): 1146–1183. PDF
  7. Colangelo, N., Assouline, S.G., & Gross, M.U.M. (2004). A Nation Deceived: How Schools Hold Back America’s Brightest Students. The Templeton National Report on Acceleration, University of Iowa. PDF
  8. Steenbergen-Hu, S., Makel, M.C., & Olszewski-Kubilius, P. (2016). What one hundred years of research says about the effects of ability grouping and acceleration on K–12 students’ academic achievement: Findings of two second-order meta-analyses. Review of Educational Research 86(4): 849–899. PDF
Try Future Proof Education

Every child taught at their level.

Book a 20-minute demo. We’ll show adaptive practice that meets each child exactly where they are, teacher dashboards that regroup on live evidence, and subject acceleration that happens the moment a child is ready.

8 citations Reviewed August 2026 Open peer review welcomed