© 2026 FUTURE PROOF™
Inside the Classroom · Classroom management

The classroom management evidence is a game.

Behaviour advice for teachers is mostly folklore — be firm, build rapport, survive. Yet classroom management owns one of education’s strongest experimental records: a 1969 team game that collapsed disruption within days, then kept paying off in randomized trials that followed first-graders all the way into adulthood.

TL;DR

The finding: The classroom management evidence has an unusual crown jewel: the Good Behavior Game. In its first test, disruption fell from roughly 96 percent of observed intervals to about 19 — and came back when the game was switched off. Randomized trials then followed first-graders for over a decade. Men from game classrooms had roughly half the rate of drug abuse or dependence disorders at ages 19–21.

The mechanism: Prevention, not reaction. Clear rules taught in advance, a team stake in quiet, silent marks instead of public scolding, and small shared rewards. Peer pressure — the force teachers fight all day — switches sides. The best-managed classrooms are the ones where discipline is rarely needed at all.

The product: Future Proof Education™ builds the same prevention logic into software: adaptive practice that keeps every child working in their success zone, teacher dashboards that flag drift before it becomes disruption, and training baked into every rollout — because the trials show dosage decides whether the effect survives.

In this article

  1. 01A fourth-grade classroom in Kansas
  2. 02The game, in one paragraph
  3. 03From one classroom to a behavioural vaccine
  4. 04The Baltimore trials
  5. 05What first grade predicted at twenty-one
  6. 06Prevention beats reaction
  7. 07Scaling it: the school-wide turn
  8. 08What the evidence doesn’t show
  9. 09Managing the room, by the evidence
© 2026 FUTURE PROOF™
The route. 9 sections, from “A fourth-grade classroom in Kansas” to “Managing the room, by the evidence”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Ask a hundred teachers what makes the job hard, and behaviour lands near the top of every list. Disruption eats lesson time in small, constant bites. It exhausts new teachers faster than any other part of the work, and it pushes many of them out of the profession. The advice offered in return is mostly folklore. Be firm but fair. Build relationships. Pick your battles. None of it is wrong, exactly. It is just unmoored — advice with no experiment behind it.

What the folklore rarely mentions is that classroom management has one of the strongest experimental records in education. The crown jewel of that record is a game. In 1969, a teacher in Kansas split her class into two teams and turned quiet, seated work into a group score. Disruption collapsed within days (Barrish, Saunders & Wolf, 1969). Decades later, randomized trials had linked the same simple game to lower rates of drug disorders in young adulthood (Kellam et al., 2008). Few interventions in any field carry a paper trail like that.

This article walks the trail in order. It starts in the original fourth-grade classroom and follows the game through dozens of replications. It then turns to the Baltimore randomized trials, which tracked first-graders into their twenties. It ends with the school-wide systems built on the same logic — and with the limits an honest reader should carry out of the room.

A fourth-grade classroom in Kansas

The Good Behavior Game began as a rescue, not a research program. In 1969, behaviour researchers from the University of Kansas were invited into a fourth-grade classroom that was, by any measure, out of control. Observers sat at the back and coded conduct in one-minute intervals, day after day. The numbers they recorded still startle. During the math period, pupils talked out of turn in roughly 96 percent of observed intervals. Someone was out of their seat without permission in roughly 82 percent (Barrish, Saunders & Wolf, 1969).

The fix cost nothing. The teacher divided the class into two teams. She posted a few concrete rules: stay in your seat, work quietly, no talking out of turn. When a pupil broke a rule, the pupil’s team took a mark on the board. Any team that stayed under a small limit won modest privileges — lining up first for lunch, stars on a chart, a slice of free time at the end of the day. Both teams could win. The game ran inside ordinary lessons and took no extra time to play.

Disruption collapsed. Talking-out fell from roughly 96 percent of intervals to about 19. Out-of-seat behaviour fell from about 82 percent to 9 (Barrish, Saunders & Wolf, 1969). Then the researchers did the thing that made the study a classic. They switched the game off, and the chaos returned. They switched it back on, and order returned with it. Within a single classroom, the design demonstrated cause and effect about as cleanly as a light switch.

The game, in one paragraph

The mechanics matter, because they explain the result. The game is what behaviour analysts call an interdependent group contingency. The phrase is ugly but the idea is plain: consequences apply to the team, and the team’s outcome depends on every member. That single move rewires the room’s incentives. In an ordinary classroom, a disruptive child performs for an audience. Under the game, the same audience wants the show to stop. Peer pressure — the force teachers battle all day — quietly switches sides.

Three smaller features do real work too. The rules are few, concrete and taught in advance, so no child has to guess what counts. The teacher answers a violation with a silent mark on the board, not a lecture, so misbehaviour stops earning the attention that fuels it. And the rewards are small, shared and near-immediate, which is exactly the kind of payoff young children respond to. Nothing in the package requires charisma, seniority or a gift for discipline.

Notice also what is absent. There are no points for individual children, so nobody is publicly ranked. Any team under the limit wins, so teams are never forced into rivalry. And the game is a structure rather than a personality, which is why it kept working when it left its inventor’s hands (Embry, 2002).

From one classroom to a behavioural vaccine

Replications accumulated for four decades. The game was rebuilt in other grades, other subjects, other countries, and — crucially — by ordinary teachers rather than researchers. Reviewing that record, Embry counted replications across ages and settings and argued the game had earned an unusual label: a “behavioral vaccine” (Embry, 2002). The metaphor is precise. A vaccine is cheap, brief, delivered early to everyone, and protective against disorders that appear much later. The claim was that a classroom game could satisfy that description.

The short-term half of the claim is now well quantified. A meta-analysis pooled the single-case literature — studies that, like the original, switch the game on and off while observing one classroom at a time. Across those studies the game produced consistent, large reductions in disruptive behaviour, with the biggest gains for the students at greatest behavioural risk (Bowman-Perrott et al., 2016).

But within-classroom studies, however many times repeated, cannot answer the vaccine question. Does a calmer first grade change anything about the adult who emerges thirteen years later? Answering that requires randomization, large cohorts, and patience measured in decades. That is what Baltimore supplied.

Talking out — baseline ≈96% Talking out — game on ≈19% Out of seat — baseline ≈82% Out of seat — game on ≈9% 0 25 50 75 100 Share of observed intervals with the behaviour (math period) © 2026 FUTURE PROOF™
Figure 1. The first Good Behavior Game, by the numbers. In the original fourth-grade classroom, observers coded behaviour in one-minute intervals during the math period. With the game on, talking-out fell from roughly 96 percent of intervals to about 19, and out-of-seat behaviour from about 82 percent to 9 — and both returned when the game was withdrawn (Barrish, Saunders & Wolf, 1969). Values are approximate phase means read from the original single-classroom reversal study. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The Baltimore trials

In the mid-1980s, Sheppard Kellam and colleagues at Johns Hopkins did something education research almost never does. They ran a true randomized field trial of a classroom management program — designed like a public-health trial, in partnership with the Baltimore city school district. Inside a set of urban elementary schools, first-grade classrooms were randomly assigned to play the Good Behavior Game, to receive an academic enrichment program, or to carry on as usual. Children were assigned to classrooms in a balanced way, and more than a thousand first-graders entered the study (Kellam et al., 2008).

Teachers in game classrooms were trained, mentored and observed, and the game was played through first and second grade. The short-term results matched the earlier literature: aggressive, disruptive behaviour fell, and it fell most among the boys who had started out most aggressive. A second-generation Baltimore trial, which folded the game into a broader classroom program, likewise reduced early risk behaviours by the spring of first grade (Ialongo et al., 1999).

So far, a familiar story. What made Baltimore different is that nobody stopped watching. The children were followed — not for a term, but for decades.

What first grade predicted at twenty-one

Between 2000 and 2003, the research team re-interviewed most of the original cohort as young adults, aged 19 to 21. The comparison of interest was simple: what became of children who had spent first and second grade in a game classroom, versus a standard one? The headline result concerned the men. Roughly 19 percent of men from Good Behavior Game classrooms had a lifetime drug abuse or dependence disorder. Among men from standard classrooms, the figure was roughly 38 percent (Kellam et al., 2008). A two-year game, played at age six, sat upstream of a halved addiction rate at twenty-one.

The number

≈19% vs ≈38% Lifetime drug abuse or dependence disorders at ages 19–21, for men who spent first and second grade in a Good Behavior Game classroom versus a standard one — roughly half the rate, from a two-year game played at age six (Kellam et al., 2008).

The pattern behind the headline matters as much as the headline. The benefit was concentrated among the boys who had been rated most aggressive and disruptive in first grade — the children the game was quietly built for. The same follow-up reported lower rates of regular smoking and of antisocial personality disorder among game-classroom men, with effects again strongest in that high-risk group. For women, whose base rates of these disorders were far lower, the differences were smaller and mostly not statistically reliable (Kellam et al., 2008).

One more finding deserves equal billing, because it is the practical warning label. The Baltimore team ran the trial again with a second cohort of first-graders the following year. Same game, same schools — but the teachers received less training, mentoring and support. In that cohort, the long-run effects largely faded (Kellam et al., 2008). The intervention is not the manual. It is the manual plus the dosage.

0% 10% 20% 30% 40% ≈38% ≈19% Standard classrooms Good Behavior Game Lifetime drug abuse or dependence disorders, men aged 19–21 (first Baltimore cohort) © 2026 FUTURE PROOF™
Figure 2. The long-run result. At ages 19–21, men who had spent first and second grade in Good Behavior Game classrooms showed roughly half the rate of lifetime drug abuse or dependence disorders of men from standard classrooms in the same schools (Kellam et al., 2008). Approximate rates from the first Baltimore cohort; benefits concentrated among boys rated most aggressive in first grade, and a second cohort whose teachers got less training showed much weaker effects. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Prevention beats reaction

Why does a scoreboard outperform a scolding? The deepest answer predates the Baltimore trials. Around the same time the game was invented, Jacob Kounin was filming and coding orderly and disorderly classrooms, expecting to find that orderly teachers were better at discipline. He found the opposite of what he expected. The reprimands — he called them “desists” — looked much the same in both kinds of room. What differed was everything that happened before misbehaviour: constant awareness of the whole room, smooth transitions, momentum, and work that kept every child engaged (Kounin, 1970).

Kounin’s conclusion has been replicated in spirit ever since: management is mostly prevention, and by the time a teacher is reacting, the useful moment has largely passed. The Good Behavior Game is prevention formalized. It states the rules before anyone breaks them, gives every child a stake in order before disorder starts, and strips the drama out of the response when a rule does break.

The wider literature agrees about direction, and is sober about size. A meta-analysis of classroom management interventions across dozens of studies found positive average effects on the order of 0.2 standard deviations — real, useful, and modest — with programs that also target children’s social and emotional skills among the stronger performers (Korpershoek et al., 2016). Management is not a miracle lever. It is a reliable one.

The catch

The second Baltimore cohort played the same game with less teacher training, mentoring and support — and the long-run effects largely faded (Kellam et al., 2008). Every school that adopts a behaviour program inherits this warning: the program is the training and the dosage, not the poster on the wall.

Scaling it: the school-wide turn

The modern descendant of this logic is Positive Behavioral Interventions and Supports, or PBIS — a school-wide framework rather than a classroom game. Its ingredients are recognisably the same. A school defines a handful of positive expectations, teaches them explicitly in every setting, acknowledges children who meet them, and reserves intensive support for the small group who need more. Behaviour data, such as office referrals, steer the system rather than serving as ammunition.

PBIS has its own randomized evidence. In a group-randomized trial across 37 elementary schools, schools assigned to PBIS implemented it with good fidelity, and over roughly five school years their office discipline referrals and suspensions fell reliably relative to control schools (Bradshaw, Mitchell & Leaf, 2010). For a whole-school reform measured against randomized controls, that is an unusually clean record.

Two honest footnotes belong here. Office referrals and suspensions are school records, so they partly measure adult decisions as well as child behaviour. And the effect sizes are modest — consistent with the wider management literature, not larger than it (Korpershoek et al., 2016). The school-wide turn did not amplify the classroom effect. It made the classroom effect an institution, which is a different and arguably harder achievement.

Barrish 1969 one school year Ialongo 1999 grade 1, spring Bradshaw 2010 ≈5 school years Kellam 2008 ≈14 yrs (age 21) 0 5 10 15 Years from classroom intervention to last reported follow-up © 2026 FUTURE PROOF™
Figure 3. How far the evidence follows the children. The original demonstration was a within-year reversal study (Barrish, Saunders & Wolf, 1969); the second-generation Baltimore trial reported proximal outcomes that spring (Ialongo et al., 1999); the PBIS trial tracked schools for roughly five years (Bradshaw, Mitchell & Leaf, 2010); and the first Baltimore cohort was re-interviewed roughly 14 years after first grade, at ages 19–21 (Kellam et al., 2008). Bar lengths are years of follow-up, at real relative scale. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Good behavior game: Effects of individual contingencies for group consequences on disruptive behavior in a classroom. Barrish, Saunders & Wolf, Journal of Applied Behavior Analysis, 1969

What the evidence doesn’t show

The Good Behavior Game record is remarkable, and it is routinely oversold. The honest boundary runs here.

  • It is not an academic curriculum. The game manages behaviour; it does not teach content. Across the management literature, academic gains are smaller and less consistent than behavioural ones (Korpershoek et al., 2016).
  • The famous long-run effects concentrate in one group. The halved drug-disorder rate belongs to men, and mostly to boys who began first grade highly aggressive. Effects for women, and for low-risk children, were small and often unreliable (Kellam et al., 2008).
  • Dosage decides survival. The replication cohort, whose teachers got less training and mentoring, showed much weaker long-run effects (Kellam et al., 2008). Cheap rollouts are testing a different intervention.
  • The headline trial is one city, one era. Baltimore first-graders of the mid-1980s are not every classroom. The within-classroom evidence generalizes widely (Bowman-Perrott et al., 2016); the adulthood evidence rests on far fewer settings.
  • School-record outcomes are partly adult behaviour. Referral and suspension counts respond to staff thresholds as well as child conduct, so PBIS effects on records need cautious reading (Bradshaw, Mitchell & Leaf, 2010).
  • Head-to-head comparisons are scarce. The literature says structured prevention works; it rarely says which program beats which, per hour of teacher training spent (Embry, 2002).

Where the evidence stops

  1. 1It is not an academic curriculum
  2. 2The long-run effects concentrate in one group
  3. 3Dosage decides survival
  4. 4The headline trial is one city, one era
  5. 5School-record outcomes are partly adult behaviour
  6. 6Head-to-head comparisons are scarce
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Managing the room, by the evidence

Read as one record, the literature compresses into a short operating manual — one any school can run without buying anything.

Set the rules before the misbehaviour. Few, concrete, observable, taught in advance. Every effective program in this literature begins with expectations children can recite, not standards they must infer (Kounin, 1970).

Make order a team sport. The interdependent group contingency is the engine of the game’s effect: the team shares the consequence, so peers recruit each other into calm rather than into chaos. Keep the teams winnable — every team under the limit wins (Barrish, Saunders & Wolf, 1969).

Replace commentary with counting. A silent mark on a scoreboard removes the audience and the argument at once. Save the teacher’s attention for work worth attending to — the withdrawn spotlight is half the intervention (Embry, 2002).

Start early, and treat it as prevention. The startling returns came from first grade, in the children at highest risk (Kellam et al., 2008). A behaviour program in the earliest grades is infrastructure, not crisis response.

Buy the training, not just the manual. The faded second cohort is the sharpest lesson in the whole literature (Kellam et al., 2008). Budget for coaching, observation and fidelity checks — at rollout scale, not pilot scale (Bradshaw, Mitchell & Leaf, 2010).

Expect behavioural gains first, academic gains second. Judge the program on incidents prevented and minutes of teaching recovered, at the modest but reliable effect sizes the meta-analytic record supports (Korpershoek et al., 2016).

Applied at Future Proof

How Future Proof Education applies this.

The management literature’s deepest finding is that prevention beats reaction — rooms stay calm when every child has clear expectations and work they can actually do. Future Proof Education™ builds both halves into the school day. The AI Tutor keeps each child practising in their success zone, so the idle frustration that feeds disruption never accumulates; the Adaptive Diagnostic finds the right level on day one. Teacher dashboards flag drift — slowing pace, sliding accuracy, fading engagement — before it becomes behaviour, the same early-warning logic PBIS runs on referral data. Parents see the same picture at home. And because the trials show dosage decides whether effects survive, every school and ministry rollout ships with structured teacher training and fidelity checks, not just licences.

See the platform
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 1969Barrish
  • 1970Kounin
  • 1999Ialongo
  • 2002Embry
  • 2008Kellam
  • 2010Bradshaw
  • 2016Bowman-Perrott
  • 2016Korpershoek
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 1969–2016, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Barrish, H.H., Saunders, M., & Wolf, M.M. (1969). Good behavior game: Effects of individual contingencies for group consequences on disruptive behavior in a classroom. Journal of Applied Behavior Analysis 2(2): 119–124. PDF
  2. Embry, D.D. (2002). The Good Behavior Game: A best practice candidate as a universal behavioral vaccine. Clinical Child and Family Psychology Review 5(4): 273–297. PDF
  3. Bowman-Perrott, L., Burke, M.D., Zaini, S., Zhang, N., & Vannest, K. (2016). Promoting positive behavior using the Good Behavior Game: A meta-analysis of single-case research. Journal of Positive Behavior Interventions 18(3): 180–190. PDF
  4. Ialongo, N.S., Werthamer, L., Kellam, S.G., Brown, C.H., Wang, S., & Lin, Y. (1999). Proximal impact of two first-grade preventive interventions on the early risk behaviors for later substance abuse, depression, and antisocial behavior. American Journal of Community Psychology 27(5): 599–641. PDF
  5. Kellam, S.G., Brown, C.H., Poduska, J.M., Ialongo, N.S., Wang, W., Toyinbo, P., Petras, H., Ford, C., Windham, A., & Wilcox, H.C. (2008). Effects of a universal classroom behavior management program in first and second grades on young adult behavioral, psychiatric, and social outcomes. Drug and Alcohol Dependence 95(Suppl 1): S5–S28. PDF
  6. Kounin, J.S. (1970). Discipline and Group Management in Classrooms. Holt, Rinehart & Winston, New York. PDF
  7. Korpershoek, H., Harms, T., de Boer, H., van Kuijk, M., & Doolaard, S. (2016). A meta-analysis of the effects of classroom management strategies and classroom management programs on students’ academic, behavioral, emotional, and motivational outcomes. Review of Educational Research 86(3): 643–680. PDF
  8. Bradshaw, C.P., Mitchell, M.M., & Leaf, P.J. (2010). Examining the effects of schoolwide positive behavioral interventions and supports on student outcomes: Results from a randomized controlled effectiveness trial in elementary schools. Journal of Positive Behavior Interventions 12(3): 133–148. PDF
For schools & ministries

Prevention, built into the school day.

Book a 20-minute demo. We’ll show you adaptive practice that keeps every child in their success zone, and teacher dashboards that flag drift before it becomes disruption.

8 citations Reviewed August 2026 Open peer review welcomed