Word problem solving instruction: teach the schema.
Every teacher knows the child who can compute but cannot decide what to compute. For decades the standard fix was a poster of keywords — “in all means add, left means subtract.” The research verdict is now clear. Keyword tricks quietly teach the wrong skill. Naming the problem’s underlying structure first is what actually moves word-problem scores.
The finding: Word problem solving instruction works best when it teaches children to recognise a problem’s underlying structure — its schema — before touching the numbers. Across classroom trials and a meta-analysis, schema-based instruction beats both ordinary teaching and general strategy advice, with gains from roughly half a standard deviation to well over one. Keyword tricks, the most common alternative, point to the right operation only about half the time on one-step problems and almost never on multi-step ones.
The mechanism: A word problem is a comprehension task before it is a calculation task. Strong solvers build a model of the situation, then attach numbers to it. Weak solvers grab numbers and keywords and translate them directly into arithmetic — a strategy that collapses whenever the wording runs against the operation. Schemas give children a small set of structures to map any problem onto, so the model comes first.
The product: Future Proof Education™ builds this into practice: the Adaptive Diagnostic classifies which schema a child cannot yet recognise, the AI Tutor teaches structure-first solving with faded diagrams and mixed problem types, and the Knowledge Map shows teachers and parents exactly which structures each child has mastered.
In this article
- 01Thirty-one remainder twelve
- 02The keyword shortcut and why it backfires
- 03What a schema actually is
- 04The classroom trials
- 05The pooled verdict
- 06Beyond special education
- 07What the evidence doesn’t show
- 08Word problems by the evidence
Word problems occupy a strange place in school mathematics. They are the point of the subject — the moment arithmetic meets the world — and they are also where the subject most reliably falls apart. Children who sail through a page of naked calculations stall the moment the same calculation arrives wearing a sentence. Teachers see it every week. The research has measured it for forty years, and it has also found something more useful: a way of teaching word problems that works, and a popular way that quietly makes things worse.
The best-known demonstration of the problem is a single test item. In the early 1980s, a national assessment in the United States asked thirteen-year-olds a plain question: an army bus holds 36 soldiers, 1,128 soldiers need transport, how many buses? Roughly seven in ten students did the long division correctly. Then the answers went strange. About 29 percent wrote “31 remainder 12” — buses, apparently, can come with a remainder. Another 18 percent rounded down to 31, stranding twelve soldiers at the depot. Fewer than one in four gave the answer a bus company would accept: 32 (Carpenter, Lindquist, Matthews & Silver, 1983).
The arithmetic was fine. What was missing was any sense that the numbers described a situation. Researchers came to call this the suspension of sense-making: inside a maths lesson, children learn that the game is to combine the numbers on the page and move on, not to reason about buses. It is a rational response to how word problems are usually taught — and it is the habit that good instruction has to break.
≈7 in 10 of tested thirteen-year-olds divided 1,128 by 36 correctly on the famous bus problem — yet fewer than 1 in 4 answered the actual question. The computation was never the bottleneck (Carpenter, Lindquist, Matthews & Silver, 1983).
The keyword shortcut and why it backfires
Faced with children who freeze at word problems, schools reached for a shortcut that felt like kindness. Teach the signal words. “In all” and “altogether” mean add. “Left” and “fewer” mean subtract. “Each” means multiply or divide. The lists still hang, laminated, in classrooms on every continent. The appeal is obvious: the trick converts a reading task into a matching task, and it produces quick wins on the simplest problems.
The research case against it has two parts. The first is a counting exercise. Researchers checked curriculum problems against the keyword lists, and the tricks pointed to the correct operation only about half the time even on one-step problems (Powell & Fuchs, 2018). On problems needing two or more steps, they almost never did. A strategy that performs near chance on easy items, then collapses on exactly the problems that matter, is not a strategy. It is a liability with good branding.
The second part is sharper: keywords do not just underperform, they actively mislead. Lewis and Mayer studied problems built around relational sentences such as “petrol at one station costs 5 cents less than at the other”. When the relational word ran against the needed operation — “less” in a problem that requires addition — errors multiplied, with students committing roughly two to three times as many reversal mistakes as on consistently worded problems (Lewis & Mayer, 1987). A child trained on keywords is trained into precisely this trap.
Eye-movement research later showed what separates solvers. Hegarty, Mayer and Monk compared successful and unsuccessful problem solvers reading the same items. The unsuccessful ones used what the authors called direct translation: they fixated on numbers and keywords and moved straight to arithmetic. Successful solvers spent their attention differently. They built a model of the situation first — who has what, what changed, what is being compared — and only then selected an operation (Hegarty, Mayer & Monk, 1995). Keyword teaching rehearses the losing strategy. The winning strategy is what schema instruction teaches directly.
“More” does not mean add. In compare problems with inconsistent wording, the keyword points to exactly the wrong operation, and errors multiply (Lewis & Mayer, 1987). Keyword posters train children to skip the one step — building a model of the situation — that distinguishes successful solvers (Hegarty, Mayer & Monk, 1995).
What a schema actually is
The alternative has a long pedigree. In the early 1980s, cognitive scientists mapped the deep structures underneath elementary arithmetic problems. Riley, Greeno and Heller showed that nearly every one-step additive word problem is one of three situations. A change problem: some amount goes up or down over time. A combine problem: two parts make a total. A compare problem: two amounts and the difference between them (Riley, Greeno & Heller, 1983). Multiplication brings a similar short list — equal groups, multiplicative comparison, proportion.
These structures are called schemas: a schema is a mental template for a situation type, with slots for the quantities and a plan for finding whichever slot is empty. The taxonomy explained something teachers had always sensed. Problems with identical arithmetic vary wildly in difficulty depending on the structure and on which slot is unknown (Riley, Greeno & Heller, 1983). “Joe had 3 marbles, then won 5” is easy. “Joe won 5 marbles and now has 8 — how many did he start with?” defeats far younger solvers, though both are solved by the same fact family. The difficulty was never the arithmetic. It was the mapping.
Schema-based instruction turns this analysis into a teaching sequence. Children learn to identify the problem type first, out loud — “this is a compare problem” — then fill a simple diagram matched to that type: two bars and a difference, a total split into parts, a before-and-after. Only then do they choose an operation, which by that point almost chooses itself. Number sentences fall out of the diagram rather than out of a keyword. And problems deliberately include extra, irrelevant numbers, so that grabbing whatever digits appear stops paying.
The classroom trials
The approach has been tested the hard way, in ordinary classrooms with control groups. Fuchs and colleagues ran a randomised trial across third-grade classrooms, teaching problem-type structures and — crucially — training transfer: showing children that a new cover story, different numbers, or an added irrelevant detail does not change the underlying type. Schema-taught classes outperformed controls on immediate and transfer measures of problem solving (Fuchs et al., 2004). Effects on taught problem types were large — often above a full standard deviation — and the advantage reached even novel-looking problems.
A fair sceptic would ask whether any structured attention to problem solving would have done the same. Jitendra and colleagues answered that with a stronger comparison. They put schema-based instruction head to head against general strategy instruction — the standard “read, plan, solve, check” approach, taught with equal time and care. Third graders taught schemas came out ahead on word-problem solving, on the order of half a standard deviation, and the advantage held weeks after teaching ended (Jitendra et al., 2007). Against an active, plausible rival — not against nothing — structure-first teaching still won.
That detail matters for buyers of programmes as much as for researchers. Most classroom interventions look good against business as usual. Very few keep their edge against a credible alternative given the same minutes. Schema instruction is one of the few, and the delayed tests suggest children keep the skill rather than borrowing it for the post-test.
The pooled verdict
Single trials can flatter. The pooled record is harder to argue with. Zhang and Xin gathered the word-problem intervention studies conducted with students with mathematics difficulties — the children for whom word problems are most punishing — and meta-analysed them. The average gains were large, on the order of a full standard deviation or more, though the pool mixes group designs with single-case studies of varying rigour (Zhang & Xin, 2012).
More telling than the headline number is what the strongest interventions shared. The approaches that worked taught representation: diagramming the problem’s structure before computing, exactly the schema move. What does not appear anywhere among the winning ingredients is the keyword list. The pooled literature and the classroom trials agree on the direction of travel — teach the structure, and the operations follow (Zhang & Xin, 2012).
It is worth pausing on who these studies serve. Much of the schema evidence was built with struggling learners, because that is where the need was sharpest. The result is an intervention with its strongest proof among the children ordinary teaching fails first — which is the opposite of most education research, where effects are demonstrated on the easy cases and hoped onto the hard ones.
Effective word-problem instruction: Using schemas to facilitate mathematical reasoning.Powell & Fuchs, Teaching Exceptional Children, 2018
Beyond special education
Because the strongest trials grew out of special education, schema-based instruction is sometimes filed as a remedial technique. The logic runs the other way. Powell and Fuchs, writing for practising teachers, lay out schema instruction as first-line teaching for every classroom (Powell & Fuchs, 2018). Name the structures, additive and multiplicative alike. Give each type its diagram. Teach an attack routine whose first question is the problem’s type — and retire the keyword posters that rehearse the losing strategy.
Two practical notes travel with the method. First, language still matters. A schema cannot rescue a child who does not know what “fewer” means, so word-problem teaching includes direct work on the small set of quantity words the problems depend on (Powell & Fuchs, 2018). Second, mixing matters. If Tuesday’s page is all compare problems, children stop reading again — the schema becomes the new keyword. The identification skill only develops when problem types arrive shuffled and the child must genuinely decide. The same logic drives the interleaving research on the corporate side of the learning-science literature, at future-proof.app/research-interleaving.
Worked examples fit naturally here too: studying a solved schema diagram before attempting one is a cheaper first step than solving cold, a pattern the adult-learning evidence maps in detail at future-proof.app/research-socratic-constraint. The K-12 word-problem literature reads as the child-sized version of the same lesson: structure shown first, then faded, beats struggle alone.
What the evidence doesn’t show
The schema literature earned its verdict, but the verdict has edges. Six limits are worth holding onto.
- It is not a reading cure. Schema instruction assumes the child can decode the sentences and knows the quantity vocabulary; it does not substitute for reading comprehension work, and weak readers need both (Powell & Fuchs, 2018).
- The base is mostly elementary arithmetic. The strong trials sit in Grades 2 to 7, on additive and proportion problems; evidence for algebra, geometry and secondary word problems is far thinner (Jitendra et al., 2007).
- Pooled numbers are inflated by design mix. The meta-analytic average blends group experiments with single-case studies, whose effect metrics run hot; the honest headline is “consistently large”, not one clean number (Zhang & Xin, 2012).
- Researcher support did some lifting. Many trials ran with researcher-designed materials and coached teachers; effects at unsupported, whole-district scale are less documented (Fuchs et al., 2004).
- Sense-making is necessary but not sufficient. Schemas fix the mapping problem, not the deeper habit of checking answers against reality that the bus problem exposed; that habit needs problems where blind combination visibly fails (Carpenter, Lindquist, Matthews & Silver, 1983).
- Real-world modelling remains untested. Trial problems are still school problems — tidy, solvable, single-answer; whether schema fluency transfers to genuinely messy applied modelling is largely unmeasured (Riley, Greeno & Heller, 1983).
Where the evidence stops
- 1It is not a reading cure
- 2The base is mostly elementary arithmetic
- 3Pooled numbers are inflated by design mix
- 4Researcher support did some lifting
- 5Sense-making is necessary but not sufficient
- 6Real-world modelling remains untested
Word problems by the evidence
Read together, the studies compress into a short set of moves any school can make this term — most of them free.
Take down the keyword posters. They rehearse direct translation, the strategy that distinguishes unsuccessful solvers (Hegarty, Mayer & Monk, 1995), and they mislead outright on inconsistently worded problems (Lewis & Mayer, 1987). Nothing in the evidence pays for the wall space.
Teach the problem types by name. Change, combine, compare for addition and subtraction; equal groups, comparison and proportion beyond (Riley, Greeno & Heller, 1983). Give each type its diagram, and make “what kind of problem is this?” the first question of every solution, before any number is touched (Powell & Fuchs, 2018).
Vary the surface, hold the structure. Change cover stories, positions of the unknown, and add irrelevant numbers, so children learn that structure — not vocabulary or digits — is the invariant. Explicitly teaching this transfer is what made whole-class schema instruction stick (Fuchs et al., 2004).
Shuffle problem types within practice. Blocked pages let children answer without reading. Mixed pages force the identification decision that schema instruction exists to train, and mixed-with-active-comparison is the condition under which the method proved itself (Jitendra et al., 2007).
Include problems the numbers cannot answer. A steady diet of solvable, tidy problems taught a generation to write “31 remainder 12” (Carpenter, Lindquist, Matthews & Silver, 1983). Ask questions with missing information, with surplus information, and with answers that must be rounded to reality — and give marks for saying so.
Reserve extra intensity for the children who need it. The pooled evidence is strongest for students with mathematics difficulties, taught in small groups with explicit diagrams and guided practice (Zhang & Xin, 2012). Schema instruction is first-line teaching for everyone and, at higher dosage, the best-evidenced intervention for the children furthest behind.
How Future Proof Education™ applies this.
The evidence says word-problem skill is structure recognition, so the platform treats it that way. The Adaptive Diagnostic does not just mark a word problem wrong — it classifies which schema the child failed to recognise, separating mapping errors from arithmetic slips. The AI Tutor teaches structure first: it names the problem type, shows the matching diagram, then fades the support while mixing problem types so the identification decision is always live. Practice sets deliberately include surplus numbers and unanswerable questions, because number-grabbing should never pay. The Memory Coach returns each schema at spaced intervals so recognition survives the term, and the Knowledge Map shows teachers — and parents — which structures each child has mastered, in plain language. The same picture scales from one classroom to a ministry dashboard.
See the platform →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above.
The evidence, by year
- 1983Riley
- 1983Carpenter
- 1987Lewis
- 1995Hegarty
- 2004Fuchs
- 2007Jitendra
- 2012Zhang
- 2018Powell
- Riley, M.S., Greeno, J.G., & Heller, J.I. (1983). Development of children’s problem-solving ability in arithmetic. In H.P. Ginsburg (Ed.), The Development of Mathematical Thinking. Academic Press: 153–196. PDF
- Carpenter, T.P., Lindquist, M.M., Matthews, W., & Silver, E.A. (1983). Results of the Third NAEP Mathematics Assessment: Secondary School. Mathematics Teacher 76(9): 652–659. PDF
- Lewis, A.B., & Mayer, R.E. (1987). Students’ miscomprehension of relational statements in arithmetic word problems. Journal of Educational Psychology 79(4): 363–371. PDF
- Hegarty, M., Mayer, R.E., & Monk, C.A. (1995). Comprehension of arithmetic word problems: A comparison of successful and unsuccessful problem solvers. Journal of Educational Psychology 87(1): 18–32. PDF
- Fuchs, L.S., Fuchs, D., Prentice, K., Hamlett, C.L., Finelli, R., & Courey, S.J. (2004). Enhancing mathematical problem solving among third-grade students with schema-based instruction. Journal of Educational Psychology 96(4): 635–647. PDF
- Jitendra, A.K., Griffin, C.C., Haria, P., Leh, J., Adams, A., & Kaduvettoor, A. (2007). A comparison of single and multiple strategy instruction on third-grade students’ mathematical problem solving. Journal of Educational Psychology 99(1): 115–127. PDF
- Zhang, D., & Xin, Y.P. (2012). A follow-up meta-analysis for word-problem-solving interventions for students with mathematics difficulties. Journal of Educational Research 105(5): 303–318. PDF
- Powell, S.R., & Fuchs, L.S. (2018). Effective word-problem instruction: Using schemas to facilitate mathematical reasoning. Teaching Exceptional Children 51(1): 31–42. PDF
Teach the structure, not the trick.
Book a 20-minute demo. We’ll show you schema-level diagnosis of word-problem errors, structure-first tutoring with faded diagrams, and a dashboard that tells every teacher which problem types each child can actually read.