© 2026 FUTURE PROOF™
Reading & Writing · Writing

Evidence-based writing instruction: what works.

Writing is the school subject taught most on tradition and least on trials — grammar drills, red ink, and hope. Yet the pooled experiments have ranked what actually improves student writing, with effect sizes attached. The list is short, teachable, and one classroom staple lands below zero.

TL;DR

The finding: Evidence based writing instruction has a measured league table. The Writing Next meta-analysis pooled experimental studies in grades 4–12 and ranked eleven elements: strategy instruction and summarization lead at ≈0.82, collaborative writing ≈0.75, product goals ≈0.70, sentence combining ≈0.50 — and traditional grammar drill sits at ≈−0.32, below zero. The elementary meta-analysis repeats the ordering, with strategy instruction near ≈1.0.

The mechanism: Writing overloads a novice’s working memory: ideas, words, sentences and spelling compete for the same limited attention. Strategy instruction works because it hands students the expert’s playbook for planning and revising, then fades support. Sentence combining automatizes syntax inside real writing. Grammar drill spends the same minutes on labeling language rather than producing it — and loses.

The product: Future Proof Education™ turns the league table into daily practice: the AI Tutor coaches planning, drafting and revising strategies against specific goals, the Memory Coach spaces sentence-level skills until they are automatic, and teacher dashboards show writing growth across a class, a school or a ministry.

In this article

  1. 01The neglected R
  2. 02What Writing Next measured
  3. 03The league table
  4. 04The grammar result
  5. 05Writing in the elementary grades
  6. 06Why strategy instruction wins
  7. 07Writing feeds reading
  8. 08What the evidence doesn’t show
  9. 09Writing instruction by the evidence
© 2026 FUTURE PROOF™
The route. 9 sections, from “The neglected R” to “Writing instruction by the evidence”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The neglected R

Of the three Rs, writing gets the least science and the most folklore. Reading has meta-analyses that changed national policy. Mathematics has decades of cognitive research. Writing, in most classrooms, is taught the way it was taught to the teacher: assign a piece, mark the errors, drill some grammar, repeat. National assessments keep returning the same verdict — most students write below the level their futures will demand.

The folklore was never necessary. Controlled experiments on writing instruction have accumulated since the 1960s, and they have been pooled twice at scale: once for adolescents (Graham & Perin, 2007), once for the elementary grades (Graham, McKeown, Kiuhara & Harris, 2012). The result is something writing pedagogy rarely gets: a ranked list of teachable practices, each with a measured average effect on the quality of what students write.

This article walks the list — what tops it, what surprises, and why the most familiar staple of language teaching comes out below zero. Then it turns the rankings into a working plan for schools.

What Writing Next measured

In 2007, Steve Graham and Dolores Perin published Writing Next, a Carnegie-commissioned meta-analysis of writing instruction for grades 4 to 12 (Graham & Perin, 2007). They gathered every true or quasi-experiment they could find in which one group of students received an identifiable instructional treatment, another did not, and writing quality was measured at the end. A companion analysis in a peer-reviewed journal reports the method in full, drawing on well over a hundred studies (Graham & Perin, 2007).

The output was eleven instructional elements, each with an average effect size — the difference treatment made, in standard deviations, on judged writing quality. An effect near 0.20 is small; 0.50 is moderate; 0.80 is large for a school intervention. Because every element faced the same yardstick, the list reads as a league table. Two cautions before the table. The elements are not rivals — several combine naturally. And effects on judged quality are the field’s standard outcome, but they carry the noise of human rating (Graham & Perin, 2007).

It helps to feel what the numbers mean in a classroom. An effect of 0.82 implies that the average student in the treatment group ends up writing better than roughly four out of five comparison students. An effect of 0.50 moves the average student past about two-thirds of them. These are not laboratory curiosities. They are differences a teacher can see in a stack of marking — and, at the top of the table, differences most schools have never deliberately bought.

The league table

Top of the table, at an average effect of roughly 0.82: strategy instruction — explicitly teaching students strategies for planning, drafting and revising, modeling each one, then fading support until students run the process alone (Graham & Perin, 2007). Tied with it, summarization at roughly 0.82: teaching students to compress a text into its essentials, a compact exercise in choosing and structuring content.

Collaborative writing follows at ≈0.75 — structured arrangements where students plan, draft and revise together rather than alone. Specific product goals land at ≈0.70: telling students precisely what the finished piece must accomplish, not just its topic. Word processing earns ≈0.55, sentence combining ≈0.50, and further down sit prewriting, inquiry and the process-writing approach around ≈0.32, with study of models at ≈0.25 (Graham & Perin, 2007).

The number

≈0.82 The average effect of explicit strategy instruction on adolescents’ writing quality — the top of the Writing Next league table, and the best-replicated result in the writing literature (Graham & Perin, 2007).

Notice what the top of the table shares. Strategy instruction, summarization, collaboration and product goals all make the invisible work of writing visible — planning, selecting, structuring, revising against a target. None of them is about correctness. The elements that police surface features cluster at the bottom. And one of them falls off the table entirely.

Strategy instruction 0.82 Summarization 0.82 Collaborative writing 0.75 Specific product goals 0.70 Word processing 0.55 Sentence combining 0.50 Inquiry activities 0.32 Study of models 0.25 0 0.2 0.4 0.6 0.8 average effect size on writing quality, grades 4–12 © 2026 FUTURE PROOF™
Figure 1. The Writing Next league table: average effect sizes for instructional elements tested in experiments across grades 4–12 (Graham & Perin, 2007). Strategy instruction and summarization lead; prewriting and the process approach (≈0.32, not all rows shown) sit mid-table. Traditional grammar instruction, at ≈−0.32, is plotted separately in Figure 2 — it is the only element below zero. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The grammar result

The most useful number in Writing Next is its most uncomfortable one. Across the studies in which traditional grammar instruction — parts of speech, sentence parsing, usage drill — served as a treatment, the average effect on writing quality was roughly −0.32 (Graham & Perin, 2007). Not zero. Negative. Students receiving grammar-focused instruction wrote measurably worse, on average, than comparison students spending the same time otherwise.

The finding was not new; it was a replication. George Hillocks’ landmark 1986 review of two decades of composition experiments had reached the same figure. Traditional grammar teaching scored near −0.3 — the worst-performing focus of instruction he examined — while sentence combining sat comfortably positive at ≈+0.35 (Hillocks, 1986). Twenty years later, a systematic review at York combed the trial evidence again and found no evidence that teaching formal grammar improves the quality or accuracy of children’s writing (Andrews et al., 2006). Three sweeps of the literature, one answer.

The catch

Grammar drill’s ≈−0.32 is not evidence that grammar knowledge harms writers. It is evidence about time: lessons spent labeling language displace lessons spent producing it (Graham & Perin, 2007). Grammar taught inside writing — as in sentence combining — reliably helps (Hillocks, 1986). The drill is the problem, not the grammar.

Why does a result this consistent keep losing to tradition? Partly habit: grammar drill is easy to set, easy to mark and visible to parents. Partly the exams: where tests reward error-spotting, teaching drifts toward it. And partly a confusion between two goals. Grammar as knowledge about language is a legitimate subject in its own right. The trials only measure its value as a route to better writing — and as a route, it points backwards (Andrews et al., 2006).

The constructive replacement is sentence combining: students take kernel sentences — The dog barked. The dog was old. — and fuse them into better ones, in the context of their own writing. That is grammar as a production skill rather than an identification skill. In a randomized study with fourth graders, peer-assisted sentence-combining practice beat grammar instruction on sentence-construction skill, with gains visible in story quality — for stronger and weaker writers alike (Saddler & Graham, 2005).

Writing in the elementary grades

Everything so far concerns grades 4 and up. In 2012, Graham, McKeown, Kiuhara and Harris pooled the true experiments for the elementary years, and the table replicated with the volume turned up (Graham, McKeown, Kiuhara & Harris, 2012). Strategy instruction again topped the list, with an average effect near ≈1.0 — and larger still when the strategies came wrapped in self-regulation: goal setting, self-monitoring, self-talk, the package known as SRSD, developed by Harris and Graham themselves.

The supporting cast repeated too, at approximate effects worth reading as a set: peer assistance around ≈0.9, specific product goals ≈0.76, prewriting activities ≈0.54, word processing ≈0.47, and simply adding writing time ≈0.30 (Graham, McKeown, Kiuhara & Harris, 2012). Even the humblest entry carries a lesson: many elementary classrooms allocate writing so little time that merely scheduling more of it produces a measurable gain.

For school leaders the elementary table removes an excuse. The strongest practices are not grade-locked. A seven-year-old can be taught a planning strategy, a revision checklist, a way to talk herself through a draft — and the measured payoff is, if anything, larger than for teenagers (Graham, McKeown, Kiuhara & Harris, 2012).

traditional grammar instruction sentence combining Hillocks (1986) ≈−0.29 ≈+0.35 Graham & Perin (2007) −0.32 +0.50 Andrews et al. (2006) review: no evidence grammar teaching improves writing −0.4 −0.2 0 +0.2 +0.4 +0.6 effect on writing quality — the same comparison, two decades apart © 2026 FUTURE PROOF™
Figure 2. One choice, measured twice. In both Hillocks’ 1986 review and the Writing Next pool, traditional grammar instruction averages below zero while sentence combining — grammar exercised inside writing — averages solidly positive (Hillocks, 1986) (Graham & Perin, 2007). Values are approximate pooled averages; the York systematic review reached the same qualitative verdict on formal grammar teaching (Andrews et al., 2006). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Why strategy instruction wins

The league table has a theory underneath it. Ronald Kellogg’s account of how writing skill develops treats composition as one of the most cognitively demanding things schools ask of anyone. Idea generation, word choice, sentence construction, spelling and handwriting all compete for the same limited working memory at the same moment (Kellogg, 2008). Experts survive the load because large parts of the process have become automatic, freeing attention for meaning. Novices drown in it.

Read the winning elements against that account and they stop looking like a grab bag. Strategy instruction hands the novice an expert’s routine — plan first, using these steps; revise against these questions — so the process no longer has to be improvised under load. Product goals shrink the target. Collaboration distributes the load across heads. Sentence combining and transcription practice automate the sentence-level machinery, which is why fluent handwriting and typing matter well beyond neatness (Kellogg, 2008). Every high-ranking element is a load-management device.

It also explains the self-regulation bonus in the elementary data: writing is long, frustrating work, and strategies survive only if the writer can monitor herself through them (Graham, McKeown, Kiuhara & Harris, 2012). Teaching the strategy without the self-management is handing over a map without the compass.

Writing feeds reading

One more pooled result belongs in the case for taking writing seriously. In 2011, Graham and Michael Hebert assembled the experiments on writing’s effect on reading. Writing about a text improved comprehension of it by roughly a third of a standard deviation on average — outperforming rereading it or discussing it (Graham & Hebert, 2011). Teaching writing, and simply increasing how much students write, also strengthened reading skill more broadly.

For curriculum planners this dissolves the zero-sum framing in which writing time competes with reading time. Summarizing, answering in writing, note-making — writing is how a reader externalizes and tests understanding. The two Rs are one system, and the league table above is also, quietly, reading instruction (Graham & Hebert, 2011).

Strategy instruction ≈1.02 Peer assistance ≈0.89 Specific product goals ≈0.76 Prewriting activities ≈0.54 Word processing ≈0.47 Extra writing time ≈0.30 0 0.25 0.5 0.75 1.0 approximate average effect size on writing quality, elementary grades © 2026 FUTURE PROOF™
Figure 3. The elementary replication. Pooled true experiments in the elementary grades repeat the adolescent ordering with larger effects at the top — strategy instruction near a full standard deviation, and even bare extra writing time earning ≈0.30 (Graham, McKeown, Kiuhara & Harris, 2012). Values are approximate pooled averages; strategy effects grow further when self-regulation (SRSD) is included. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Writing Next: Effective strategies to improve writing of adolescents in middle and high schools. Graham & Perin, Carnegie Corporation report, 2007

What the evidence doesn’t show

The writing literature is unusually consistent, but its findings are narrower than the slogans they spawned. The honest limits:

  • That grammar knowledge is worthless. The negative effect attaches to decontextualized drill displacing writing practice; grammar exercised inside composition — sentence combining — is one of the table’s solid positives (Hillocks, 1986) (Saddler & Graham, 2005).
  • A ready-made curriculum. The meta-analyses rank elements, not sequences; how much strategy instruction, combined with what, in which genre and grade, remains professional design work (Graham & Perin, 2007).
  • Durability and transfer. Most trials measure writing soon after treatment, in genres close to those taught; evidence on maintenance over years, and transfer to unfamiliar genres, is far thinner (Graham & Perin, 2007).
  • An objective outcome. Effects are on judged writing quality — human ratings with real reliability limits; the numbers inherit that noise (Graham & Perin, 2007).
  • That devices improve writing. Word processing’s ≈0.5 reflects drafting and revising fluidity, not hardware; a laptop scheme without instructional change buys the tool effect at most (Graham, McKeown, Kiuhara & Harris, 2012).
  • AI-era evidence. Every pooled study predates generative AI; how drafting with a model reshapes the practice that builds writing skill is, so far, unmeasured.

Where the evidence stops

  1. 1That grammar knowledge is worthless
  2. 2A ready-made curriculum
  3. 3Durability and transfer
  4. 4An objective outcome
  5. 5That devices improve writing
  6. 6AI-era evidence
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Writing instruction by the evidence

Turned into practice, the league table becomes six moves a school can make this term.

Teach the process explicitly. Name a planning strategy, model it aloud, practise it with support, fade the support — for every major genre. This is the table’s top row, in both age bands (Graham & Perin, 2007) (Graham, McKeown, Kiuhara & Harris, 2012). Add the self-regulation wrapper: goals, self-checks, self-talk.

Swap drill hours for sentence combining. Take the timetable slots that grammar drill occupies and refill them with combining, expanding and imitating sentences inside students’ own drafts (Hillocks, 1986) (Saddler & Graham, 2005). Same grammatical content, opposite sign.

Give every assignment a product goal. Not “write about the trip” but “persuade a reader who disagrees, with three reasons and a counterargument answered”. Specific targets earn ≈0.7 on their own (Graham & Perin, 2007).

Schedule writing, and some of it together. Volume alone helps in the elementary years, collaboration helps everywhere — structured pairs planning, drafting and revising jointly (Graham, McKeown, Kiuhara & Harris, 2012).

Automate the machinery early. Fluent transcription — handwriting and typing — frees working memory for composing; the load account predicts, and the data confirm, that sentence-level automaticity is a prerequisite, not a nicety (Kellogg, 2008).

Write across the curriculum. Summaries and written responses in science and history lift reading comprehension while building writing mileage — one intervention, two Rs (Graham & Hebert, 2011).

Applied at Future Proof Education

How Future Proof Education™ applies this.

The writing evidence favours explicit strategies, specific goals, sentence-level automaticity and steady volume — exactly the shape of the platform’s writing practice. The AI Tutor coaches students through planning, drafting and revising against a named product goal, modeling each strategy before fading support, and gives feedback on the writing itself rather than a grade. Sentence-combining exercises replace drill, drawn from each learner’s own drafts, and the Memory Coach spaces them until constructions come automatically. The Adaptive Diagnostic tracks growth in judged quality over time, teacher dashboards surface who is stalled at which stage of the process, and ministries get the same picture across a system. Grammar appears everywhere in the practice — and nowhere as a drill sheet.

See the platform
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 1986Hillocks
  • 2005Saddler
  • 2006Andrews
  • 2007Graham
  • 2007Graham
  • 2008Kellogg
  • 2011Graham
  • 2012Graham
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 1986–2012, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Graham, S., & Perin, D. (2007). Writing Next: Effective strategies to improve writing of adolescents in middle and high schools. Carnegie Corporation of New York / Alliance for Excellent Education, Washington, DC. PDF
  2. Graham, S., & Perin, D. (2007). A meta-analysis of writing instruction for adolescent students. Journal of Educational Psychology 99(3): 445–476. PDF
  3. Hillocks, G. (1986). Research on Written Composition: New Directions for Teaching. ERIC Clearinghouse on Reading and Communication Skills / National Conference on Research in English, Urbana, IL. PDF
  4. Andrews, R., Torgerson, C., Beverton, S., Freeman, A., Locke, T., Low, G., Robinson, A., & Zhu, D. (2006). The effect of grammar teaching on writing development. British Educational Research Journal 32(1): 39–55. PDF
  5. Saddler, B., & Graham, S. (2005). The effects of peer-assisted sentence-combining instruction on the writing performance of more and less skilled young writers. Journal of Educational Psychology 97(1): 43–54. PDF
  6. Graham, S., McKeown, D., Kiuhara, S., & Harris, K.R. (2012). A meta-analysis of writing instruction for students in the elementary grades. Journal of Educational Psychology 104(4): 879–896. PDF
  7. Kellogg, R.T. (2008). Training writing skills: A cognitive developmental perspective. Journal of Writing Research 1(1): 1–26. PDF
  8. Graham, S., & Hebert, M. (2011). Writing to read: A meta-analysis of the impact of writing and writing instruction on reading. Harvard Educational Review 81(4): 710–744. PDF
Try Future Proof Education

Writing practice with the drill removed.

Book a 20-minute demo. We’ll show you strategy coaching, goal-driven drafting and sentence-combining practice built from your students’ own writing — for a classroom or a country.

8 citations Reviewed August 2026 Open peer review welcomed