
AI-generated writing samples for KS2 literacy: UK DfE pilot cuts costs but raises ethical concerns
Published by AINave Editorial • Reviewed by Ramit
The UK Department for Education (DfE) is piloting AI-generated writing samples for KS2 literacy moderation, using ChatGPT's GPT-5 to replace some human-authored scripts. The move could cut the annual moderation budget from around £100,000 by about 95%, but it also introduces serious questions about bias, authenticity, and governance that education technology builders should track closely.
The pilot: GPT-5 replaces some human-written scripts
Around 2,000 moderators currently cross-check pupils' writing against samples drawn from real children's work to standardise grades at the end of key stage 2. Starting this year, some moderators will instead compare against AI-generated output created by GPT-5. The DfE says GPT-5 is being used to create three collections of writing for one standardisation exercise, while one full exercise in 2026-27 and 2027-28 will continue to include scripts written by children obtained under the previous contract.
About 20 experienced local authority moderation managers will review the AI-derived material for authenticity before it is used. The department plans to decide in spring 2027 whether to continue producing samples this way or return to procuring them from an external supplier. Its transparency record lists no formal impact assessment for the system.
Why this matters for AI builders and education tech teams
This is a concrete case of AI-generated exemplars entering a high-stakes official assessment process. For teams building AI tools for education, the pilot highlights both the cost incentive and the governance gap. The DfE's own risk assessment notes that LLMs produce texts that often exclude atypical vocabulary and sentence structures used by neurodivergent or non-native English-speaking students. The department plans to mitigate this through thorough review, but no formal impact assessment has been published.
Rebecca Clarkson at Anglia Ruskin University, who studies KS2 writing assessment and moderation, told New Scientist that using AI-generated samples creates a philosophical and ethical issue. The worry is that AI shifts our perspective of what is acceptable in writing or what a good standard looks like. Jo-Anne Baird at the University of Oxford raised a similar concern: AI-generated materials could become the standards that teachers try to get pupils to emulate, leading to strange places.
The risks: bias, authenticity, and governance gaps
Because AI can produce flat or bland writing, children who speak English as a second language or who are neurodivergent may be discounted if their writing does not match the synthetic samples. The DfE says AI outputs will be heavily reviewed and edited, but the lack of a published impact assessment makes it hard to evaluate the safeguards.
The consequences of not using real children's writing are impossible to predict, Clarkson said. The DfE did not respond to New Scientist's request for comment. For now, the experiment is limited, but the decision in spring 2027 will determine whether AI-generated samples become a permanent fixture in UK literacy assessment.
For AI builders, this pilot is a reminder that deploying LLMs in assessment requires more than cost savings. Authenticity checks, bias mitigation, and independent governance are not optional extras. They are the difference between a useful tool and a system that distorts educational outcomes.
FAQs
Sources
- Test moderators use AI-generated writing to judge literacy standards
- blader/humanizer: Agent skill that removes signs of AI-generated...
- AI Detector: Ranked #1 Free AI Checker for ChatGPT
- Free AI Humanizer – 100% Human Text & No Sign-up, Unlimited
- Free AI Text Humanizer
- Humanize AI: Guaranteed 100% Human Score & Unlimited Free Words
- Why don't moderators test their AI content detection abilities?
- Use AI





















