GCSE English Language
Marking.ai marked as closely as a second examiner - and agreed with examiners more closely than a general AI model given the same mark scheme. Here is the study.
How accurately does Marking.ai mark GCSE English?
On GCSE English Language reading responses, Marking.ai landed within 10% of the question's total marks on 71% of answers. The two human examiners who marked the same work managed that with each other on 72% of them. Across all eight questions in the corpus our agreement score was 0.833, against the examiners' own 0.811 - so Marking.ai marked this work about as closely as a second examiner would have. The same AI model handed the same mark scheme, with none of Marking.ai around it, scored 0.805, and compared answer by answer we came out ahead by a margin that excludes zero.
What the study found
Marking.ai's figures, with the two examiners' own agreement on the same answers beside each one.
- of reading-response marks within 10% of the examiner's. Examiners: 72%
- 71%
- agreement across all eight questions. Examiners: 0.811
- 0.833
- real GCSE answers, each marked by two examiners
- 240
Marking.ai, two examiners, and a general AI model
Every arm marked the identical 240 answers. The third column is the same AI model Marking.ai runs on, given the same question and the same mark scheme in a single message, with nothing we built around it - not ChatGPT, and not any other chat product.
| Measure | Marking.ai | Two examiners | A general AI model |
|---|---|---|---|
| Reading responses within 10% of the question's total marks | 71% | 72% | 64% |
| Reading responses, agreement score | 0.853 | 0.849 | 0.803 |
| All eight questions, agreement score | 0.833 | 0.811 | 0.805 |
| Average distance from the examiner's mark | 2.12 marks | 2.11 marks | 2.39 marks |
Why a second examiner is the measure that matters
Most accuracy questions assume there is a right mark and the only question is whether the machine finds it. Marking does not work like that. Every answer in this corpus was marked independently by two qualified examiners, and they disagree with each other constantly - on the 40-mark essays they gave the same mark 13% of the time.
That second examiner is what makes the study worth running. A figure like "agrees with the examiner 70% of the time" has nothing to sit beside, and a reader cannot tell whether it is good. With a human bar on the same work, it can be read.
What the study covers
It measures marking: an answer that has already been read off the page, marked against a mark scheme. Each answer was supplied on its own page. A teacher uploads a scanned script with many questions on it, and this study does not measure reading one. Reading a scanned script and marking the answer on it are separate skills, and this study measures the second.
The answers were chosen to cover the full range of marks, not to look like a real class. And the corpus has been public since June 2026, so an AI model trained after that date may have seen it - which applies equally to every arm, so the comparison between them holds.
It is one subject and one corpus, so it says nothing about mathematics. That has its own study, and it reaches a different conclusion.
And it measures agreement rather than correctness. There is no ground-truth mark on work like this - the examiners disagree with each other - so a figure here says how close our mark sits to a qualified human's, not how often it is right.
How we measured it
Every answer was marked through the production marking path: the same routing rules that decide which strategy marks a question, the same prompts and models those strategies carry, the same parsing applied to the result. Nothing about the marking was re-implemented for the study, which is what makes the outcome evidence about the product rather than about a research tool.
Agreement is reported as quadratic weighted kappa, the measure the published benchmark uses, computed per question and then averaged. It runs from 0 for chance agreement to 1 for identical marks, and being one mark apart counts far less against a marker than being ten apart. Each arm was scored against both examiners and averaged, so no arm can score well by happening to match whichever examiner was listed first.
The corpus is the Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence. We did not mark the work and we did not choose the answers.
What this study has not measured
A measure we have not run, or one we will not stand behind yet, is listed here rather than left out.
| Measure | Where it stands |
|---|---|
| Extended writing (the two 40-mark questions) | Not reported. Two questions is too small a base for a claim in either direction, and our figure on them is the flattering one |
| Direction of error - generous or harsh when we differ | Not yet reported. The run measured distance but did not separate the two |
| Consistency - how far the same script moves when marked again | Not yet reported. Needs repeat marking of the same answers |
| Scanned, handwritten scripts | Out of scope. This corpus is typed, and each answer was supplied on its own page |
| Equations and mathematical working | Covered by the GCSE Mathematics study, which reports separately |
In short
- Marking.ai marked GCSE English about as closely as a second examiner: 0.833 agreement against the examiners' own 0.811.
- On reading responses it landed within 10% of the question total on 71% of answers, against the examiners' 72%.
- It agreed with examiners more closely than the same AI model given the same mark scheme, and the difference holds when compared answer by answer.
- 240 real GCSE answers, each marked independently by two examiners before we saw them, from a public benchmark we did not build.
- The work is typed and supplied one answer per page, so this measures marking rather than reading a scanned script.
Common questions
Where does the data come from?
The Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence: Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI. github.com/medlyai/medly-marking-benchmark. It is a public subset of a larger study covering 32,534 answers, which has not been released. We tested everything that was made public rather than choosing a favourable slice.
Could the AI have seen this data already?
Possibly. The corpus has been public since June 2026, so an AI model trained after that date may have seen it. That applies equally to every arm of the test, so the comparison between them holds. We would rather say so than leave you to find it.
Does this tell me how well it will mark my class?
Partly. It measures marking against a public mark scheme on typed answers, so it is a good guide to how the marking behaves on this kind of question and no guide at all to how well a scanned script is read. The most reliable check is still the direct one: run it on a set you have already marked and compare.
Assessment terms explained
Plain-English definitions of the vocabulary used on this page.
- AI markingAI marking is the use of artificial intelligence to assess student work against defined criteria - a mark scheme or a teacher's own marking guidelines - and to draft marks and feedback for a teacher to review.
- Assessment for learningAssessment for learning is the practice of using evidence of what students currently understand to decide what to teach next, and to help students see what they need to do to improve.
- Assessment objectivesAssessment objectives are the categories of skill a qualification tests - commonly labelled AO1, AO2 and AO3 - each carrying a defined share of the available marks.
- Comparative judgementComparative judgement is an assessment method in which markers repeatedly choose the better of two pieces of work, and those paired decisions are combined statistically into a rank order.
- Criterion-referenced assessmentCriterion-referenced assessment judges work against defined standards of what a student should know or be able to do, rather than against how other students performed.
- ExemplarAn exemplar is a piece of work - often a real student answer - presented with its marks and commentary to show what a particular standard looks like in practice.
The evidence your DPO will ask for
Ask us for the pack - our DPIA, our data processing agreement and the security overview - and a person will send it. If you would rather be walked through it first, book a demo for your school.

