GCSE English Language
Across 240 real GCSE English Language answers, Marking.ai marked about as closely as a second qualified examiner. Here is every figure, beside the examiners' own.
How accurately does Marking.ai mark GCSE English?
On GCSE English Language reading responses, Marking.ai landed within 10% of the question's total marks on 71% of answers. The two human examiners who marked the same work managed that with each other on 72% of them. Across all eight questions in the corpus our agreement score was 0.833, against the examiners' own 0.811 - so Marking.ai marked this work about as closely as a second examiner would have. The same AI model handed the same mark scheme, with none of Marking.ai around it, scored 0.805, and compared answer by answer we came out ahead by a margin that excludes zero.
What the study found
Marking.ai's figures, with the two examiners' own agreement on the same answers beside each one.
- of reading-response marks within 10% of the examiner's. Examiners: 72%
- 71%
- agreement across all eight questions. Examiners: 0.811
- 0.833
- real GCSE answers, each marked by two examiners
- 240
Marking.ai, two examiners, and a general AI model
Every arm marked the identical 240 answers. The first column is Marking.ai, the second is the agreement between the two examiners, and the third is the same AI model Marking.ai runs on, given the same question and mark scheme, with nothing we built around it.
| Measure | Marking.ai | Two examiners | A general AI model |
|---|---|---|---|
| Reading responses within 10% of the question's total marks | 71% | 72% | 64% |
| Reading responses, agreement score | 0.853 | 0.849 | 0.803 |
| All eight questions, agreement score | 0.833 | 0.811 | 0.805 |
| Average distance from the examiner's mark | 2.12 marks | 2.11 marks | 2.39 marks |
What are we comparing Marking.ai against?
The most useful benchmark for AI marking is not perfection. It is another qualified examiner.
Every answer in this study had already been marked independently by two qualified examiners. That matters because examiners do not always give the same mark to the same piece of work.
We therefore compared three things on exactly the same student answers:
- Marking.ai vs the examiners
- One examiner vs the other examiner
- A general AI model vs the examiners
Across all eight questions, Marking.ai achieved an agreement score of 0.833. The two examiners achieved 0.811 agreement with each other, while the general AI model achieved 0.805.
In simple terms: Marking.ai agreed with the qualified examiners at least as closely as another qualified examiner did - and more closely than the general AI model.
What did we test?
We tested 240 real GCSE English Language answers, covering eight questions and a range of marks.
The answers came from the independently published Medly AI GCSE marking benchmark. Every answer had been marked by two qualified examiners before Marking.ai saw it, and we tested the entire publicly available dataset rather than selecting particular answers.
Marking.ai marked every answer using the same production system our customers use - including the same routing, prompts, models and processing.
The general AI comparison used the same underlying AI model and the same question and mark scheme, but without Marking.ai's marking system around it. This helps separate the performance of the underlying AI model from the additional value provided by Marking.ai.
What does "accuracy" mean here?
For this kind of marking, there is not always one objectively correct mark.
The two qualified examiners themselves disagreed on many answers - including giving exactly the same mark on only 13% of the 40-mark responses.
So rather than pretending one examiner's mark is absolute ground truth, we measure agreement: how closely each marker's judgement matches the qualified examiners.
We explain why examiner agreement matters when evaluating AI marking in more detail, including what the 13% figure can and cannot be used to claim.
Our main agreement measure is quadratic weighted kappa. It ranges from 0 for chance agreement to 1 for identical marking, while recognising that being one mark apart is much less significant than being ten marks apart.
Each system was compared independently against both examiners, so the result does not depend on treating either examiner as the "correct" one.
What does this study tell us?
This study provides evidence that, on this set of GCSE English Language answers:
Marking.ai performs at approximately the level of a second qualified examiner and outperforms the same general AI model used without Marking.ai's marking system.
It does not mean that every mark Marking.ai produces will match an examiner.
It also does not establish one universal accuracy figure for every subject, assessment or type of student work. We test different subjects separately because marking behaviour - and the appropriate benchmark - can differ substantially between them.
What are the limitations?
This study measures marking, not document reading.
The benchmark contains typed answers supplied individually. It therefore tests how accurately Marking.ai applies the mark scheme once it has the student's response - not how accurately handwriting or a multi-page scanned script is extracted.
The answers were also selected by the benchmark creators to cover the range of available marks rather than to represent the mark distribution of a typical class.
The dataset has been publicly available since June 2026, so it is possible that AI models trained subsequently could have encountered it. The same dataset was used for every AI comparison in this study, so this does not favour Marking.ai over the general AI baseline.
Study source
The study uses the Medly AI GCSE marking benchmark, published under a CC BY 4.0 licence:
Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI.
Marking.ai did not create the benchmark, select the answers or provide the original examiner marks.
What this study has not measured
A measure we have not run, or one we will not stand behind yet, is listed here rather than left out.
| Measure | Where it stands |
|---|---|
| Extended writing (the two 40-mark questions) | Not reported. Two questions is too small a base for a claim in either direction, and our figure on them is the flattering one |
| Direction of error - generous or harsh when we differ | Not yet reported. The run measured distance but did not separate the two |
| Consistency - how far the same script moves when marked again | Not yet reported. Needs repeat marking of the same answers |
| Scanned, handwritten scripts | Out of scope. This corpus is typed, and each answer was supplied on its own page |
| Equations and mathematical working | Covered by the GCSE Mathematics study, which reports separately |
In short
- Marking.ai marked GCSE English about as closely as a second examiner: 0.833 agreement against the examiners' own 0.811.
- On reading responses it landed within 10% of the question total on 71% of answers, against the examiners' 72%.
- It agreed with examiners more closely than the same AI model given the same mark scheme, and the difference holds when compared answer by answer.
- 240 real GCSE answers, each marked independently by two examiners before we saw them, from a public benchmark we did not build.
- The work is typed and supplied one answer per page, so this measures marking rather than reading a scanned script.
Common questions
Where does the data come from?
The Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence: Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI. github.com/medlyai/medly-marking-benchmark. It is a public subset of a larger study covering 32,534 answers, which has not been released. We tested everything that was made public rather than choosing a favourable slice.
Does this tell me how well it will mark my class?
Partly. It measures marking against a public mark scheme on typed answers, so it is a good guide to how the marking behaves on this kind of question and no guide at all to how well a scanned script is read. The most reliable check is still the direct one: run it on a set you have already marked and compare.
Assessment terms explained
Plain-English definitions of the vocabulary used on this page.
- AI markingAI marking is the use of artificial intelligence to assess student work against defined criteria - a mark scheme or a teacher's own marking guidelines - and to draft marks and feedback for a teacher to review.
- Assessment for learningAssessment for learning is the practice of using evidence of what students currently understand to decide what to teach next, and to help students see what they need to do to improve.
- Assessment objectivesAssessment objectives are the categories of skill a qualification tests - commonly labelled AO1, AO2 and AO3 - each carrying a defined share of the available marks.
- Comparative judgementComparative judgement is an assessment method in which markers repeatedly choose the better of two pieces of work, and those paired decisions are combined statistically into a rank order.
- Criterion-referenced assessmentCriterion-referenced assessment judges work against defined standards of what a student should know or be able to do, rather than against how other students performed.
- ExemplarAn exemplar is a piece of work - often a real student answer - presented with its marks and commentary to show what a particular standard looks like in practice.
The evidence your DPO will ask for
Ask us for the pack - our DPIA, our data processing agreement and the security overview - and a person will send it. If you would rather be walked through it first, book a demo for your school.

