Evidence & accuracy

GCSE English Language

Marking.ai marked as closely as a second examiner - and agreed with examiners more closely than a general AI model given the same mark scheme. Here is the study.

How accurately does Marking.ai mark GCSE English?

On GCSE English Language reading responses, Marking.ai landed within 10% of the question's total marks on 71% of answers. The two human examiners who marked the same work managed that with each other on 72% of them. Across all eight questions in the corpus our agreement score was 0.833, against the examiners' own 0.811 - so Marking.ai marked this work about as closely as a second examiner would have. The same AI model handed the same mark scheme, with none of Marking.ai around it, scored 0.805, and compared answer by answer we came out ahead by a margin that excludes zero.

What the study found

Marking.ai's figures, with the two examiners' own agreement on the same answers beside each one.

of reading-response marks within 10% of the examiner's. Examiners: 72%
71%
agreement across all eight questions. Examiners: 0.811
0.833
real GCSE answers, each marked by two examiners
240
The results

Marking.ai, two examiners, and a general AI model

Every arm marked the identical 240 answers. The third column is the same AI model Marking.ai runs on, given the same question and the same mark scheme in a single message, with nothing we built around it - not ChatGPT, and not any other chat product.

GCSE English Language marking study, measured 15 September 2026. Data: Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI. github.com/medlyai/medly-marking-benchmark
MeasureMarking.aiTwo examinersA general AI model
Reading responses within 10% of the question's total marks71%72%64%
Reading responses, agreement score0.8530.8490.803
All eight questions, agreement score0.8330.8110.805
Average distance from the examiner's mark2.12 marks2.11 marks2.39 marks

Why a second examiner is the measure that matters

Most accuracy questions assume there is a right mark and the only question is whether the machine finds it. Marking does not work like that. Every answer in this corpus was marked independently by two qualified examiners, and they disagree with each other constantly - on the 40-mark essays they gave the same mark 13% of the time.

That second examiner is what makes the study worth running. A figure like "agrees with the examiner 70% of the time" has nothing to sit beside, and a reader cannot tell whether it is good. With a human bar on the same work, it can be read.

What the study covers

It measures marking: an answer that has already been read off the page, marked against a mark scheme. Each answer was supplied on its own page. A teacher uploads a scanned script with many questions on it, and this study does not measure reading one. Reading a scanned script and marking the answer on it are separate skills, and this study measures the second.

The answers were chosen to cover the full range of marks, not to look like a real class. And the corpus has been public since June 2026, so an AI model trained after that date may have seen it - which applies equally to every arm, so the comparison between them holds.

It is one subject and one corpus, so it says nothing about mathematics. That has its own study, and it reaches a different conclusion.

And it measures agreement rather than correctness. There is no ground-truth mark on work like this - the examiners disagree with each other - so a figure here says how close our mark sits to a qualified human's, not how often it is right.

How we measured it

Every answer was marked through the production marking path: the same routing rules that decide which strategy marks a question, the same prompts and models those strategies carry, the same parsing applied to the result. Nothing about the marking was re-implemented for the study, which is what makes the outcome evidence about the product rather than about a research tool.

Agreement is reported as quadratic weighted kappa, the measure the published benchmark uses, computed per question and then averaged. It runs from 0 for chance agreement to 1 for identical marks, and being one mark apart counts far less against a marker than being ten apart. Each arm was scored against both examiners and averaged, so no arm can score well by happening to match whichever examiner was listed first.

The corpus is the Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence. We did not mark the work and we did not choose the answers.

Scope

What this study has not measured

A measure we have not run, or one we will not stand behind yet, is listed here rather than left out.

Measures not reported by the GCSE English Language marking study, as at 15 September 2026
MeasureWhere it stands
Extended writing (the two 40-mark questions)Not reported. Two questions is too small a base for a claim in either direction, and our figure on them is the flattering one
Direction of error - generous or harsh when we differNot yet reported. The run measured distance but did not separate the two
Consistency - how far the same script moves when marked againNot yet reported. Needs repeat marking of the same answers
Scanned, handwritten scriptsOut of scope. This corpus is typed, and each answer was supplied on its own page
Equations and mathematical workingCovered by the GCSE Mathematics study, which reports separately

In short

  • Marking.ai marked GCSE English about as closely as a second examiner: 0.833 agreement against the examiners' own 0.811.
  • On reading responses it landed within 10% of the question total on 71% of answers, against the examiners' 72%.
  • It agreed with examiners more closely than the same AI model given the same mark scheme, and the difference holds when compared answer by answer.
  • 240 real GCSE answers, each marked independently by two examiners before we saw them, from a public benchmark we did not build.
  • The work is typed and supplied one answer per page, so this measures marking rather than reading a scanned script.
The GCSE Mathematics study
Answers

Common questions

Still stuck?

Ask us anything - a real person answers, in term time.

Ask a question

Where does the data come from?

The Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence: Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI. github.com/medlyai/medly-marking-benchmark. It is a public subset of a larger study covering 32,534 answers, which has not been released. We tested everything that was made public rather than choosing a favourable slice.

Could the AI have seen this data already?

Possibly. The corpus has been public since June 2026, so an AI model trained after that date may have seen it. That applies equally to every arm of the test, so the comparison between them holds. We would rather say so than leave you to find it.

Does this tell me how well it will mark my class?

Partly. It measures marking against a public mark scheme on typed answers, so it is a good guide to how the marking behaves on this kind of question and no guide at all to how well a scanned script is read. The most reliable check is still the direct one: run it on a set you have already marked and compare.

The evidence your DPO will ask for

Ask us for the pack - our DPIA, our data processing agreement and the security overview - and a person will send it. If you would rather be walked through it first, book a demo for your school.