Evidence & accuracy

GCSE Mathematics

Marking.ai marked 240 handwritten GCSE Maths answers as closely as a second examiner did. Here is the study, question type by question type.

How accurately does Marking.ai mark GCSE Maths?

On 240 handwritten GCSE Maths answers, Marking.ai gave the same mark as the examiner 73% of the time and came within one mark 91% of the time. The two examiners who marked the same work agreed exactly with each other 72% of the time and came within one mark 91% of the time - so on the measures a teacher would use, Marking.ai marked this work as closely as a second examiner. On short answers it gave the exact mark on 85% of them, against the examiners' 82%. On the agreement score, which corrects for how easy a one-mark question is to agree on by chance, we read below a general AI model given the same mark scheme, and the full table is below.

What the study found

Marking.ai's figures, with the two examiners' own agreement on the same answers beside each one.

of marks matched the examiner exactly. Examiners: 72%
73%
of marks within one mark of the examiner. Examiners: 91%
91%
of short answers marked exactly right. Examiners: 82%
85%
The results

Marking.ai, two examiners, and a general AI model

Every arm marked the identical 240 answers. The third column is the same AI model Marking.ai runs on, given the same question and the same mark scheme in a single message, with nothing we built around it - not ChatGPT, and not any other chat product.

GCSE Mathematics marking study, measured 15 September 2026. Data: Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI. github.com/medlyai/medly-marking-benchmark
MeasureMarking.aiTwo examinersA general AI model
Same mark as the examiner73%72%75%
Within one mark of the examiner91%91%93%
Short answers marked exactly right85%82%84%
Average distance from the examiner's mark0.41 marks0.40 marks0.35 marks
All twelve questions, agreement score0.7410.7710.795

Maths is not the easy, deterministic subject

It is tempting to assume marking Maths is simple because arithmetic has right answers. The data says otherwise. The two examiners agreed with each other less closely on Maths than on English - 0.771 against 0.811 on the agreement score - because GCSE Maths marks are mostly method marks, and whether a student's approach earned one is a judgement rather than a calculation.

Raw agreement rates point the other way, and that is the trap. The examiners gave the same Maths mark 72% of the time against 29% on English, but most questions here are worth one to three marks, where agreeing by chance is easy. The agreement score corrects for that; the raw percentage does not.

How it compares with a general AI model

On the measures a teacher would recognise, Marking.ai and the general model both sit at the examiner bar and the differences are small. On the agreement score we read below it: 0.741 against 0.795, with the examiners at 0.771. Compared answer by answer the difference is not statistically separable from zero, so the honest summary is that we are level with a bare model on this corpus rather than behind it. On English we were measurably ahead. Here we were not, and we would rather you read that from us than find it.

Four answers explain almost all of the gap, and the explanation is not flattering to the alternative. In those four the student's working was not in the image at all: the canvas holds a single figure such as 50.5 degrees, because the working sat on the question paper, which the corpus does not capture. Marking.ai awarded one mark out of five and recorded why - no working shown, so the method marks could not be evidenced. The general model awarded five out of five, for working it could not have seen.

Remove those four answers and the order reverses: 0.780 for Marking.ai against the general model's 0.770. We publish the unadjusted figure, because dropping four inconvenient answers to improve a number is the practice this page exists to argue against. But it is worth knowing which behaviour you are buying: a marker that will not award a mark it cannot evidence.

One question resists that explanation. On a three-mark constrained optimisation question both arms scored far below the examiners, and we were wrong in both directions on different answers. The examiners were unstable there too - their own agreement was the second lowest of the twelve questions, and two answers with identical working were given different marks. It is a contested question rather than a defect with a known fix, and we have not fixed it.

What the study covers

It measures marking: an answer that has already been read off the page, marked against a mark scheme. Each answer was supplied on its own page. A teacher uploads a scanned script with many questions on it, and this study does not measure reading one. Keeping thirty scripts attached to the right questions is a different job, and we have not measured it.

The answers were chosen to cover the full range of marks, not to look like a real class. And the corpus has been public since June 2026, so an AI model trained after that date may have seen it - which applies equally to every arm, so the comparison between them holds.

As with every study in this section, it measures agreement rather than correctness. There is no ground-truth mark here, and on this subject the examiners disagree with each other more than they do on English.

How we measured it

Every answer was marked through the production marking path - the same routing, prompts, models and parsing a teacher's marking runs through. Nothing was re-implemented for the study. Each arm was scored against both examiners and averaged, so no arm can score well by happening to match whichever examiner was listed first. The corpus is the Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence. We did not mark the work and we did not choose the answers.

Scope

What this study has not measured

A measure we have not run is listed here rather than left out.

Measures not reported by the GCSE Mathematics marking study, as at 15 September 2026
MeasureWhere it stands
Direction of error - generous or harsh when we differNot yet reported. The run measured distance but did not separate the two
Consistency - how far the same script moves when marked againNot yet reported. Needs repeat marking of the same answers
Scanned, multi-question scriptsOut of scope. Each answer here was supplied on its own page, so extraction and question-matching are untested
A-level and beyondOut of scope. This corpus is GCSE, on questions worth up to five marks

In short

  • Marking.ai gave the same mark as the examiner 73% of the time and came within one mark 91% of the time - matching the two examiners' own agreement with each other.
  • On short answers it gave the exact mark 85% of the time, against the examiners' 82%.
  • On the agreement score it read below a general AI model given the same mark scheme, and the table says so.
  • Four answers where the student's working was never in the image account for almost all of that gap. The figures are published unadjusted anyway.
  • Examiners agree less closely on Maths than on English, because GCSE Maths marks are mostly method marks.
  • Each answer was supplied on its own page, so this does not measure reading a scanned, multi-question script.
The GCSE English Language study
Answers

Common questions

Still stuck?

Ask us anything - a real person answers, in term time.

Ask a question

Where does the data come from?

The Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence: Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI. github.com/medlyai/medly-marking-benchmark. The same corpus is behind our GCSE English Language study.

Does this mean equations are marked accurately?

It means we have evidence where we previously had none, on this corpus and this question set: handwritten working on GCSE questions worth up to five marks. It does not cover A-level, and it does not cover every notation a student might use. We would rather you checked it on your own class than took one study as settled.

Could the AI have seen this data already?

Possibly. The corpus has been public since June 2026, so an AI model trained after that date may have seen it. That applies equally to every arm of the test, so the comparison between them holds.

The evidence your DPO will ask for

Ask us for the pack - our DPIA, our data processing agreement and the security overview - and a person will send it. If you would rather be walked through it first, book a demo for your school.