GCSE Mathematics
Marking.ai marked 240 handwritten GCSE Maths answers as closely as a second examiner did. Here is the study, question type by question type.
How accurately does Marking.ai mark GCSE Maths?
On 240 handwritten GCSE Maths answers, Marking.ai gave the same mark as the examiner 73% of the time and came within one mark 91% of the time. The two examiners who marked the same work agreed exactly with each other 72% of the time and came within one mark 91% of the time - so on the measures a teacher would use, Marking.ai marked this work as closely as a second examiner. On short answers it gave the exact mark on 85% of them, against the examiners' 82%. On the agreement score, which corrects for how easy a one-mark question is to agree on by chance, we read below a general AI model given the same mark scheme, and the full table is below.
What the study found
Marking.ai's figures, with the two examiners' own agreement on the same answers beside each one.
- of marks matched the examiner exactly. Examiners: 72%
- 73%
- of marks within one mark of the examiner. Examiners: 91%
- 91%
- of short answers marked exactly right. Examiners: 82%
- 85%
Marking.ai, two examiners, and a general AI model
Every arm marked the identical 240 answers. The third column is the same AI model Marking.ai runs on, given the same question and the same mark scheme in a single message, with nothing we built around it - not ChatGPT, and not any other chat product.
| Measure | Marking.ai | Two examiners | A general AI model |
|---|---|---|---|
| Same mark as the examiner | 73% | 72% | 75% |
| Within one mark of the examiner | 91% | 91% | 93% |
| Short answers marked exactly right | 85% | 82% | 84% |
| Average distance from the examiner's mark | 0.41 marks | 0.40 marks | 0.35 marks |
| All twelve questions, agreement score | 0.741 | 0.771 | 0.795 |
Maths is not the easy, deterministic subject
It is tempting to assume marking Maths is simple because arithmetic has right answers. The data says otherwise. The two examiners agreed with each other less closely on Maths than on English - 0.771 against 0.811 on the agreement score - because GCSE Maths marks are mostly method marks, and whether a student's approach earned one is a judgement rather than a calculation.
Raw agreement rates point the other way, and that is the trap. The examiners gave the same Maths mark 72% of the time against 29% on English, but most questions here are worth one to three marks, where agreeing by chance is easy. The agreement score corrects for that; the raw percentage does not.
How it compares with a general AI model
On the measures a teacher would recognise, Marking.ai and the general model both sit at the examiner bar and the differences are small. On the agreement score we read below it: 0.741 against 0.795, with the examiners at 0.771. Compared answer by answer the difference is not statistically separable from zero, so the honest summary is that we are level with a bare model on this corpus rather than behind it. On English we were measurably ahead. Here we were not, and we would rather you read that from us than find it.
Four answers explain almost all of the gap, and the explanation is not flattering to the alternative. In those four the student's working was not in the image at all: the canvas holds a single figure such as 50.5 degrees, because the working sat on the question paper, which the corpus does not capture. Marking.ai awarded one mark out of five and recorded why - no working shown, so the method marks could not be evidenced. The general model awarded five out of five, for working it could not have seen.
Remove those four answers and the order reverses: 0.780 for Marking.ai against the general model's 0.770. We publish the unadjusted figure, because dropping four inconvenient answers to improve a number is the practice this page exists to argue against. But it is worth knowing which behaviour you are buying: a marker that will not award a mark it cannot evidence.
One question resists that explanation. On a three-mark constrained optimisation question both arms scored far below the examiners, and we were wrong in both directions on different answers. The examiners were unstable there too - their own agreement was the second lowest of the twelve questions, and two answers with identical working were given different marks. It is a contested question rather than a defect with a known fix, and we have not fixed it.
What the study covers
It measures marking: an answer that has already been read off the page, marked against a mark scheme. Each answer was supplied on its own page. A teacher uploads a scanned script with many questions on it, and this study does not measure reading one. Keeping thirty scripts attached to the right questions is a different job, and we have not measured it.
The answers were chosen to cover the full range of marks, not to look like a real class. And the corpus has been public since June 2026, so an AI model trained after that date may have seen it - which applies equally to every arm, so the comparison between them holds.
As with every study in this section, it measures agreement rather than correctness. There is no ground-truth mark here, and on this subject the examiners disagree with each other more than they do on English.
How we measured it
Every answer was marked through the production marking path - the same routing, prompts, models and parsing a teacher's marking runs through. Nothing was re-implemented for the study. Each arm was scored against both examiners and averaged, so no arm can score well by happening to match whichever examiner was listed first. The corpus is the Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence. We did not mark the work and we did not choose the answers.
What this study has not measured
A measure we have not run is listed here rather than left out.
| Measure | Where it stands |
|---|---|
| Direction of error - generous or harsh when we differ | Not yet reported. The run measured distance but did not separate the two |
| Consistency - how far the same script moves when marked again | Not yet reported. Needs repeat marking of the same answers |
| Scanned, multi-question scripts | Out of scope. Each answer here was supplied on its own page, so extraction and question-matching are untested |
| A-level and beyond | Out of scope. This corpus is GCSE, on questions worth up to five marks |
In short
- Marking.ai gave the same mark as the examiner 73% of the time and came within one mark 91% of the time - matching the two examiners' own agreement with each other.
- On short answers it gave the exact mark 85% of the time, against the examiners' 82%.
- On the agreement score it read below a general AI model given the same mark scheme, and the table says so.
- Four answers where the student's working was never in the image account for almost all of that gap. The figures are published unadjusted anyway.
- Examiners agree less closely on Maths than on English, because GCSE Maths marks are mostly method marks.
- Each answer was supplied on its own page, so this does not measure reading a scanned, multi-question script.
Common questions
Where does the data come from?
The Medly AI GCSE marking benchmark, published openly under a CC BY 4.0 licence: Fox, M., Samra, K. & Jung, P. (2026). LLM Performance on a Real, Double-Marked GCSE Benchmark. Medly AI. github.com/medlyai/medly-marking-benchmark. The same corpus is behind our GCSE English Language study.
Does this mean equations are marked accurately?
It means we have evidence where we previously had none, on this corpus and this question set: handwritten working on GCSE questions worth up to five marks. It does not cover A-level, and it does not cover every notation a student might use. We would rather you checked it on your own class than took one study as settled.
Could the AI have seen this data already?
Possibly. The corpus has been public since June 2026, so an AI model trained after that date may have seen it. That applies equally to every arm of the test, so the comparison between them holds.
Assessment terms explained
Plain-English definitions of the vocabulary used on this page.
- AI markingAI marking is the use of artificial intelligence to assess student work against defined criteria - a mark scheme or a teacher's own marking guidelines - and to draft marks and feedback for a teacher to review.
- Assessment for learningAssessment for learning is the practice of using evidence of what students currently understand to decide what to teach next, and to help students see what they need to do to improve.
- Assessment objectivesAssessment objectives are the categories of skill a qualification tests - commonly labelled AO1, AO2 and AO3 - each carrying a defined share of the available marks.
- Comparative judgementComparative judgement is an assessment method in which markers repeatedly choose the better of two pieces of work, and those paired decisions are combined statistically into a rank order.
- Criterion-referenced assessmentCriterion-referenced assessment judges work against defined standards of what a student should know or be able to do, rather than against how other students performed.
- ExemplarAn exemplar is a piece of work - often a real student answer - presented with its marks and commentary to show what a particular standard looks like in practice.
The evidence your DPO will ask for
Ask us for the pack - our DPIA, our data processing agreement and the security overview - and a person will send it. If you would rather be walked through it first, book a demo for your school.

