Security

Evidence & accuracy

A dedicated hub to showcase how Marking.ai compares against common accuracy benchmarks.

How accurate is Marking.ai?

There isn't one meaningful accuracy percentage for marking, because qualified human examiners do not always award exactly the same mark to the same piece of work. So we benchmark Marking.ai against a more useful standard: how closely it agrees with qualified examiners compared with how closely those examiners agree with each other. These studies use real student answers that had already been independently marked by two qualified examiners before Marking.ai saw them, and the work is marked through the same production system teachers use. We publish our accuracy evidence subject by subject rather than claiming a single universal accuracy rate.

How often two qualified examiners give the same mark

Before asking how close a machine gets, it is worth knowing how close two humans get. These are the examiners in our corpus, marking the same work as each other.

same mark on a 40-mark English essay
13%
same mark across all the English questions
29%
same mark on GCSE Maths, mostly 1 to 3-mark questions
72%
The studies

Accuracy studies

Browse our portfolio of accuracy studies across subjects. Each study uses real student answers that had already been independently marked by two qualified examiners

Information about our studies

There is rarely one unquestionable “correct” mark in judgement-based assessment. Qualified examiners can legitimately award different marks to the same work.

That is why we benchmark Marking.ai against human-to-human agreement, not just a single examiner mark.

Where possible, our studies use student work that has already been independently marked by multiple qualified examiners. This lets us ask the more useful question:

How closely does Marking.ai agree with qualified examiners compared with how closely those examiners agree with each other?

How we measure accuracy

Our studies are designed to test the same production marking system teachers actually use against independently marked student work.

We aim to use:

  • real student responses;
  • multiple independent examiner marks;
  • the original assessment and marking guidelines;
  • a broad range of question types and performance levels.

We also compare the full Marking.ai system with the same underlying AI model where appropriate, helping us isolate the value added by our specialised marking strategies and workflow.

Results are always reported in context by subject, question type and measurement method - rather than reduced to one universal accuracy percentage.

Check a mark on your own class set

Mark one assessment you would have marked anyway, and look at what each mark rests on. No card required to start.

Answers

Common questions

Still stuck?

Ask us anything - a real person answers, in term time.

Ask a question

Isn't AI marking inaccurate?

The useful comparison isn't whether AI always matches one examiner exactly - qualified examiners don't always match each other exactly either. That's why we benchmark Marking.ai against independently marked student work and compare how closely it agrees with qualified examiners against how closely those examiners agree with each other. Across our published accuracy studies, we report performance by subject and question type rather than relying on a single headline percentage. We then add further safeguards through specialised marking strategies, evidence behind marks, confidence indicators, and teacher review before results are shared.

Why is there no single accuracy percentage?

Because different marking tasks behave differently. Agreement on a short, points-based question is not directly comparable with agreement on an extended response where professional judgement plays a larger role. We therefore report accuracy by subject, assessment and question type, state the measurement being used, and compare Marking.ai with the relevant human examiner benchmark. As new studies are published, they are added to the evidence base rather than combined into one potentially misleading number.

Why not just use ChatGPT?

General-purpose AI can be prompted to mark student work. The difference is the system built around the model. Marking.ai applies specialised marking strategies, uses the relevant marking guidelines, provides evidence and feedback, manages real class marking workflows, and gives teachers a structured review process. Where our studies include a controlled comparison with the underlying model, we publish those results too. This allows us to measure what the full Marking.ai system adds rather than simply claiming that it performs better.

How old are these figures?

Every published study is dated because both AI models and the Marking.ai marking system continue to evolve. Each study reflects the production system used at the time of measurement. As the system changes, we re-run relevant benchmarks and publish updated results rather than allowing old figures to stand indefinitely.

Why not quote a figure you already have?

Marking.ai does more than send a student's answer and marking guidelines to an AI model. Depending on the assessment, the production system can: route different questions through specialised marking strategies; evaluate work against the relevant criteria; connect marks to evidence in the student's response; use confidence indicators to help focus review. We then test that system against independently marked student work and publish the results by subject and assessment type. The aim is not to claim perfect marking. It is to build a marking system whose performance can be measured transparently against the human standards schools already rely on.

Check a mark on your own class set

Mark one assessment you would have marked anyway, and look at what each mark rests on. No card required to start.