Can AI mark accurately? A framework for evaluating AI marking
There is no single meaningful percentage that answers whether AI can mark accurately. A credible answer depends on the assessment being marked, the benchmark used, the way agreement is measured, and the decision the school is actually trying to make.
Founder, Marking.ai
5 October 2026 · 9 min read

As a founder of an AI-assisted assessment, marking and feedback platform, I get asked a version of the same question all the time: "Can AI mark accurately?"
It comes from teachers, school leaders, funders, researchers and friends who are simply trying to understand whether AI can be trusted with something as consequential as a student's work.
I have become relatively programmed to answer by saying that we prefer the word alignment to accuracy. Accuracy sounds as though there is a single, pre-defined correct mark sitting somewhere in the background. In many assessments, particularly those involving professional judgement, that simply is not how marking behaves.
But the more I have had to defend that answer, the more I have wondered whether it can also become a convenient escape hatch for companies like ours. If an AI system struggles to match the mark a human gave, it is tempting to say that the benchmark was subjective anyway.
So rather than avoiding the accuracy question, I think we should ask it more rigorously.
The useful question is not simply "Can AI mark accurately?" It is: "How should we evaluate whether an AI marking system is accurate enough for this assessment, against a credible human benchmark, for the outcome we are trying to achieve?"
A four-question framework for evaluating AI marking accuracy
When a school is shown an AI marking accuracy claim, I think four questions should come before the headline percentage.
The framework at a glance
- Relevance: Was the AI tested on the same kind of assessment, subject and student work you care about?
- Benchmark: Who or what is being treated as the reference point - one marker, multiple independent markers, or an adjudicated standard?
- Measurement: What does 'agreement' actually mean - exact marks, distance from the mark, grade band, or a reliability statistic?
- Outcome: What decision are you really trying to make - theoretical accuracy, safe deployment, or whether AI reduces the teacher's review workload?
1. Is the evidence relevant to the assessment you actually want to mark?
Asking whether AI can mark accurately in the abstract is a little like asking whether an experienced engineer can accurately mark an A Level Chemistry paper. They may be intelligent and analytically capable, but that does not tell you whether they understand the assessment objectives, the mark scheme, the subject-specific conventions or the nuances in the answers students actually produce.
The same applies to AI. Evidence from a multiple-choice science test tells you very little about a 40-mark English response. Evidence from typed answers may tell you something about marking behaviour, while telling you almost nothing about how reliably a system reads a multi-page handwritten script.
This is also why I would be wary of any provider claiming one universal accuracy percentage. Our own evidence at Marking.ai is published subject by subject for exactly this reason. Our GCSE English Language study and GCSE Mathematics study do not behave identically, and they should not be collapsed into one number.
Independent research points in the same direction. A 2026 University of Cambridge project testing AI marking of written university exam responses found that performance varied materially across institutions and warned against assuming that results in one setting generalise to another.
2. What is the benchmark?
The most common school-level test I hear is straightforward: mark the mocks manually, run the same papers through the AI, and compare the results.
There is nothing inherently wrong with doing that. In fact, local testing on your own material is useful. The problem comes when the human mark is treated as unquestionable ground truth.
In judgement-based assessment, disagreement with one marker does not automatically mean the AI is wrong. The human marker may also have made a different professional judgement from another qualified marker.
That is why, where professional judgement is involved, multiple independent human marks create a much stronger benchmark than a single mark.
What two examiners taught me about "ground truth"
The statistic that changed how I think about this came from the corpus underlying our English accuracy work. On the two 40-mark English responses in that corpus, the two qualified examiners gave exactly the same mark just 13% of the time. Across all eight English questions in the public dataset, their exact-match rate was 29%.
That 13% figure is not a universal estimate of human examiner agreement, and the extended-writing sample is small. I would not use it to claim that English examiners "only agree 13% of the time". What it does demonstrate very clearly is that a single examiner's mark cannot automatically be treated as an objective truth on every question.
This changes the interpretation of an AI result. Suppose an AI gives 27/40, examiner A gives 26/40 and examiner B gives 29/40. Saying the AI was "wrong by one mark" because it did not match examiner A is much less informative than looking at the pattern of agreement across all three markers.
The public benchmark we used was created from real GCSE answers that had already been marked independently by two qualified examiners. Medly's wider benchmark contains more than 30,000 twice-marked answers across English, Maths and the Sciences. That kind of double-marked data makes it possible to compare AI-to-human agreement with human-to-human agreement.
3. How is agreement being measured?
Even once the benchmark is credible, the phrase "90% accurate" can hide a remarkable amount.
Does 90% mean the AI gave exactly the same mark as one examiner? That it came within one mark? That it placed the response in the same grade band? That it achieved a particular inter-rater reliability score? Was the assessment made up mainly of one-mark questions, where exact agreement is much easier, or extended responses where markers have more room to differ?
The metric has to fit the marking task. Raw exact-match percentages can be intuitive, but they are not always comparable between subjects or question types. A credible study should therefore show more than one flattering headline number and explain what each measure does and does not mean.
4. What decision are you actually trying to make?
This is the question that is easiest to skip.
When a teacher or school leader asks me whether AI marking is accurate, they are rarely conducting a philosophical investigation into whether a machine can replicate human judgement. Usually they are trying to decide whether using AI will make their marking process better.
That often means they are really asking a second question: "Will this save us time without introducing unacceptable risk?"
That distinction matters. An AI system could achieve an impressive aggregate accuracy statistic and still create a terrible workflow if its mistakes are unpredictable, difficult to find, or concentrated in the exact kinds of answers teachers care about most. Equally, a system that sometimes differs from a teacher by one mark may still be extremely useful if the reasoning is transparent, the areas of uncertainty are easy to review, and the teacher can approve or change the result quickly.
The practical value of AI marking is not determined by an abstract accuracy percentage alone. It is determined by whether the pattern of disagreement is predictable and reviewable enough that professional oversight remains efficient.
What should a credible AI marking study tell you?
You do not need to become an assessment researcher to challenge an accuracy claim. A useful study should make it possible to answer a fairly simple checklist:
- What student work was tested, and how similar is it to the work I want to mark?
- Was the full dataset tested, or was a favourable sample selected?
- Who produced the human marks, and were there multiple independent markers?
- What marking guidelines or mark scheme were used?
- What does the reported accuracy metric actually measure?
- How does AI-to-human agreement compare with human-to-human agreement where that comparison is available?
- What parts of the workflow were not tested - for example handwriting recognition, script splitting or question matching?
- What limitations or unfavourable results are disclosed?
- When was the study run, and which system or model was actually tested?
I would also look for evidence that challenges the provider's preferred story. In our own Maths study, for example, a general AI model performed as well as or better than Marking.ai on some statistical measures. We publish that rather than hiding it. If the purpose of an evidence programme is genuinely to understand a system, an inconvenient result is still a result.
This is not an argument that AI marking is automatically safe or accurate
There is a danger of taking human disagreement and using it as a way to lower the bar for AI. That is not what I am arguing.
If two humans disagree, it does not follow that any AI answer inside that range is acceptable. Assessment validity, fairness, transparency, consistency and accountability still matter. Ofqual's 2026 working paper on AI use in marking makes this point particularly well: technical performance is only one part of judging whether AI has a legitimate role in a marking process.
The Cambridge university study is a useful counterweight too. Its researchers found only moderate agreement in the university contexts they tested and recommended local evaluation rather than assuming an AI system that works in one institution will work in another.
Those findings do not undermine the case for AI-assisted marking. They undermine the case for broad, context-free claims.
So, can AI mark accurately?
Sometimes, on some assessments, against some benchmarks, yes. On others, not yet. A single percentage cannot answer the question responsibly.
For me, the more useful way to think about it is:
- Test the kind of work you actually care about.
- Use a benchmark that reflects normal professional variation rather than assuming one marker is always ground truth.
- Understand exactly how agreement is being measured.
- Judge accuracy in the context of the outcome you are trying to achieve - including the amount and type of teacher review that remains necessary.
As AI-assisted marking becomes more common, I think the industry will need to get much more comfortable publishing evidence this way: subject by subject, question type by question type, with the human benchmark visible alongside the AI and the limitations written down.
Schools should expect that level of transparency. Providers should welcome it.
At Marking.ai, we are building our Evidence & Accuracy hub around that principle. The goal is not to prove that our system is perfect. It is to make the evidence clear enough that teachers and school leaders can decide where it is strong enough to help, where human review matters most, and where we still have work to do.
Sources and further reading
- Ofqual - Principles of AI use in marking (2026) - Working paper on technical capability, assessment validity, fairness, transparency and accountability in high-stakes marking.
- Marking.ai - Evidence & Accuracy - Marking.ai's subject-by-subject evidence hub.
- Marking.ai - GCSE English Language accuracy study - 240 public GCSE English answers, each independently marked by two examiners.
- Marking.ai - GCSE Mathematics accuracy study - 240 public GCSE Maths answers, each independently marked by two examiners.
- Medly AI - How well does AI grade real student answers? - Overview of a 30,000+ answer, twice-marked benchmark across five subjects.
- University of Cambridge - AI in University Assessment (2026) - Independent research testing automated marking across three UK universities.
Founder, Marking.ai
Edtech founder. Focused on AI assisted assessment, marking and feedback. Dedicated to helping teachers and school leaders understand how the use of AI can reduce workload and improve student outcomes.
Assessment terms explained
Plain-English definitions of the vocabulary used here.
- AI markingAI marking is the use of artificial intelligence to assess student work against defined criteria - a mark scheme or a teacher's own marking guidelines - and to draft marks and feedback for a teacher to review.
- Assessment for learningAssessment for learning is the practice of using evidence of what students currently understand to decide what to teach next, and to help students see what they need to do to improve.
- Assessment objectivesAssessment objectives are the categories of skill a qualification tests - commonly labelled AO1, AO2 and AO3 - each carrying a defined share of the available marks.
- Comparative judgementComparative judgement is an assessment method in which markers repeatedly choose the better of two pieces of work, and those paired decisions are combined statistically into a rank order.
- Criterion-referenced assessmentCriterion-referenced assessment judges work against defined standards of what a student should know or be able to do, rather than against how other students performed.
- ExemplarAn exemplar is a piece of work - often a real student answer - presented with its marks and commentary to show what a particular standard looks like in practice.
See it mark a real answer
Book a demo for your school, or start free and mark your own work - no card required.
