AI in schools

Aligning AI Marking with UK Curriculum Standards

How AI marking fits with AQA, Edexcel and OCR mark schemes: the difference between points-based and levels-based marking, what the Ofqual and DfE rules actually require, and why the teacher remains the marker.

The Marking.ai team

Marking.ai

9 August 2025 · 5 min read

Marking GCSE and A-level work is one of the most time-consuming things a teacher does, and it is also the part where consistency matters most. The question a head of department asks about any AI marking tool is not whether it is impressive. It is whether it applies the same mark scheme the exam board does, and who is accountable when it gets one wrong.

The UK exam landscape

In England, GCSEs and A-levels are awarded by several boards:

All of them are regulated by Ofqual, whose job is to hold national standards steady across boards and across years.

Why mark schemes matter

A mark scheme is the rule book, and its central demand is consistency rather than correctness. Pearson's general marking guidance, printed at the front of its mark schemes, puts it plainly: all candidates must receive the same treatment, and examiners must mark the first candidate in exactly the same way as they mark the last.

That is a harder standard than it sounds. It is the standard a tired human marker on script twenty-eight is least able to meet, and it is the one specific thing a machine is structurally well-suited to hold.

Two kinds of mark scheme

Points-based (maths, sciences)

A mark for each correct element. A three-mark calculation might award one for the correct formula, one for substitution and one for the final answer - with method marks surviving an arithmetic slip. These schemes are the most mechanical to apply and the most checkable.

Levels-based (English, essays, extended responses)

Answers are placed in a level or band against descriptors, and the marker works up until they find the best fit. This is where marking becomes genuinely a judgement, where two experienced teachers most often disagree, and where any marker - human or otherwise - should be treated as producing a proposal rather than a verdict.

How an AI marker is pointed at a mark scheme

This is the part most often described badly, including by us in an earlier version of this post. Here is what actually happens in Marking.ai:

  • Your marking guidance and the extracted per-question criteria are attached to the request. Not a general sense of what a good answer looks like - the specific scheme you supplied.
  • The model marks criterion by criterion and returns the evidence for each mark, anchored to the student's own words and the page it appeared on.
  • A teacher reviews and approves every mark before it reaches a student.

What does not happen is training. Student work is never used to train a model - not ours, and not the provider's, whose terms bar it contractually on the paid tier the product runs on. You can read how we handle student data on our security pages. If a tool tells you it gets better by learning from the scripts you upload, that is a different arrangement with your students' data, and it is worth asking exactly what it means.

What the rules actually require

Two sets of expectations apply, and they say compatible things.

Ofqual has been explicit that its regulations do not permit AI to be used as the sole marker of a regulated qualification, and that schools and colleges must not use AI as the sole marker when marking a non-exam assessment component. Its published principles of AI use in marking and its approach to regulating AI in the qualifications sector set out the reasoning, and transparency is central to it: a system nobody can interrogate is a problem for trust in the qualification regardless of how well it scores.

The Department for Education's guidance on generative AI in education lands in the same place for day-to-day school marking: AI can help teachers plan, draft, adapt and review, but it must not make unsupervised decisions about learners, must not replace teacher assessment, and must not handle personal data without a lawful and secure process.

Read together, the requirement is not "do not use it". It is that a person remains the marker and can account for the mark. That is a design constraint, and it is the one worth checking any tool against.

No black box

Which is why transparency is not a nice-to-have here. A marking tool that is useful under those rules has to show its working:

  • Question-by-question marks, not a single overall score.
  • Each mark tied to the criterion in the scheme that earned it.
  • The specific evidence in the student's answer highlighted.

That is what makes a mark defensible in a moderation meeting or a conversation with a parent - "two marks for AO1 here, one for AO3 evaluation there", with the text in front of you. A number with no reasoning attached is not reviewable, and under the rules above it is not usable either.

Supporting moderation

The practical benefit for a department is a common starting point. Everyone begins from the same criteria-driven first pass, moderators can see the reasoning behind each mark rather than reconstructing it, and the record of what was awarded and why survives a staffing change.

None of that removes the moderation meeting. It removes the twenty minutes at the start of it that everyone spends working out what was actually marked.

What we will not tell you

We do not publish an accuracy figure for Marking.ai, because there is not one we would stand behind. Numbers produced on a research harness measure the harness's prompts rather than the marking path teachers actually use, and we would rather say so than quote something flattering. How we intend to measure, and what we will measure on, sits with the rest of our security and evidence material - and figures will appear there when studies report, not before.

If that is unsatisfying, it should be. It is also the honest position, and a supplier who gives you a confident percentage without telling you what was measured and on which paper has not given you more information than we just did.

Marking smart

Aligned with AQA, Edexcel and OCR standards, an AI marker can apply criteria consistently across a whole class, produce transparent and detailed feedback, and meet the Ofqual and DfE expectation that a teacher stays in charge of the mark. The future of marking is not a machine replacing a professional judgement. It is the repetitive first pass coming off a Sunday evening, so the judgement gets the attention it deserves.

The Marking.ai team

Marking.ai

We build Marking.ai. Several of us taught; all of us have watched a colleague lose a Sunday to a stack of scripts.

Answers

Common questions

Still stuck?

Ask us anything - a real person answers, in term time.

Ask a question

Can AI legally mark GCSE and A-level work in the UK?

For regulated qualifications, Ofqual's regulations do not permit AI to be used as the sole marker, and schools and colleges must not use AI as the sole marker for a non-exam assessment component. For internal and formative marking, the DfE's guidance allows AI to support marking provided a teacher reviews the output and it does not make unsupervised decisions about learners. In both cases a person remains the marker.

Does AI marking use the real exam board mark schemes?

It uses the marking guidance you give it. In Marking.ai, your marking guide and the per-question criteria extracted from your assessment are attached to every marking request, and each mark is returned against a specific criterion with the evidence from the student's answer. It is your scheme applied consistently, not a general model of what a good answer looks like.

Is student work used to train the AI?

No. Student work is never used to train a model - neither ours nor the model provider's, whose terms contractually bar it on the paid tier the product runs on. Any accuracy testing is done separately and disclosed. If another tool says it improves by learning from the scripts you upload, ask precisely what that means for your students' data.

What is the difference between points-based and levels-based mark schemes?

A points-based scheme awards a mark for each correct element - a formula, a substitution, a final answer - and is common in maths and the sciences. A levels-based scheme places an answer in a band against descriptors and is used for essays and extended responses. Points-based schemes are more mechanical to apply; levels-based ones involve real judgement, which is where a marker's output should be treated as a proposal to check.

See it mark a real answer

Book a demo for your school, or start free and mark your own work - no card required.