Comparative judgement
Comparative judgement is an assessment method in which markers repeatedly choose the better of two pieces of work, and those paired decisions are combined statistically into a rank order.
What is comparative judgement?
Comparative judgement is a method where markers are shown two pieces of work at a time and simply decide which is better, rather than awarding marks against a scheme. Many such decisions, from several markers, are combined statistically into a single rank order and a score for each piece.
The method rests on a well-established finding: people are markedly better at judging which of two things is better than at placing a single thing on an absolute scale. Asked whether an essay is worth 19 or 21 marks, experienced markers disagree. Asked which of two essays is stronger, they agree far more often. Comparative judgement builds an assessment method on the second question instead of the first.
How it works in practice
Markers are shown pairs of scripts and choose the better one, with no rubric to apply and no marks to award. Each script appears in many pairings, judged by different markers. A statistical model then converts the accumulated decisions into a scale, producing a rank order and a score for every piece - along with a measure of how consistently the judges agreed.
Where it fits, and where it does not
It suits open tasks where quality is holistic and hard to itemise: extended writing, portfolios, design work. It is less useful where answers are right or wrong, because a mark scheme already handles those cleanly and far more cheaply. Its practical costs are real, too - it needs enough judges and enough judgements to be stable, and it produces a rank order rather than the criterion-referenced explanation a student can act on, so it is usually paired with feedback from another source.
Key takeaways
- Comparative judgement asks markers which of two pieces is better, rather than what mark each deserves.
- Many paired decisions are combined statistically into a rank order and scores.
- It suits holistic, open tasks such as extended writing and portfolios.
- It produces a rank order rather than criterion-referenced feedback, so it is usually paired with another source of feedback.
Frequently asked questions
Is comparative judgement more reliable than marking to a mark scheme?
For open, holistic tasks it typically achieves higher inter-rater reliability, because relative judgements are more consistent than absolute ones. For questions with defined creditworthy content a mark scheme is at least as reliable and considerably cheaper, so the comparison depends on the task.
What are the drawbacks of comparative judgement?
It needs a substantial number of judgements from several markers before the scale is stable, which is an organisational cost. It also yields a position in a rank order rather than an explanation of what was missing, so students need feedback from elsewhere if the assessment is to be formative.
Related terms
Inter-rater reliability
Inter-rater reliability is the extent to which different markers award the same mark to the same piece of work - a measure of how far a mark depends on who marked it.
Moderation
Moderation is the process by which teachers compare samples of marked work to check that a mark scheme has been applied consistently, so the same work would receive the same mark whoever marked it.
Levels of response marking
Levels of response marking awards marks by deciding which band descriptor an answer best fits as a whole, rather than by adding up credit for individual points.
See it mark a real answer
Book a demo for your school, or start free and mark your own work - no card required.

