Inter-rater reliability
Inter-rater reliability is the extent to which different markers award the same mark to the same piece of work - a measure of how far a mark depends on who marked it.
What is inter-rater reliability?
Inter-rater reliability is the degree to which independent markers agree when marking the same work. High reliability means a student's mark reflects their answer rather than which marker they happened to get. It is the property that standardisation and moderation are designed to protect.
Every marking system makes an implicit promise: that the mark describes the work, not the marker. Inter-rater reliability is the extent to which that promise holds. It is measured by having two or more markers assess the same scripts independently and comparing the results - how often they agree exactly, how far apart they are when they differ, and whether one marker is consistently more generous than another.
Why it is never perfect
Reliability is highest where mark schemes are most specific. Points-based marking of short answers achieves close agreement; levels-based marking of extended writing does not, because the judgement is holistic and band descriptors require interpretation. This is a property of the task, not a failing of the markers - assessing the quality of an argument involves judgement, and judgement varies.
How departments improve it
Standardisation raises reliability before marking begins, by having markers agree the standard against exemplars. Moderation checks it afterwards, by sampling and comparing. Clearer descriptors and worked exemplars help; so does reducing the number of independent first passes, since differences that arise in separate first readings have to be reconciled later.
Key takeaways
- Inter-rater reliability measures how much a mark depends on who marked it.
- It is measured by having markers assess the same work independently and comparing results.
- Reliability is higher for points-based marking than for levels-based extended writing.
- Standardisation raises it beforehand; moderation verifies it afterwards.
Frequently asked questions
How is inter-rater reliability measured?
By having two or more markers mark the same scripts independently and comparing the outcomes - the proportion of exact agreement, the average size of the difference when they disagree, and any systematic bias where one marker is consistently harsher. Which measure matters most depends on the assessment.
Why is agreement lower on essays than on short answers?
Because the judgement is different in kind. A short answer either contains the creditworthy content or does not. An essay is placed in a band on a holistic reading of its quality, and band descriptors have to be interpreted against real writing - where two experienced markers can reasonably differ.
Related terms
Moderation
Moderation is the process by which teachers compare samples of marked work to check that a mark scheme has been applied consistently, so the same work would receive the same mark whoever marked it.
Standardisation
Standardisation is the process of agreeing how a mark scheme should be applied before marking begins - typically by marking exemplar work together - so that every marker works to the same standard.
Comparative judgement
Comparative judgement is an assessment method in which markers repeatedly choose the better of two pieces of work, and those paired decisions are combined statistically into a rank order.
See it mark a real answer
Book a demo for your school, or start free and mark your own work - no card required.

