Glossary

Inter-rater reliability

Inter-rater reliability is the extent to which different markers award the same mark to the same piece of work - a measure of how far a mark depends on who marked it.

What is inter-rater reliability?

Inter-rater reliability is the degree to which independent markers agree when marking the same work. High reliability means a student's mark reflects their answer rather than which marker they happened to get. It is the property that standardisation and moderation are designed to protect.

Every marking system makes an implicit promise: that the mark describes the work, not the marker. Inter-rater reliability is the extent to which that promise holds. It is measured by having two or more markers assess the same scripts independently and comparing the results - how often they agree exactly, how far apart they are when they differ, and whether one marker is consistently more generous than another.

Why it is never perfect

Reliability is highest where mark schemes are most specific. Points-based marking of short answers achieves close agreement; levels-based marking of extended writing does not, because the judgement is holistic and band descriptors require interpretation. This is a property of the task, not a failing of the markers - assessing the quality of an argument involves judgement, and judgement varies.

How departments improve it

Standardisation raises reliability before marking begins, by having markers agree the standard against exemplars. Moderation checks it afterwards, by sampling and comparing. Clearer descriptors and worked exemplars help; so does reducing the number of independent first passes, since differences that arise in separate first readings have to be reconciled later.

Key takeaways

  • Inter-rater reliability measures how much a mark depends on who marked it.
  • It is measured by having markers assess the same work independently and comparing results.
  • Reliability is higher for points-based marking than for levels-based extended writing.
  • Standardisation raises it beforehand; moderation verifies it afterwards.
Answers

Frequently asked questions

Still stuck?

Ask us anything - a real person answers, in term time.

Ask a question

How is inter-rater reliability measured?

By having two or more markers mark the same scripts independently and comparing the outcomes - the proportion of exact agreement, the average size of the difference when they disagree, and any systematic bias where one marker is consistently harsher. Which measure matters most depends on the assessment.

Why is agreement lower on essays than on short answers?

Because the judgement is different in kind. A short answer either contains the creditworthy content or does not. An essay is placed in a band on a holistic reading of its quality, and band descriptors have to be interpreted against real writing - where two experienced markers can reasonably differ.

See it mark a real answer

Book a demo for your school, or start free and mark your own work - no card required.