Skip to content
AI.info

Evaluation

Human Evaluation, Rubrics, and Annotation Reliability

Build reliable human-evaluation protocols through clear constructs, sampling, blinding, rater training, agreement analysis, and adjudication.

By the end you can

“Which answer is better?” is not yet a rubric

One reviewer may prefer brevity, another factual detail, and a third a friendly tone. Their disagreement can reflect an underspecified task rather than careless work.

Human evaluation starts by defining the construct and separating properties that can conflict. That choice is not cosmetic. Rescore the same outputs under a different protocol and the leaderboard can come back in a different order.

That is what happened in 2021 to the top WMT 2020 systems in two language pairs. Professional translators replaced crowd workers. MQM error annotation replaced a rating. Full document context replaced the isolated segment. Freitag and colleagues published the rescoring. Nothing about the systems had changed. The protocol had.

Raters cannot apply criteria that the protocol never made explicit.

Example

Same outputs, a narrower task, a different rater pool — and a different winner

Two pairs of published studies did what a pilot is supposed to do: they changed one part of the protocol and watched the verdict move.

  • Unit of judgment: raters preferred human over machine translation more strongly when they judged whole documents rather than isolated sentences. Läubli and colleagues showed that in 2018.
  • Rater pool and rubric: the top WMT 2020 systems in two language pairs were rescored with MQM error annotation, carried out by professional translators reading full documents. Freitag and colleagues, 2021.
  • Outcome: the ranking moved. From their abstract: “We analyze the resulting data extensively, finding among other results a substantially different ranking of evaluated systems from the one established by the WMT crowd workers, exhibiting a clear preference for human over machine output.”
  • Reference standard: 7–8 ophthalmologists graded each image, and a simple majority of them defined the truth. Gulshan and colleagues, in JAMA in 2016: “A simple majority decision (an image was classified as referable if ≥50% of ophthalmologists graded it referable) served as the reference standard for both referability and gradability.” Against that standard the algorithm reached 90.3% sensitivity and 98.1% specificity on EyePACS-1.
  • Adjudication: the same task was re-graded to full consensus among retinal specialists, and a small set of adjudicated grades allowed substantial improvements in algorithm performance. Krause and colleagues, Ophthalmology, 2018.

Comparison

Four judgment formats

Format changes how hard the task is and how raters read it. The four are not interchangeable views of one underlying quality. Absolute ratings, pairwise preferences and rankings ask for a verdict. Error annotation asks what is wrong, where, and how badly.

The WMT rescoring is the demonstration that the choice decides the outcome. Swap the crowd rating for MQM error annotation by professional translators reading full documents, and the same top WMT 2020 systems in two language pairs come back in a substantially different ranking. Freitag and colleagues, 2021. The unit alone does part of that work: in 2018 Läubli and colleagues found the preference for human over machine translation was stronger when raters judged whole documents rather than isolated sentences.

Choose the format that matches the claim you intend to make, and report which one you used. A reader cannot reconstruct it from the scores.

FigureComparison · 4 columns

Absolute rating

Assign a score to one output on a scale.

  • Supports per-item scores
  • Scale use varies by rater
  • Needs anchors
  • Can suffer central tendency

Pairwise preference

Choose which of two outputs better satisfies a criterion.

  • Often easier
  • Supports comparative models
  • Order can bias choices
  • Does not give absolute adequacy

Ranking

Order several outputs from best to worst.

  • Efficient for small sets
  • Cognitive load grows quickly
  • Ties need policy
  • Can support preference models

Error annotation

Mark specific omissions, unsupported claims, boundary errors, or policy violations.

  • Actionable diagnosis
  • Requires detailed taxonomy
  • Higher training cost
  • Supports severity weighting

Visual

From construct to observable decision

A rubric connects abstract goals to examples, and it has to say what happens when the examples run out. Two diabetic-retinopathy studies mark the two ends of that decision. One resolved grader disagreement by simple majority over 7–8 ophthalmologist grades per image: Gulshan and colleagues, JAMA, 2016. The other re-graded the same task by adjudication to full consensus among retinal specialists: Krause and colleagues, Ophthalmology, 2018.

Both are defensible rules. They are not the same rule, and the second changed what the measured algorithm performance could reach.

FigureProcess · 5 steps
  1. 1. Name one property

    Factual support, usefulness, fluency, severity, or another construct.

  2. 2. Define inclusion

    Describe what counts and what does not.

  3. 3. Add anchors

    Provide examples near each rating boundary.

  4. 4. Handle ambiguity

    Specify ties, insufficient evidence, and unanswerable cases.

  5. 5. Pilot and revise

    Use disagreements to improve wording before the main study.

Protect the study from avoidable bias

Blind model identity where possible, randomize output order, balance left–right presentation, control context length, and avoid revealing scores from other raters. Track fatigue and learning over the session.

Oncology has already written those precautions into a rulebook. Tumour assessments generally should be verified by central reviewers blinded to study treatments — that is the FDA's December 2018 endpoints guidance. The same passage offers a cheaper route: “Alternatively, a random sample-based blinded central review auditing approach could be used with a detailed auditing plan prespecified including a strategy to detect potential assessment bias.”

How often does the independent re-read actually change the answer? Across 49 randomized Roche oncology trials and over 32,000 patients, blinded independent central review and local evaluation reached the same statistical inference in 87% of progression-free-survival comparisons, with a BICR/LE hazard-ratio ratio of 1.044 in double-blind studies. Lian and colleagues published that in The Oncologist in 2024. The point is not that blinding is usually redundant. It is that the disagreement rate is a measurable quantity, and someone measured it.

For sensitive topics, recruit raters with relevant language, cultural, or professional expertise and compensate them appropriately. Convenience annotators may not represent affected users.

The evaluation interface can change the judgment it claims merely to record.

Case

Teachers told the machine from the human; crowdworkers did not

Mechanical Turk workers could not tell machine-written text from human-written text. Teachers could. Karpinska and colleagues ran that comparison in 2021, after surveying 45 open-ended text generation papers and finding the crucial task details mostly unreported.

The filters did not rescue the workers. Even behind strict qualification filters, “AMT workers (unlike teachers) fail to distinguish between model-generated text and human-generated references”. One change did help: showing the human reference beside the model output improved the workers' calibration.

Key idea

Agreement is not validity — and the coefficient is a choice

High agreement can arise from an easy but irrelevant task, shared bias, or overly coarse labels. Low agreement can reveal genuine ambiguity, inadequate instructions, or heterogeneous expertise. It does not automatically destroy the comparison the ratings were collected for.

The TREC document pools are the long-running demonstration. NIST's Ellen Voorhees re-judged the TREC-4 and TREC-6 pools with two extra assessors per topic. Assessor overlap was only about 0.4. The system rankings barely moved: “The high correlations indicate that the comparative evaluation of retrieval performance is stable despite substantial differences in relevance judgments, and thus reaffirm the use of the TREC collections as laboratory tools.”

A quarter of a century later the result held. Parry and colleagues re-annotated 43 topics of TREC Deep Learning 2019 with 8 annotators, in 2025. Kendall's tau between system orderings was 0.897 in-sample and 0.879 across annotator combinations. Their reading: “despite low annotator agreement, under re-annotation, system ordering is highly correlated with the original DL'19 annotations”.

The threshold most readers reach for has a provenance rather than a proof. Jean Carletta imported it into computational linguistics in 1996, citing Krippendorff (1980): “content analysis researchers generally think of K > .8 as good reliability, with .67 < K < .8 allowing tentative conclusions to be drawn.” CL researchers went on to follow it, as Artstein and Poesio's 2008 survey records. The same survey notes that deciding an adequate level of agreement “is still little more than a black art”.

The coefficient also moves with decisions the protocol has to declare. Take one worked example of 40 units and 5 observers, from Hayes and Krippendorff in 2007. Treat the data as nominal and alpha is .4765. Treat the same data as ordinal and alpha is .7598, with a bootstrap 95% confidence interval of .7078 to .8078 and a probability q = .9473 of failing an alpha-min of .800. Artstein and Poesio put the same problem more bluntly for a single anaphora annotation experiment: “This is because depending on the way we measure agreement, we can report α values ranging from 0.122 to 0.998 for the very same experiment!”

So report the statistic, the unit, prevalence, missing ratings, and the confidence interval. Then inspect the disagreements. One number chosen from that range is not a verdict.

Raters can agree consistently on the wrong construct.

Case

Evaluators at chance level, agreeing with one another

Untrained evaluators “distinguished between GPT3- and human-authored text at random chance level”. Clark, August and colleagues tested the raters, not the systems, in 2021. Three quick training methods were tried. Accuracy improved only “up to 55%”.

That is what an agreement coefficient cannot tell you. Raters at chance can still agree with one another.

Analogy

Calibrating several judges at a competition

Before a competition begins, its judges review examples that define each deduction and every borderline case. Their consistency improves further when the rubric separates execution, difficulty, and safety.

A gymnastics panel has a routine that was actually performed, so a judge who drifts from the others is simply wrong. Many AI properties have no such performance underneath them, and a rater who dissents may be reporting a real difference in language, culture, or domain. Shared anchors are what make that difference readable instead of noisy.

Calibration aligns interpretation; it does not erase legitimate disagreement.

Steps

Run a defensible human study

Treat the protocol as an experiment, and write down in advance what oncology writes into its guidance: who re-reads, blinded to what, and what fraction is audited. The FDA's December 2018 endpoints guidance asks that tumour assessments generally be verified by central reviewers blinded to study treatments, with a prespecified random-sample audit as the stated alternative. What the second reading buys has been counted: the same statistical inference in 87% of progression-free-survival comparisons, and a BICR/LE hazard-ratio ratio of 1.044 in double-blind studies, across 49 randomized Roche oncology trials and over 32,000 patients. Lian and colleagues, The Oncologist, 2024.

Step five is where the reference standard is decided rather than discovered. A simple majority over 7–8 ophthalmologist grades per image and adjudication to full consensus among retinal specialists are two answers to one disagreement. Gulshan and colleagues took the first route in 2016. Krause and colleagues took the second in 2018, on the same task, and the adjudicated grades allowed substantial improvements in algorithm performance. Choose the rule before the ratings arrive, and report it beside the result.

FigureProcess · 5 steps
  1. 1. Sample cases

    Represent traffic, severity, languages, and known failure slices.

  2. 2. Assign raters

    Define expertise, independence, overlap, and workload.

  3. 3. Pilot the rubric

    Measure comprehension and revise ambiguous instructions.

  4. 4. Collect blinded judgments

    Randomize order and capture confidence or uncertainty when useful.

  5. 5. Analyze and adjudicate

    Report agreement, effect size, uncertainty, disagreements, and final decision rules.

Key takeaways