Evaluation / Methods
verifiedHuman Evaluation
Ask people whether the output is good. It is the ground truth every automatic metric is trying to approximate, and it is expensive, slow, and much less reliable than its status suggests — two annotators given the same pair often disagree, and the same annotator on a different day sometimes disagrees with themselves.
What makes a study trustworthy is mostly procedural: a rubric specific enough that two people reading it converge, several annotators per item, an agreement statistic reported, and the raters' background stated because it changes the answer. Studies missing these are common and their numbers are not comparable with anything. LLM as Judge exists because this is expensive — and it is calibrated against exactly the human ratings whose reliability is in question, so it inherits every weakness here plus its own.
The interval around a mean rating depends on the number of raters as well as the number of items, and the rater term usually dominates because rater disagreement exceeds item difficulty variation. So twenty raters on ten items is a different measurement from ten raters on twenty, and a result reporting only the item count has reported the smaller half of its own uncertainty. Agreement statistics correct for chance because raw agreement on a two-option choice starts at 50% for coin flips.
rater-disagreement holds 17% of the budget; rest holds the remaining 83%.
Variance from raters disagreeing with each other, against variance from the items differing in difficulty, in equal units. Drag the rater disagreement up to watch it dominate — which is why the rater count matters more than the item count and is usually the one left unreported.
Reviewed by opendroid · 2026-08-18
- arXiv:2202.06935 — Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text
- arXiv:2009.01325 — Learning to summarize from human feedback