the.ai

Evaluation / Methods

verified

Human Evaluation

Ask people whether the output is good. It is the ground truth every automatic metric is trying to approximate, and it is expensive, slow, and much less reliable than its status suggests — two annotators given the same pair often disagree, and the same annotator on a different day sometimes disagrees with themselves.

Viz primitive · budget-splitrater-disagreement = 8

rater-disagreement holds 17% of the budget; rest holds the remaining 83%.

Variance from raters disagreeing with each other, against variance from the items differing in difficulty, in equal units. Drag the rater disagreement up to watch it dominate — which is why the rater count matters more than the item count and is usually the one left unreported.

8

Reviewed by opendroid · 2026-08-18

  • arXiv:2202.06935 — Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text
  • arXiv:2009.01325 — Learning to summarize from human feedback