the.ai

Data / Evaluation

verified

Annotation Quality

Labels are produced by people who disagree, get tired, and interpret instructions differently. On many datasets the gap between two annotators is larger than the gap between the top models, which means the leaderboard is partly measuring who best fits the annotation noise.

Viz primitive · budget-splitcontested-items = 6

contested-items holds 25% of the budget; rest holds the remaining 75%.

Items annotators disagreed on against items they agreed on, in items. Drag the disagreement up to watch the ceiling come down — a model cannot be right about an item where right was never established.

6

Reviewed by opendroid · 2026-08-18

  • arXiv:2406.11794 — DataComp-LM: In search of the next generation of training sets for language models