the.ai

Data / Evaluation

verified

Training Corpus Contamination

Evaluation sets end up in training corpora because both are scraped from the same internet. The model then answers from memory and the benchmark reports understanding. Unlike Benchmark Contamination, which is about a benchmark's integrity, this is about what a crawl swept up before anyone looked.

Viz primitive · budget-splitcontaminated = 4

contaminated holds 13% of the budget; rest holds the remaining 87%.

Evaluation items with a match in the training corpus against items with none, in items. Drag the contamination up to watch the score stop measuring the model — a small share at the left is enough to move a leaderboard.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:2107.06499 — Deduplicating Training Data Makes Language Models Better
  • arXiv:2101.00027 — The Pile: An 800GB Dataset of Diverse Text for Language Modeling