the.ai

Evaluation / Validity

verified

Benchmark Contamination

A model trained on most of the public internet has probably seen the test set. When that happens a benchmark stops measuring ability and starts measuring memorisation, and the score goes up either way. The problem is not that anyone cheated — it is that the training corpus is too large for anyone to know what is in it.

Viz primitive · budget-splitcontaminated = 10

contaminated holds 10% of the budget; rest holds the remaining 90%.

Share of a benchmark score attributable to memorised items against genuine ability. Drag the contamination rate to watch a weak model close the gap.

10

Reviewed by opendroid · 2026-08-04

  • arXiv:2311.04850 — Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
  • arXiv:2403.07974 — LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code