the.ai

Alignment / Evaluation

verified

Safety Benchmark Validity

A benchmark labelled "safety" is only measuring safety if a model can score well on it without simply being more capable. Many do not clear that bar — their scores track general capability closely enough that a bigger model improves on them for free, which means they measure progress that would have happened anyway.

Viz primitive · budget-splitcapability-explained = 12

capability-explained holds 29% of the budget; rest holds the remaining 71%.

Benchmark variance explained by general capability, against the variance that could be measuring safety, in equal units. Drag the explained share up to watch capability account for the whole score — at the right the benchmark has a different title and the same content.

12

Reviewed by opendroid · 2026-08-18

  • arXiv:2407.21792 — Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
  • arXiv:2403.13793 — Evaluating Frontier Models for Dangerous Capabilities