Alignment / Evaluation
verifiedSafety Benchmark Validity
A benchmark labelled "safety" is only measuring safety if a model can score well on it without simply being more capable. Many do not clear that bar — their scores track general capability closely enough that a bigger model improves on them for free, which means they measure progress that would have happened anyway.
The check is straightforward and rarely run: correlate the benchmark against a general capability measure across a range of models. A safety benchmark whose scores are nearly a linear function of capability is a capability benchmark with a different title, and reporting improvement on it as safety progress is the failure the term "safetywashing" names. Benchmarks that survive the check are the ones where the two come apart — where a more capable model can and does score worse.
The quantity is the share of variance in the benchmark not explained by capability, because that is the only part that could be measuring something else. A high correlation does not prove a benchmark is worthless — safety and capability are genuinely entangled in places — but it does mean the benchmark cannot distinguish them, and a measure that cannot distinguish its subject from a confound is not measuring its subject.
capability-explained holds 29% of the budget; rest holds the remaining 71%.
Benchmark variance explained by general capability, against the variance that could be measuring safety, in equal units. Drag the explained share up to watch capability account for the whole score — at the right the benchmark has a different title and the same content.
Reviewed by opendroid · 2026-08-18
- arXiv:2407.21792 — Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
- arXiv:2403.13793 — Evaluating Frontier Models for Dangerous Capabilities