Alignment / Evaluation
verifiedDangerous Capability Evaluation
Before deploying a model, test whether it can do the specific things that would be harmful if it could — not whether it will, but whether it is able to. It is a different question from alignment, and a more tractable one, because capability is demonstrable in a way that intent is not.
The evaluations are adversarial by construction: the model is given tools, scaffolding and multiple attempts, because a deployed system will have those and an evaluation that withholds them measures the wrong thing. The hard part is the negative result — showing a model cannot do something requires having elicited it properly, and a failed elicitation is indistinguishable from an absent capability. This is why elicitation effort is reported alongside the score, and why the score alone means little.
A negative result is a statement about the elicitation, not about the model, and its strength scales with how hard anyone tried. So the useful reading is the highest capability found by the best attempt, treated as a lower bound that later attempts can only raise — a bound that gets tighter with effort and never becomes a ceiling. Reporting it as a ceiling is the error the framing exists to prevent.
292 of 1000 flagged. 41% of them were right and 173 were false alarms; 79% of what should have been caught was, leaving 31 missed.
Models that genuinely hold a capability, against ones that do not, scored by an evaluation. Drag the elicitation effort up to watch them separate — at the left a negative result says the attempt was weak, and cannot be told apart from an absent capability.
Reviewed by opendroid · 2026-08-18
- arXiv:2403.13793 — Evaluating Frontier Models for Dangerous Capabilities
- arXiv:2212.09251 — Discovering Language Model Behaviors with Model-Written Evaluations