the.ai

Alignment / Evaluation

verified

Dangerous Capability Evaluation

Before deploying a model, test whether it can do the specific things that would be harmful if it could — not whether it will, but whether it is able to. It is a different question from alignment, and a more tractable one, because capability is demonstrable in a way that intent is not.

Viz primitive · threshold-sweepseparation = 1.6 · threshold = 0.8 · base-rate = 0.15
let throughcutflagged

292 of 1000 flagged. 41% of them were right and 173 were false alarms; 79% of what should have been caught was, leaving 31 missed.

Models that genuinely hold a capability, against ones that do not, scored by an evaluation. Drag the elicitation effort up to watch them separate — at the left a negative result says the attempt was weak, and cannot be told apart from an absent capability.

1.6

Reviewed by opendroid · 2026-08-18

  • arXiv:2403.13793 — Evaluating Frontier Models for Dangerous Capabilities
  • arXiv:2212.09251 — Discovering Language Model Behaviors with Model-Written Evaluations