the.ai

Alignment / Evaluation

verified

Model-Written Evaluation

Writing evaluations by hand is slow, and the behaviours worth measuring keep multiplying. So use a model to generate the test items, and a model to filter them, and check a sample by hand. It scales the one thing that had been the bottleneck, and it introduces a dependency worth naming: the tests now share a source with the thing being tested.

Viz primitive · budget-splitshared-blind-spots = 8

shared-blind-spots holds 17% of the budget; rest holds the remaining 83%.

Failures the generator and the model under test share, and so cannot detect, against failures the evaluation can catch, in failures. Drag the shared region up to watch it grow — a score from a test that shares provenance carries less than the same score from an independent one.

8

Reviewed by opendroid · 2026-08-18

  • arXiv:2212.09251 — Discovering Language Model Behaviors with Model-Written Evaluations
  • arXiv:2407.21792 — Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?