the.ai

Reasoning / Selection

verified

Verifier Model

Generate many candidate answers and train a second model to say which one is right. It works for the same reason marking is easier than solving: checking a chain of reasoning is a different and usually simpler task than producing one, so a verifier can be smaller than the generator and still beat it at telling good from bad.

Viz primitive · threshold-sweepseparation = 1.8 · threshold = 0.8 · base-rate = 0.05
let throughcutflagged

236 of 1000 flagged. 18% of them were right and 194 were false alarms; 84% of what should have been caught was, leaving 8 missed.

Candidate answers that are correct, against ones that are not, scored by the verifier. Drag its quality up to watch them separate — at the left, with the generator right one time in twenty, most of what the verifier accepts is still wrong however many samples you draw.

1.8

Reviewed by opendroid · 2026-08-18