the.ai

Reasoning / Training

verified

Reasoning RL

Rather than prompting a model to reason, train it to, by rewarding trajectories that reach verifiable answers. The surprise was how much emerged without being asked for: models trained this way lengthen their own reasoning, start checking their work, and backtrack when a line fails — behaviours nobody wrote a reward for.

Viz primitive · budget-splitunverifiable-problems = 40

unverifiable-problems holds 40% of the budget; rest holds the remaining 60%.

Problems with no automatic checker, against the ones a test suite or a numeric comparison can score, in problems. Drag the domain wider to watch the checkable set become the small part — which is the boundary of where this training signal reaches at all.

40

Reviewed by opendroid · 2026-08-18