the.ai

Reasoning / Selection

verified

Process Supervision

Reward each step of the reasoning rather than only the final answer. It is more expensive to label — someone has to read every line — and it fixes a specific problem that outcome rewards create, which is that a model reaching the right answer through nonsense gets exactly the same credit as one that reasoned correctly.

Viz primitive · budget-splitstep-labels = 3

step-labels holds 13% of the budget; rest holds the remaining 87%.

Feedback signals a step-labelled trajectory carries, against the single bit an outcome-only reward gives it, in signals. Drag the reasoning length up to watch process supervision's advantage grow — outcome feedback stays at one bit however long the chain gets.

3

Reviewed by opendroid · 2026-08-18