the.ai

Reasoning / Foundations

verified

Test-Time Compute

There are two ways to make a model answer harder questions: build a bigger one, or let the one you have think for longer. The second is a scaling axis in its own right, and for a while nobody was treating it as one — inference was a fixed cost per query rather than a dial. Turning it into a dial is the change that reorganised the field.

Viz primitive · budget-splitreasoning-tokens = 40

reasoning-tokens holds 29% of the budget; rest holds the remaining 71%.

Serving budget spent on tokens the model generates before answering, against the budget spent on parameters, in equal units. Drag the thinking up to watch it take the budget — worth it while the problem is within reach, and worth nothing past that boundary.

40

Reviewed by opendroid · 2026-08-18