the.ai

Reasoning / Foundations

verified

Inference Scaling Law

Accuracy improves predictably with compute spent at inference, in the same regular way it improves with compute spent at training — and the two are substitutes. That means a serving budget is a choice between a larger model and a longer-thinking one, and the choice has a right answer that depends on the problem rather than on taste.

Viz primitive · budget-splitinference-compute = 30

inference-compute holds 23% of the budget; rest holds the remaining 77%.

Compute paid per query at inference, against the amortised share of training compute per query, in equal units. Drag the thinking budget up to watch the recurring cost overtake the one-off — which is why the crossover moves with traffic volume.

30

Reviewed by opendroid · 2026-08-18