the.ai

Optimization / Generalization

verified

Scaling Laws

Loss falls predictably as models, data and compute grow — smoothly, along a straight line on a log-log plot, over many orders of magnitude. That predictability is what makes it possible to justify a training run costing millions before it starts: the result can be extrapolated from small ones.

Viz primitive · loss-curvesteps = 1000 · params = 1
loss
step 01000

Loss over 1000 training steps, starting near 6.0. It falls to about 2.15, with 88% of the total improvement arriving in the first half.

Loss against training compute for one model size. Drag the parameter count to watch the whole curve shift down along the power law.

1

Reviewed by opendroid · 2026-08-04