the.ai

Optimization / Regimes

verified

Pretraining

Pretraining is the long, expensive first phase where a model learns language itself by predicting the next token over an enormous corpus. No labels, no task, no human supervision — just the text, and an objective that turns out to require learning a great deal in order to do well at.

Viz primitive · loss-curvesteps = 1000 · lr = 0.001 · params = 1
loss
step 01000

Loss over 1000 training steps, starting near 6.0. It falls to about 2.14, with 87% of the total improvement arriving in the first half.

Loss across a full pretraining run, falling fast and then slowly. Drag model size to watch the floor fall — how little arrives in the final stretch does not.

1

Reviewed by opendroid · 2026-08-04