the.ai

Evaluation / Validity

verified

Emergent Abilities

Some capabilities appear to switch on abruptly with scale: flat at chance across several model sizes, then suddenly working. That claim shaped a great deal of thinking about what scaling buys. It is also disputed — the sharpness may be an artefact of grading answers as right or wrong rather than of anything happening inside the model.

Viz primitive · loss-curvesteps = 1000 · params = 1
loss
step 01000

Loss over 1000 training steps, starting near 6.0. It falls to about 2.14, with 87% of the total improvement arriving in the first half.

Loss against training step. Drag model size to watch the floor fall smoothly — the cliff the benchmarks show is made by the metric, and it is not in this curve.

1

Reviewed by opendroid · 2026-08-04