Evaluation / Validity
verifiedEmergent Abilities
Some capabilities appear to switch on abruptly with scale: flat at chance across several model sizes, then suddenly working. That claim shaped a great deal of thinking about what scaling buys. It is also disputed — the sharpness may be an artefact of grading answers as right or wrong rather than of anything happening inside the model.
The original observation was that certain benchmarks stay at random until a threshold and then rise steeply. The rebuttal is that those benchmarks use discontinuous metrics — exact match, multiple choice — and that swapping to a continuous one turns the same runs into smooth curves. Both readings fit the same data, which is why this entry cites both.
Under exact match, a task requiring L correct tokens scores p super L for per-token accuracy p, so smooth improvement in p looks like a threshold in p super L . The metric manufactures the discontinuity. Whether anything discontinuous also happens in the model is a separate question the metric cannot answer.
Loss over 1000 training steps, starting near 6.0. It falls to about 2.14, with 87% of the total improvement arriving in the first half.
Loss against training step. Drag model size to watch the floor fall smoothly — the cliff the benchmarks show is made by the metric, and it is not in this curve.
Reviewed by opendroid · 2026-08-04
- arXiv:2206.07682 — Emergent Abilities of Large Language Models
- arXiv:2304.15004 — Are Emergent Abilities of Large Language Models a Mirage?