the.ai

Landscape / Regimes

verified

Grokking

A network fits the training set perfectly, sits at chance on held-out data for thousands of steps, and then abruptly generalises. Nothing about the training loss predicts the moment. It was found on small algorithmic tasks and it is the clearest evidence that memorising and understanding are different states a model passes between.

Viz primitive · budget-splitgeneralising-weight = 4

generalising-weight holds 13% of the budget; rest holds the remaining 87%.

Weight on the generalising circuit against weight on the memorised lookup, in equal units. Drag the generalising share up to watch the transition happen — during the plateau this bar is moving while the training loss is not.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:2201.02177 — Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets