the.ai

Provenance / Foundations

verified

Memorisation

Models reproduce parts of their training data verbatim. Not as a bug in a few pathological cases, but routinely and predictably, and more as they get larger. It is the same capacity that lets a model recall a fact, working on a passage instead — which is why it cannot simply be trained away without losing the thing it is a side effect of.

Viz primitive · budget-splitduplicated-sequences = 8

duplicated-sequences holds 17% of the budget; rest holds the remaining 83%.

Training sequences that appeared many times in the corpus, against the ones that appeared once, in sequences. Drag the duplication up to watch the memorisable share grow — which is why deduplication is the cheapest intervention against extraction.

8

Reviewed by opendroid · 2026-08-18