the.ai

Data / Methods

verified

Data Mixing

A training corpus is several corpora — web text, code, books, papers, forums — and how much of each you use is a decision with large effects. Code improves reasoning on non-code tasks; too much of any single source costs breadth. The proportions are tuned like hyperparameters and reported far less often.

Viz primitive · update-spectrumbalancing = 0.2 · bars = 16

16 values. The left group decays steeply; the right group is 36% of the way to flat, and reads flatter than the left.

How training mass is spread across source domains, as sampled and as it is balanced. Drag the balancing up to watch the mixture even out — a natural crawl is steeply uneven, and both ends of this drag cost something.

0.2

Reviewed by opendroid · 2026-08-18

  • arXiv:2305.10429 — DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining