the.ai

Data / Methods

verified

Data Filtering

Most of a web crawl is not worth training on — boilerplate, spam, machine translation, pages that are mostly navigation. Filtering decides what survives, and the decision moves benchmark scores more than most architectural changes do. It is also the least published part of any frontier model.

Viz primitive · budget-splitdiscarded = 8

discarded holds 50% of the budget; rest holds the remaining 50%.

Documents the pipeline throws away against documents it keeps, in documents. Drag the filtering up to watch most of the crawl go — and past some point what leaves is the diversity that made the crawl worth having.

8

Reviewed by opendroid · 2026-08-18

  • arXiv:2406.11794 — DataComp-LM: In search of the next generation of training sets for language models
  • arXiv:2101.00027 — The Pile: An 800GB Dataset of Diverse Text for Language Modeling