the.ai

Provenance / Foundations

verified

Data Provenance

Knowing what is actually in a training corpus, and where each part came from. It sounds like bookkeeping and turns out to be research: web-scale corpora are assembled from crawls and dumps whose contents nobody has read, so what a model trained on is genuinely unknown to the people who trained it until somebody goes and looks.

Viz primitive · budget-splitunattributed-documents = 20

unattributed-documents holds 33% of the budget; rest holds the remaining 67%.

Corpus documents whose origin was never recorded, against the ones with a known source, in documents. Drag the unattributed count up to watch it take over — provenance is recorded at collection or not at all, so this bar only moves one way.

20

Reviewed by opendroid · 2026-08-18