the.ai

Provenance / Rights

verified

Membership in Training Data

Given a document and a model, was that document in the model's training set? Authors want to know it about their books, benchmark maintainers about their test sets, and everyone about their own writing. It is a simple question with no reliable answer from outside, which is why so much rests on provenance being recorded rather than inferred.

Viz primitive · threshold-sweepseparation = 1 · threshold = 0.9 · base-rate = 0.05
let throughcutflagged

196 of 1000 flagged. 14% of them were right and 169 were false alarms; 54% of what should have been caught was, leaving 23 missed.

Documents that were in the training set, against ones that were not, scored by how surprised the model is. Drag the distinctiveness up to watch them separate — ordinary writing sits at the left, inside the overlap, and that is where the question is usually asked.

1

Reviewed by opendroid · 2026-08-18