the.ai

Deployment / Measurement

verified

Online Experiment

Split live traffic between the current model and the new one, and measure which does better on the thing you actually care about. It is the only measurement that is not a proxy — and it costs real traffic, real time, and exposes real users to whichever arm turns out to be worse.

Viz primitive · threshold-sweepseparation = 1 · threshold = 1.6 · base-rate = 0.3
let throughcutflagged

122 of 1000 flagged. 65% of them were right and 43 were false alarms; 26% of what should have been caught was, leaving 221 missed.

Experiments where the new model is genuinely better, against ones where the difference is noise, with the significance cut through both. Drag the effect size up to watch them separate — small effects sit inside the overlap, and no threshold gets them out; only more traffic does.

1

Reviewed by opendroid · 2026-08-18