the.ai

Uncertainty / Decisions

verified

Selective Prediction

Give the model the option to say nothing. A classifier that answers everything at 92% may be less useful than one that answers four fifths of the time at 99% and hands the rest to a person — if the escalation is cheap enough. The design question is not how accurate the model is but where the threshold goes, and that is a question about the cost of being wrong rather than about the model.

Viz primitive · threshold-sweepseparation = 1.8 · threshold = 0.2 · base-rate = 0.15
let throughcutflagged

494 of 1000 flagged. 29% of them were right and 353 were false alarms; 94% of what should have been caught was, leaving 9 missed.

Inputs the model gets wrong, against inputs it gets right, ranked by its own confidence; everything left of the cut goes to a person. Drag the cut right to send fewer — the false alarms fall and so does the share of real mistakes caught. No cut removes both.

0.2

Reviewed by opendroid · 2026-08-18