Fairness / Foundations
verifiedImpossibility Result
You cannot have all of them. The fairness criteria people most want simultaneously turn out to be mathematically incompatible whenever the groups have different base rates — which is nearly always, since differing base rates are usually the trace of the historical unfairness that prompted the question. This is not a limitation of current methods. It is a theorem, and no amount of engineering removes it.
The practical consequence is that a fairness argument which begins "the model should be calibrated AND have equal error rates" has already ended, and nobody in the room realises. What is available is a choice about which failure to accept, made explicitly and written down. The ProPublica–Northpointe dispute is the canonical case: both sides were right about their own criterion, and the disagreement was never resolvable by looking harder at the model.
Chouldechova's identity makes it exact. With prevalence p, false-negative rate FNR and positive predictive value PPV, the false-positive rate is forced: FPR = (p/(1−p))·((1−PPV)/PPV)·(1−FNR). Hold PPV equal across two groups whose p differs, and FPR and FNR cannot both match — the identity determines the third quantity from the other two. Kleinberg, Mullainathan and Raghavan prove the parallel result for score calibration with balance in both classes: the three hold together only when prediction is perfect or the base rates are equal.
330 of 1000 flagged. 28% of them were right and 238 were false alarms; 92% of what should have been caught was, leaving 8 missed.
The same detector at the same cut, with only the base rate changing. Drag it to watch precision move while the scores and the threshold stay exactly where they were — that dependence is the whole mechanism, and it is why two groups with different prevalence cannot match on everything at once.
Reviewed by opendroid · 2026-08-18
- arXiv:1609.05807 — Inherent Trade-Offs in the Fair Determination of Risk Scores
- arXiv:1703.00056 — Fair prediction with disparate impact: A study of bias in recidivism prediction instruments