the.ai

Deployment / Monitoring

verified

Production Monitoring

The thing that makes monitoring a served model hard is that you usually do not know whether it was right. A label arrives days later, or after a human reviews it, or never. So the metric you most want — accuracy — is the one you cannot have in real time, and everything shipped is a proxy for it.

Viz primitive · budget-splitunlabelled-predictions = 20

unlabelled-predictions holds 33% of the budget; rest holds the remaining 67%.

Predictions that will never get a label, against the ones that eventually will, in predictions. Drag the label latency up to watch the unlabelled share take everything — and note the labelled remainder is not a random sample, but the cases somebody escalated.

20

Reviewed by opendroid · 2026-08-18

  • arXiv:2004.05785 — Learning under Concept Drift: A Review
  • arXiv:1906.02530 — Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift