the.ai

Deployment / Practice

verified

Regression Suite

A fixed set of cases the model must still get right before a new version ships. It is not a benchmark and it is not trying to measure quality — it is trying to stop a specific thing that broke once from breaking again, which is a much narrower and much more achievable goal.

Viz primitive · budget-splitincident-cases = 6

incident-cases holds 17% of the budget; rest holds the remaining 83%.

Cases added after an incident, against the cases someone wrote up front, in cases. Drag the incidents up to watch the suite become mostly scar tissue — which is the healthy shape, because it grows toward the failures the system actually has.

6

Reviewed by opendroid · 2026-08-18

  • arXiv:2011.03395 — Underspecification Presents Challenges for Credibility in Modern Machine Learning
  • arXiv:2306.05685 — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena