the.ai

Inference / Adaptive

verified

Model Cascade

Ask a cheap model first. If its answer looks good enough, keep it; if not, escalate to an expensive one. Most requests are answered by the cheap model, most of the cost is avoided, and the quality is close to the expensive model's — provided the check that decides is any good.

Viz primitive · budget-splitescalated-requests = 8

escalated-requests holds 17% of the budget; rest holds the remaining 83%.

Requests that escalate to the expensive model, against the ones the cheap one settles, in requests. Drag the escalation rate up to watch the saving disappear — past the crossing the cascade costs more than the expensive model alone, because it has paid for both, and where that crossing sits is set by the price ratio.

8

Reviewed by opendroid · 2026-08-18

  • arXiv:2305.05176 — FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
  • arXiv:2403.02419 — Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems