the.ai

Recommenders / Evaluation

verified

Offline-Online Gap

A change that improves every offline metric can lose an A/B test, and this happens often enough that mature teams treat offline numbers as a filter rather than as evidence. The reason is that offline evaluation replays a world the new model would not have produced.

Viz primitive · budget-splitin-support = 6

in-support holds 25% of the budget; rest holds the remaining 75%.

Decisions the logs can speak to against decisions outside anything ever shown, in decisions. Drag the coverage up to watch offline evaluation become informative — outside this range, no correction recovers the answer.

6

Reviewed by opendroid · 2026-08-18

  • arXiv:1907.06902 — Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches
  • arXiv:1902.10730 — Degenerate Feedback Loops in Recommender Systems