the.ai

Context / Evaluation

verified

Needle in a Haystack

Hide one distinctive sentence in a very long document and ask the model to find it. It is the standard long-context demonstration, it produces a satisfying green grid, and it is much easier than almost anything a reader actually wants to do with a long context.

Viz primitive · threshold-sweepseparation = 3.2 · threshold = 1 · base-rate = 0.5
let throughcutflagged

565 of 1000 flagged. 86% of them were right and 77 were false alarms; 98% of what should have been caught was, leaving 12 missed.

Contexts where the model finds what it was asked for, against ones where it does not, by how distinctive the target is. Drag the distinctiveness up to a single odd sentence and watch the task become trivial — which is where this benchmark sits, and why saturating it proves little.

3.2

Reviewed by opendroid · 2026-08-18

  • arXiv:2404.06654 — RULER: What's the Real Context Size of Your Long-Context Language Models?
  • arXiv:2307.03172 — Lost in the Middle: How Language Models Use Long Contexts