the.ai

Behaviour / Failure

verified

Hallucination

A model states something false with exactly the fluency it uses for something true. There is no tell — no hedging, no change of register — because nothing in how the sentence was produced distinguishes the two. That absence of a signal is the whole problem, and it is why the word stuck.

Viz primitive · threshold-sweepseparation = 0.9 · threshold = 0.6 · base-rate = 0.2
let throughcutflagged

337 of 1000 flagged. 37% of them were right and 212 were false alarms; 63% of what should have been caught was, leaving 75 missed.

False statements against true ones, scored by the model's own confidence. Drag the separation up to watch them come apart — at the left, where the corpus actually sits, confidence tells you almost nothing, which is why detection needs a signal from outside the generation.

0.9

Reviewed by opendroid · 2026-08-18

  • arXiv:2311.05232 — A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
  • arXiv:2202.03629 — Survey of Hallucination in Natural Language Generation