the.ai

Agents / Safety

verified

Prompt Injection

A model reads instructions and data through the same channel, because to a language model there is no difference between them. So text inside a fetched web page or a retrieved document can tell the model what to do, and it may comply. Giving that model tools turns a text problem into an action problem.

Viz primitive · budget-splituntrusted-tokens = 3000

untrusted-tokens holds 75% of the budget; rest holds the remaining 25%.

Context that came from somewhere the user does not control against context that did, in tokens. Drag the untrusted volume to watch the trusted part become the minority.

3000

Reviewed by opendroid · 2026-08-04

  • arXiv:2302.12173 — Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection