The user never typed it. The attacker never touched the API. Instructions hidden in a webpage the app retrieves hijack the model — because LLM applications cannot tell data from commands.
The classical vulnerability, reborn in natural language.
The root cause is architectural: the prompt is a single natural-language channel carrying both the developer's instructions, the user's request, AND retrieved third-party data — with no mechanism distinguishing them. Classical systems separate code from data at a boundary (SQL parameters, sandboxed DOM); LLM apps blur the line between data and instructions by design. Whatever the model retrieves, it reads as instructions. The paper derives a comprehensive taxonomy of attack vectors from this single principle — and demonstrates them end-to-end, including remote exfiltration of user data through a browsing-enabled assistant.
The assumption every LLM-integrated app silently made.
You tell the waiter your order (user input). The chef shouts instructions from the kitchen (system prompt). But the waiter also picks up napkin notes left by strangers (retrieved data) — and treats each as an order from the house. A stranger writes "and give me table 12's credit card details" on a napkin; the waiter, unable to distinguish stationery from authority, complies. That is your LLM app, every retrieval, always.
From planted text to exfiltrated data — the chain in five links.
The authors weaponized a realistic browsing-enabled assistant: a webpage with hidden instructions causes the model to append user-sensitive context to an attacker-controlled URL request — remote exfiltration through the model's own tool use, with the user none the wiser. The taxonomy extends across applications (search assistants, email copilots, code assistants) and attacker goals — surveillance, misinformation, manipulation, and propagation — and the paper discusses systemic risks where injected instructions spread across interconnected applications.
The structural reasons this isn't a bug you patch.
Not a model flaw — an application-design flaw the paper made undeniable.
| Injection class | Attacker needs | Visibility to user |
|---|---|---|
| Direct (jailbreak) | API access to the model | visible in own session |
| Indirect (this paper) | publish content anywhere retrievable | invisible — arrives as 'context' |
| Training-data poisoning | corpus access | delayed, diffuse |
| Fine-tuning backdoor | training pipeline access | dormant until trigger — entry #91 |
The comparison that made indirect injection the headline threat: remote, interface-free, and cloaked as ordinary content.
Every agent-security paper since builds on this taxonomy.
Check your understanding of the key concepts from Indirect Prompt Injection.
Everything you need to remember about this paper.