Prompt Injection

Prompt injection is an attack where untrusted text causes a language model to ignore its intended instructions and follow the attacker's instead. It's the #1 risk in the OWASP Top 10 for LLM applications, and it becomes serious once models can take actions: an AI agent that reads email, browses the web, or processes documents can be steered by instructions hidden in that content, then misuse its tools to leak data, send messages, or change records.

Unlike SQL injection, there's no reliable escaping mechanism: models process instructions and data as the same kind of tokens. No filter or prompt fully prevents injection today. Defense is therefore architectural. Limit what a compromised model can do, separate privileges, require confirmation for consequential actions, and treat model output as untrusted.

TL;DR

Quick Example

An indirect injection hidden in a web page an agent is asked to summarize:

If the agent can read the inbox (private data), reads this page (untrusted content), and renders markdown images (external communication), the page can exfiltrate data. That's all three legs of the trifecta.

A safer design for the same agent:

Core Concepts

Direct vs Indirect Injection

Indirect injection is the more dangerous class for agents, because the victim isn't the attacker: a user innocently asks for a summary, and the content attacks them.

Why It's Hard to Fix

Model providers keep improving robustness through training, which meaningfully raises the bar, but systems must be designed assuming some injections will succeed.

Impact Channels

The Lethal Trifecta

An agent is at high risk when it combines:

  1. Access to private data (inboxes, documents, databases, credentials).
  2. Exposure to untrusted content (web, email, uploads, third-party tools).
  3. An exfiltration channel (sending messages, making web requests, rendering external images or links).

Remove any one leg for a given task, and the worst-case outcome shrinks dramatically.

Defense in Depth

Architectural Controls (Most Important)

Input and Context Handling

Output Handling

Detection and Monitoring

Best Practices

Threat-Model Every Agent Capability

For each tool, ask what an attacker could make the agent do with it after injecting content, and whose data it could reach. Design mitigations per tool, not just per app.

Assume the Model Can Be Convinced

Security guarantees must come from code: permissions, confirmations, allowlists, sandboxes. Prompt instructions like "never reveal X" are helpful but not a security boundary.

Red-Team Continuously

Maintain a suite of injection attacks (direct, indirect, encoded, multi-step) relevant to your tools, run it in CI against prompt and model changes, and add real incidents to it. See LLM evaluation.

Educate Users About Confirmations

Confirmations only help if users read them. Show clear, specific previews of actions, and avoid confirmation fatigue from trivial prompts.

Common Mistakes

Relying on a System Prompt Rule

It helps against casual attempts, but determined attackers routinely bypass it. It must be backed by architectural limits.

Granting Broad Credentials to Agents

An agent with an admin API key turns any successful injection into full compromise. Scope credentials per user and per task.

Rendering Untrusted Markdown

Rendering model output that includes !img silently sends data to the attacker when the image loads. Restrict image and link domains, or render plain text.

FAQ

Can prompt injection be completely prevented?

Not with current technology. Models can't reliably distinguish trusted instructions from instructions embedded in data. Robustness training, detection, and careful prompting reduce success rates, but secure systems limit the impact of successful injections through permissions, isolation, and human approval.

What's the difference between prompt injection and jailbreaking?

Jailbreaking aims to bypass a model's safety policies, to get disallowed content. Prompt injection aims to subvert an application's instructions, often through third-party content, to make it misuse data or tools. They overlap in technique, but differ in target and impact.

Is indirect prompt injection a risk if my app only summarizes documents?

The risk is lower if the summarizer has no tools, no private data beyond the document, and no way to send data out, though injected content can still manipulate the summary a user relies on. Risk rises sharply once the same agent can also read private data or take actions.

How do I test my application for prompt injection?

Build test cases embedding malicious instructions in every untrusted input channel (user messages, documents, web pages, tool results, file names, image text), and check whether the agent performs unauthorized actions or leaks data. Use automated red-teaming tools, public injection datasets, and periodic manual testing by security staff.

Related Topics

References