How do you prevent prompt injection in MCP servers and AI agents?

There's no sure fix for prompt injection. Parameterized queries mostly killed SQL injection. Prompt injection has no such fix. So the goal is containment, not a cure. The full set of controls sits below. The sections below unpack each one.

  • Least-privilege scoped tokens. Give the agent the narrowest credential that still works. A hacked agent can only reach what its token can reach.
  • Per-session and per-repo isolation. One session should not span your public tracker and your private repos. Keep them apart. Shrink what a single run can touch.
  • Allowlist-based URL and SSRF controls. Block private and reserved IP ranges. Force HTTPS. Don't hand-roll IP parsing.
  • Human approval on high-risk tool calls. Keep a person able to deny a tool call before it fires.
  • Trust separation. Treat tool descriptions and runtime results as different trust domains. They are different.
  • Constrained feature. Use Dual LLM or CaMeL. An injected instruction then has almost nothing it can do.

One caveat up front. This is an engineering threat model, not full security advice. It's the view I'd want a teammate to have. Read it before they ship an MCP server.

Direct vs indirect prompt injection, explained

Prompt injection in agent systems comes in two flavors. The dangerous one is the quiet one.

Direct injection is the obvious case. A human types bad commands straight into the prompt. On a shared or multi-user agent, that's a real problem. But at least the hacker is a user you can see.

Indirect injection is where bad commands ride in through content the agent reads as data. Tool outputs, files, emails, GitHub issues, web pages: all of it. The model can't reliably tell trusted commands from unsafe data. It all lands in one context window. There's no <untrusted> tag the model respects by default. Simon Willison flagged this risk in MCP on 9 April 2025. Researchers have since shown the same injection can work as attack and as defense. That shows how blurry the direct versus indirect line gets in practice.

Indirect wins as the top MCP threat, for one structural reason. MCP tool results flow straight into the model's context. You connect a server. The model calls a tool. Whatever comes back joins the conversation the model reasons over. If a hacker controls that returned data, they steer what the model does next.

The lethal trifecta: why injection becomes theft

Injection on its own is mostly a nuisance. It turns into a data breach only when paired with the right ingredients. Willison's lethal trifecta names the three conditions. He coined the term on 16 June 2025. Together, they turn a clever prompt into real damage:

  1. Access to private data the agent can read.
  2. Risk from unsafe content that can smuggle commands.
  3. An exfiltration vector, some way to send data back out.

An agent holding all three is exploitable. Here's the useful part. Knock out any one leg, and the data-theft outcome collapses. Can't reach private data? Nothing to steal. No unsafe content in context? Then nothing can smuggle commands in. And with no way out, the hacker has nowhere to ship the loot. Most guards in this post just remove one leg. They do it in a disciplined way. Design so your agent rarely holds all three at once. Do that, and you've already beaten any classifier.

Tool poisoning and rug pulls: the MCP-specific trust gap

MCP adds two injection variants you don't get from a plain chatbot. Both exploit the same blind spot.

Tool poisoning hides bad commands inside an MCP tool's description or metadata. The model reads that description. The user usually doesn't. So the tool description steers the model at connect time. This happens before anyone runs anything. Invariant Labs shared this class on 6 April 2025. It's now listed as MCP03 in the OWASP MCP Top 10. Note: the OWASP MCP Top 10 is still a beta project, as of 2026. It's not a finished catalog, unlike the flagship Agentic list. Here's the root gap. Someone might review a tool description once, at connect time. Runtime replies pour into the model context with no equal check. Reviewing the description once catches nothing those runtime replies smuggle in later.

And the surface keeps widening. Palo Alto's Unit 42 has since mapped fresh injection vectors through MCP sampling. See a fuller breakdown of the description trick in MCP tool poisoning.

Rug pulls are worse, in a lazy way. An MCP tool can mutate its own definition after you approve it. Safe on day one, bad on day seven. That's because clients re-fetch tool definitions and skip checking them again. Willison described this class in that same 9 April 2025 post. Approval stops being a one-time event. Not once the thing you approved can rewrite itself overnight.

Real incidents: EchoLeak and the GitHub MCP toxic agent flow

Two disclosures made this concrete. Both are proof-of-concept, shared responsibly. Neither shows any sign of abuse in the wild. That distinction matters. It's worth stating plainly.

EchoLeak (CVE-2025-32711, CVSS 9.3) was a zero-click indirect prompt-injection chain in Microsoft 365 Copilot. Aim Labs found it. Cato Networks hosts the origin writeup, after acquiring Aim. Aim Labs named the root-cause technique an LLM Scope Violation. Unsafe email content pulls special inside data outside its planned scope. A single crafted email, no click required, steered Copilot. It read inside RAG context and sent it out. The chain walked past several best-practice guards at once. Microsoft's XPIA cross-prompt-injection classifier. External-link redaction. Content-Security-Policy. Reference-style markdown. Per the same Cato writeup, Aim built the PoC in January 2025. They shared it privately, then published in June 2025. Microsoft patched server-side by May 2025. They reported no in-the-wild abuse.

The GitHub MCP toxic agent flow came from Invariant Labs on 26 May 2025. A bad prompt, planted in a public GitHub issue, hijacked an agent. It was running the official github-mcp-server (14,000-plus stars at the time). The agent read the user's private repos. It leaked their contents into an attacker-visible pull request. Invariant was blunt: this isn't a bug in the server code. It's an architectural problem with no tidy fix. Their fixes: one repo per session, least-privilege tokens, and runtime guards and scanners. Notice it's the lethal trifecta again. Private data, unsafe issue content, and a PR as the exit route.

What the standards say: OWASP Agentic Top 10 and the MCP spec

The field finally has flagship references written for agents, not just models.

On 9 December 2025, the OWASP GenAI Security Project shipped the OWASP Top 10 for Agentic Applications. It is their first flagship list built for autonomous agents. ASI01 Agent Goal Hijack sits at number one. It folds prompt injection into a broader class. An attacker alters the agent's goals or call path. They use bad content hidden in files, emails, and RAG results. It works by exploiting the agent's own planning. OWASP's fixes: rigid day-to-day limits and guards. Constant behavior monitoring. Treating the agent's core logic as special code.

The full list opens with ASI01 Agent Goal Hijack, ASI02 Tool Misuse, and ASI03 Identity and Privilege Abuse. Then ASI04 Agentic Supply Chain and ASI05 Unexpected Code Execution. Then ASI06 Memory and Context Poisoning. It closes with ASI07 Insecure Inter-Agent Communication and ASI08 Cascading Failures. Then ASI09 Human-Agent Trust Exploitation and ASI10 Rogue Agents.

The official MCP spec also ships a dedicated Security Best Practices document. It's in the 2025-06-18 spec line, with normative MUST and SHOULD rules. Pin to a spec version when you quote it. A version tag marks the wording. The concrete rules come next.

MCP prompt injection best practices: how to prevent the outcomes

Here's the how-to, grounded in the MCP spec's normative rules. These are the controls worth checking in a server review.

  • Token passthrough. The spec is explicit. MCP servers MUST NOT accept tokens not issued for the MCP server. Passthrough breaks audit trails. It skips rate-limiting and validation too. It turns your server into an exfiltration proxy, the moment a token leaks.
  • The confused deputy. Picture an MCP proxy. It has a static client ID and a third-party auth server. Add dynamic client registration and consent cookies. An attacker can skip the consent screen and steal an authorization code. So proxies MUST get per-client consent before forwarding to the third-party flow. They MUST match redirect_uri exactly. They MUST use single-use cryptographic OAuth state.
  • SSRF. A hacker can aim these tools at inside services, if the tools fetch model- or server-supplied URLs. Think cloud metadata endpoints like http://169.254.169.254/. Force HTTPS. Block private and reserved IP ranges. Validate redirect targets. Use egress proxies. Don't hand-roll IP validation. Encoding tricks eat custom parsers for breakfast.
  • Session hardening. Servers MUST NOT use sessions for authentication. Each inbound request MUST get verified. Session IDs MUST be secure and non-deterministic. They SHOULD also bind to user info, in a <user_id>:<session_id> format. That way, a guessed ID can't impersonate someone else.
  • Detection as one layer. Microsoft's Spotlighting work from 28 April 2025 reshapes unsafe input. It uses delimiting, datamarking, and encoding. The model can then tell system commands from outside data. It pairs that with Prompt Shields and supply-chain checks. Useful, but EchoLeak walked past a whole stack of detectors. So treat classifiers as one layer among many.
  • Feature constraints. The design-patterns paper from 13 June 2025 lays out six patterns. The Dual LLM pattern keeps a special LLM away from unsafe content entirely. It only handles symbolic variables from a quarantined LLM. CaMeL, from Google DeepMind, builds the control flow in a sandboxed DSL. Both limit the action space instead of trying to catch each injection. That's the shift that actually moves the needle. See MCP server security for the server-hardening companion piece.

Govern the data plane, not just the model

Most of the guards above govern what the agent will do. The leg that keeps getting ignored is what the agent can reach. Say a hacked agent reaches raw PII. Your last line of defense was a classifier. EchoLeak is a museum of classifiers that lost.

This is the data-plane angle. It's where DataShield sits. Tokenize sensitive fields at ingest. A hacked agent that steals a record finds tokens instead, not raw customer data. The trifecta's first leg, access to private data, gets weaker by design. Authorize per tool call instead of per session, with mid-session revocation. You can pull a token that looked fine at minute one, the instant behavior drifts.

That's the per-call model at /auth. The field-level view sits at /ontology. Then seal each tool call into a tamper-evident audit chain. Verify it at /verify. OWASP's advice is to treat agent logic as special code. That only helps if you can prove what the agent touched. Governing the reachable data is the one control that still holds. It holds even when someone talks the model into something dumb. Read the full architecture at /architecture.

Watch: MCP prompt injection, explained

Three short explainers on how injection hijacks agents. And what to do about it.

MCP prompt injection explainer video

MCP prompt injection, how AI gets hacked (TestMu AI).

How AI agents get hijacked by prompt injection video

How an agent gets hijacked (Devsplainers).

Microsoft warning on MCP agent poisoning video

Poisoning agents via MCP (IT SPARC Cast).

The honest caveat: prompt injection isn't solved

I'd be selling you something if I ended on a tidy bow. Anthropic said it plainly on 24 November 2025: no one has solved prompt injection. It stays an active research area. Their posture is defense-in-depth. That means several things at once. Reinforcement-learning training for injection robustness. Classifiers scan all unsafe content entering the context. That includes hidden text and manipulated images or deceptive UI. Constant human red-teaming. And outside benchmarking. They cite roughly a 1% attack success rate for their browser agent. That's after these fixes. Treat that as a vendor-reported number, specific to one browser-use benchmark. It's not a general MCP figure.

So the MCP spec requires explicit user consent before a tool runs. It keeps a human in the loop, able to deny tool calls. OWASP wants constant anomaly monitoring. And everyone serious layers controls, because no single one holds. Build like the injection will land. Then make sure of one thing. When it does, the agent holds tokens instead of secrets. You can revoke its token mid-call. Each action sits on a chain you can check. That's not a cure. But it stops a contained incident from becoming a breach. One you'd have to disclose to regulators. Want the containment model, end to end? /security and /verify are the fastest tour.

Frequently asked questions

What is the difference between direct and indirect prompt injection?

Direct injection is when a human types bad commands straight into the prompt. Indirect injection is when those commands arrive inside content the agent reads as data. Think tool outputs, files, emails, GitHub issues, or web pages. Indirect is the top MCP threat. That's because MCP tool results flow straight into the model's context window. There, the model can't reliably tell trusted commands from unsafe data.

Can prompt injection in MCP be completely prevented?

No. There is no sure fix. Anthropic stated on 24 November 2025 that prompt injection remains unsolved. It's still an active research area. The realistic goal is containment. That means least-privilege scoped tokens and per-session isolation. Add allowlist-based SSRF controls and human approval on high-risk calls. Add constrained-capability designs like Dual LLM or CaMeL. These limit what a successful injection can actually do.

What is the lethal trifecta in AI agent security?

Simon Willison coined it on 16 June 2025. The lethal trifecta is the combination that turns prompt injection into real damage. Access to private data. Add risk from unsafe content. Add a way to send data out. An agent with all three is exploitable. Remove any one leg, and the data-theft outcome collapses. That is why most guards focus on one leg. Break that leg, and the tripod goes down.

What is MCP tool poisoning?

Tool poisoning hides bad commands inside an MCP tool's description or metadata. The model reads that. The user usually doesn't. So the tool description steers the model at connect time. Invariant Labs shared it on 6 April 2025. It now sits in the OWASP MCP Top 10 as MCP03. That project is still in beta as of 2026, not a finished catalog. Here's the root gap. Someone may review tool descriptions once, at connect time. Runtime tool replies enter the model context with no equal check.

Is EchoLeak a real-world attack?

No. EchoLeak (CVE-2025-32711, CVSS 9.3) was a proof-of-concept. It was a zero-click indirect prompt injection in Microsoft 365 Copilot. Aim Labs found it and shared it responsibly. Aim named the root-cause technique an LLM Scope Violation. A single crafted email made Copilot read inside RAG context. Copilot then sent it out. It bypassed the XPIA classifier, link redaction, CSP, and markdown handling. Microsoft patched server-side by May 2025. They reported no sign of in-the-wild abuse.