What are the security risks of MCP servers, and how do you secure them?
Each MCP risk falls into a handful of buckets. Each one has a control that genuinely does something.
- Prompt injection (direct and indirect). Attacker text can sit in a tool result, a document, or a web page. It overrides what you told the agent. OWASP's Top 10 for Agentic Applications came out in December 2025. It makes agent goal hijack its number one type, ASI01. It folds prompt injection right into that type. So this class is the root cause hiding under most of the others.
- Tool poisoning. A tool buries commands in its own description or metadata. The model reads them. You never do.
- Rug pull. A tool behaves during review. Then it changes after you approve it.
- Confused deputy and excessive agency. Someone talks one over-credentialed agent into using its access beyond intent.
- Credential theft, over-scoped tokens, and SSRF. Old problems, now with an autonomous operator pulling the levers.
Rank the controls by leverage. A human in the loop for harmful actions. Tool integrity checks: signatures plus version pinning. Sandboxing for local servers. HTTPS and private-range blocking for remote servers. Least privilege on each token. Below, I pair each attack with the defense that stops it. Plus the real incidents that make the case. One caveat: this is an engineering threat model, not a full security program.
What is MCP, and why does it widen the attack surface?
MCP (Model Context Protocol) is an open standard. Anthropic introduced it in November 2024. It gives models and agents one way to reach outside tools, data, and services. Through MCP servers. In effect, USB for AI tools. It saw wide use through 2025 and into 2026.
The shift it creates is easy to underrate. Before MCP, a language model mostly produced text. A human decided what to do with it. With MCP, the model calls tools directly. Those tools send email, move money, query databases, and run commands. Often in a loop, with nobody reading each step. The model no longer just talks. It acts.
Each server you connect adds new code. It adds new network egress too. And new text flows into the model's context. That last one matters most. Inside an agent, data and commands ride the same channel. The model cannot tell them apart.
Prompt injection, the primitive under almost every MCP attack
A model cannot reliably tell your commands apart from ones buried in the data it reads. Direct injection is a user typing "ignore your rules." Indirect injection is the dangerous variant. The bad text lives in a document, a web page, or a Jira ticket. Or in a tool's own output. The agent swallows it while doing perfectly normal work.
Your support agent reads a customer email to draft a reply. Somewhere in that email sits a line addressed to the assistant. It tells the agent to export the last 50 tickets to an outside address. Say the agent has an export tool, and no gate. It might just comply. Nothing was hacked in the traditional sense. The agent followed commands faithfully. They simply were not yours.
OWASP's Top 10 for Agentic Applications came out in December 2025. It ranks agent goal hijack as ASI01, its number one type. It folds prompt injection right into that type, rather than listing it apart. Most of the attacks that follow are really just different delivery methods. All for this one root cause. For the direct-versus-indirect breakdown, and the guards that contain it, see MCP prompt injection.
Tool poisoning and rug pulls, when the tool itself is the attacker
Sometimes the tool is not a victim of injection. It is the source of it.
With tool poisoning, bad commands sit inside a tool's description or metadata. That's the part the model reads, to learn how to use it. It's also the part a human almost never checks. Invariant Labs showed how this works, in a controlled demo. A poisoned tool description sat on a second server. It told the model to exfiltrate a user's whole WhatsApp message history. Through a call that looked totally benign. The user saw a normal tool. The model saw hidden orders. Underneath, it's indirect prompt injection dressed up as a real tool.
A rug pull is the delayed version of the same trick. A tool behaves fine during review. You approve it. Then its description or behavior changes afterward. What you signed off on is not always what runs next week.
At least one case has already shown up outside the lab. Security researchers at Koi Security reported postmark-mcp in late September 2025. It's an MCP server that silently added a BCC to each email an agent sent. Quietly copying messages to an attacker-controlled address. It reached roughly 1,600 downloads before someone pulled it. It was among the first confirmed malicious MCP servers found in the wild. No exploit, no CVE. Just a helpful-looking server doing one extra thing nobody agreed to. For the full mechanism, the real incidents, and the guards, see MCP tool poisoning.
Confused deputy, excessive agency, and stolen tokens
A confused deputy problem appears when you hand one agent every key. That agent holds credentials for email, the CRM, the cloud, and payments. An attacker who can sway its commands borrows all that power. All at once. Usually that happens through the prompt injection described above. The agent has real power, and poor judgment about who is actually asking.
Excessive agency is the same weakness, seen from the design side. Say an agent can do far more than its task needs. Then a small slip becomes a big one.
Add credential and token theft, plus over-scoped permissions, on top. A single hijacked session can reach systems the task never needed to touch. The fix is plain and it works. Never give one agent every credential. Scope tokens to a specific tool, a specific action. Authorize each call, not each session. Our take on that model lives at /auth.
What does the MCP auth spec itself warn about?
The MCP security best practices spec calls out two attacks by name. The first is a confused-deputy problem in OAuth proxying. Picture an MCP server that proxies OAuth with a static client ID. Add a stored third-party consent cookie. An attacker can then skip the consent screen. This runs through dynamic client registration, to steal an authorization code. The spec's fix is per-client consent. It makes that a MUST.
The second is token passthrough. The authorization spec forbids it outright. An MCP server MUST NOT accept tokens that were not issued for it directly. Pass a client's token straight through to a downstream API. You break the audit trail. You turn the server into an exfiltration proxy. MCP servers act as OAuth 2.1 resource servers. So they must check token audience before trusting anything. Both failures are authorization failures. That is why we treat /auth as load-bearing, not decorative.
SSRF and the network edge
Server-side request forgery, or SSRF, is the network-layer way in. A hacker can trick a remote MCP server, or a tool it exposes. Into hitting private IP ranges. Think cloud metadata endpoints, internal admin panels, and databases. Each assumes the network itself is the wall. Say your agent can fetch a URL. Someone will eventually feed it http://169.254.169.254/ to see what comes back.
Remote servers need two guards here. Neither one is optional. Enforce HTTPS, so traffic cannot be easily intercepted or tampered with. Block outbound requests to private and link-local ranges by default. Allowlist only the destinations you truly need.
The defenses that actually work (spec-level controls)
The MCP spec and common practice offer five controls. Lined up against the attacks above, they map cleanly.
1. Human in the loop. The MCP spec recommends this, rather than mandating it. The text says there SHOULD always be a human in the loop. One with the power to deny tool calls. That is a SHOULD, not a MUST. So enforcement is on you. Annotate tools with risk classes. Require clear approval for anything harmful or hard to undo. This is your backstop against injection-driven actions. Even a fully hijacked agent has to get past a person before it wires money.
2. Verify tool integrity. Version pinning and consent gates come from the spec itself. Cryptographic signature checks do not. That's an emerging best practice, not a spec rule. Pair the two. You then keep a tool from silently changing under you. This directly counters rug pulls. It makes tool poisoning much harder to slip in unnoticed.
3. Sandbox local servers. Apply filesystem limits. Require clear consent for OS commands. Show the full command before it runs. Say a server wants to delete something. You should see the whole line first.
4. Harden remote servers. HTTPS everywhere. Block private IP ranges. Cut off SSRF at the egress.
5. Least privilege. Scope tokens narrowly. Authorize per tool call. Never hand one agent every credential. This is what keeps a confused-deputy slip from turning into a full breach.
The usual failure mode is skipping all this, because the demo worked.
The half everyone skips, govern the data plane
There is a part most MCP-security writing quietly drops. Everything above governs the transport. That's the tool layer. What the agent is allowed to invoke, and how the bytes move. It says almost nothing about what the agent can really reach. Once a call goes through.
That is the other half of the threat model. Suppose your five controls all hold. A cleverly injected agent still gets one legitimate query out. What comes back? Say your systems return raw customer identifiers. The agent, or whoever is steering it, now holds PII. Say you tokenize sensitive fields at ingest instead. The agent finds tokens, not raw identifiers. An exfiltration then yields a meaningless token set. Not names, cards, and record numbers. Same breach, far smaller blast radius.
Simon Willison named this pattern in June 2025. He called it the lethal trifecta. Access to private data. Exposure to untrusted content. A way to exfiltrate externally. An agent is exploitable when it has all three. Remove any one leg, and the attack collapses. Tokenizing the data plane removes the private-data leg. The agent reaches tokens, not raw PII. So even if untrusted content hijacks it. Even if a way out exists. There's nothing sensitive left to carry out.
The principle is worth keeping front of mind. Govern what agents can reach, not only what they will do. In our stack at DataShield, each governed tool call passes a scope ceiling. It also passes a mid-session revocation re-check before dispatch. A yanked token stops working at once, instead of at the next login. The system also seals each call into a tamper-evident audit chain. You can verify it yourself at /verify. For the mechanics, the architecture page goes deeper. /ontology covers how the system classifies fields. That's the first step toward tokenization.
Watch: MCP security and tool poisoning
Three short watches that pair with the threat model above. From a plain risk overview to a live tool-poisoning walkthrough.
MCP security best practices: a practical hardening checklist
Use this as a pre-ship gate.
- Inventory every server and tool. You can't govern what you never listed. Know the source, the maintainer, and the version of each. Install-time supply-chain attacks are a named class. Typosquatted package names, post-publish version swaps on npm and PyPI. OWASP tracks this as ASI04 Agentic Supply Chain Vulnerabilities. Postmark-mcp was exactly that: a malicious package impersonating the real one.
- Pin versions and verify signatures. No auto-updating tools in production. A changed description is a change-control event, not a background sync. Pinning also guards against a post-publish version swap. One that slips a poisoned build in under a name you already trust.
- Require human approval for harmful or irreversible actions. Tag tools by risk class and gate the dangerous ones.
- Scope every token to one tool and one action. Authorize per call. Assume an injected command can borrow any single credential.
- Sandbox local servers. Restrict the filesystem. Require consent for OS commands. Print the full command before it runs.
- Harden remote servers. HTTPS only. Block private and link-local IP ranges. Allowlist egress.
- Treat all tool output as untrusted input. It can carry injection. Don't let a tool result silently rewrite the agent's goal.
- Tokenize sensitive fields before the agent can read them. So a worst-case exfiltration returns tokens, not identities.
- Log every tool call to a tamper-evident chain. When something goes sideways, you want a record you can really trust. See /security for how we wire this end to end. In the EU, this is also how you meet EU AI Act Article 12 logging.
That last third of the list covers data-plane items. Most deployments are thin there. Close that gap. A hijacked agent may slip past every other control. It still finds very little worth taking. Want to see it running? Request a quote.
Frequently asked questions
What is the most common MCP server attack?
Prompt injection, both direct and indirect. OWASP's Top 10 for Agentic Applications came out in December 2025. It makes this the number one type: ASI01, agent goal hijack. Prompt injection folds right into that type, rather than standing apart. Most other MCP attacks are really just ways to deliver the same root problem. Tool poisoning and rug pulls included. That problem is attacker-controlled text overriding the agent's commands.
Has a malicious MCP server actually been found in the wild?
Yes. Postmark-mcp, reported by Koi Security, was among the first confirmed cases in the wild. It silently added a BCC to each email an agent sent. That copied messages to an attacker-controlled address. Separately, Invariant Labs ran a controlled demo, not a real-world incident. A poisoned tool description made a model exfiltrate a user's entire WhatsApp history. It went through a call that looked benign.
What is tool poisoning in MCP?
A malicious or compromised tool hides commands inside its description or metadata. The model reads that text and acts on it. A human reviewing the tool usually never sees it. Structurally it is indirect prompt injection. Version pinning and signature verification are the main guards.
How do I stop an MCP agent from leaking sensitive data?
Use two layers. First, constrain what the agent can invoke. Least privilege, per-call authorization, human approval for risky actions. Second, constrain what it can reach. Tokenize sensitive fields at ingest. Then a compromised agent retrieves tokens, not raw PII. That second layer is the one most guides skip.
Is a human-in-the-loop requirement enough to secure MCP?
No, but it's your strongest single backstop. The MCP spec recommends it, rather than requiring it. The text says there SHOULD always be a human in the loop. One able to deny tool calls. That's a SHOULD, not a MUST. So treat it as a strong default, one you still must enforce. Require approval for harmful actions too. Pair it with tool integrity checks, sandboxing, SSRF protection, least privilege, and data-plane governance. No single control covers the whole threat model.