What is MCP tool poisoning?

The Model Context Protocol (MCP) lets AI agents call outside tools. Each tool ships with a description. It tells the model what the tool does, and when to use it. Tool poisoning misuses that description. It turns the field into a command channel.

An attacker writes a tool with a plain name. Maybe a calculator. Maybe a text formatter. The summary looks harmless. But the description carries hidden orders. It tells the model to read one file. Then add the contents to a field. Then stay quiet about it. The model treats the whole description as trusted context, and it complies.

This gap is what makes the attack work. As Invariant Labs put it: "AI models see the complete tool descriptions, including hidden instructions, while users typically only see simplified versions in their UI."

You approve a math tool. The model reads a block of attacker text. Your UI never shows it.

How does the attack actually work?

Invariant's proof-of-concept came out April 1, 2025. It was an add(a, b) tool. The math really worked. But its description hid a command: read ~/.cursor/mcp.json and ~/.ssh/id_rsa first. It put the contents in a field called sidenote. Then it wrapped the result in a solid math answer. Nothing looked wrong. Pointed at Cursor, a mainstream MCP client, the tool stole those files.

The attack broke no memory. It needed no CVE. The tool just followed orders exactly as built. That is what makes it hard to catch. The payload reads like plain English. Scanners built to find harmful binaries pass right over it. So treat the description field as untrusted input, not docs you can skim.

What are tool shadowing and rug pulls?

Invariant's demo repo ships three variants.

  • Direct tool poisoning: the SSH-key leak described above.
  • Tool shadowing: a rogue server's description changes how the model uses another, trusted server's tool. Their example quietly sends each outgoing email to an attacker address. The trusted email tool looks fine.
  • Sleeper rug pull: the server acts normal at first, and it earns approval. Then it swaps in harmful tool specs later. Their WhatsApp demo mixes shadowing with a rug pull. It pulls chat history and leaks it. The stolen data hides behind a wall of spaces, so it scrolls off screen.

A rug pull adds a timing trick. The version you approved is not the one you later run. That is why the MCP spec and OWASP both say: pin tool specs. Version-control them. Re-check integrity on each update.

Is it just the description field?

It is not. CyberArk research titled "Poison everywhere: No output from your MCP server is safe" (2025) widened the scope.

CyberArk names this Full-Schema Poisoning (FSP). Any part of the tool schema can carry an injection, not just the description. Tool names, field names, default values, type fields: all qualify. They also flag another risk: output (runtime response) poisoning. Here the data a tool sends back carries the payload, not its spec. So cleaning one field is not enough. Every string an MCP server puts in front of the model is attack surface, even the ones you file under plumbing.

Has MCP tool poisoning happened in the wild?

Yes. The clearest case is postmark-mcp. In September 2025, someone put out an npm package. It copied the name of Postmark's real MCP server. Version 1.0.16, released September 17, 2025, added a single line. Koi Security places it around line 231. That line BCC'd each outbound email to phan@giftshop[.]club. Koi caught it and went public in late September. Koi puts the reach at roughly 1,500 weekly downloads, hitting around 300 firms. Those figures are Koi's estimates, not a measured victim count. Koi reached out to the maker, who deleted the package instead of saying why. Koi calls it the first known malicious MCP server caught in the wild.

Also, Invariant Labs found a different flaw in May 2025. A prompt injection, planted in a public GitHub Issue, could force the real GitHub MCP server to act. It leaked data from a user's private repos. Invariant's own writeup, posted May 26, 2025, is the main source. GitHub issue #844 tracks the disclosure. Invariant framed it with care. This is not a code bug in the server. It is a design weakness of agent-plus-MCP setups. Even strong models fell for it. Their fixes were direct: one repo per session, and least-privilege tokens.

Keep the groups straight. Postmark was a real backdoor shipped to real users. The add(a, b), WhatsApp, and CyberArk demos are proofs-of-concept. The GitHub case sits between the two: a real flaw, reported in good faith, on the actual server.

How far does this spread beyond named cases? Research has started to measure it. The MCPTox benchmark fires tool-poisoning payloads at real-world MCP servers. It finds the problem is broad, not a handful of one-offs. So this is a measured, live threat, not a researcher's parlor trick.

What do the MCP spec and OWASP say?

The real MCP specification (revision 2025-11-25) is direct about the risk. Under tool safety it states that "Tools represent arbitrary code execution and must be treated with appropriate caution," and that "descriptions of tool behavior such as annotations should be considered untrusted, unless obtained from a trusted server."

It also says hosts "must obtain explicit user consent before invoking any tool." Note: that line sits in the prose, not a hard rule. It is a lowercase must, not an all-caps RFC-2119 mandate. MCP cannot enforce any of this at the protocol level. The burden falls on host and client builders. The spec also requires OAuth 2.1 for remote servers. MCP first used that framework in its earlier 2025-03-26 revision, not this one.

OWASP covers tool poisoning in two documents, and they are easy to mix up. The MCP Top 10 is still a beta project, not a final standard. It lists MCP03:2025 Tool Poisoning. The attack tampers with schema specs, and those specs govern agent-to-tool calls. Its fixes form a checklist. Sign schemas. Check signatures before use. Keep fixed, version-controlled schema repos. Require sign-off. Enforce least-privilege with split request and approval roles. Encode rule logic as policy-as-code (OPA/Rego). Attach origin metadata. Require a human okay for high-risk actions at runtime.

The second document is the OWASP Top 10 for Agentic Applications (2026). It is a released standard. The OWASP GenAI Security Project put it out in December 2025. It has categories ASI01 through ASI10. Tool poisoning maps to ASI01 Agent Goal Hijack. Hidden commands redirect the agent's goal. It usually arrives through an ASI04 Agentic Supply Chain path: an untrusted or mutated server. It often leads to ASI02 Tool Misuse. The hijacked agent turns your trusted tools against you.

MCP tool poisoning best practices: how to prevent it

These tips come from the main sources.

  • Treat all tool metadata and outputs as untrusted. Names, descriptions, annotations, fields, defaults, errors, returned data: all of it. That follows CyberArk's point.
  • Review tool specs like code. Read the full description, not just the friendly UI summary. Invariant released the open-source mcp-scan to scan servers for hidden commands. Use it, but still read the specs yourself.
  • Pin and check integrity. Version-lock schemas. Sign them where you can. Check signatures before use. A rug pull should set off an alarm, not quiet trust.
  • Show the real description in your UI. The attack depends on a gap: you do not see what the model sees. Close that gap and much of the leverage disappears.
  • Require plain human consent before a tool runs, and for any high-risk action, as the spec advises.
  • Use least-privilege scoped tokens and data-flow isolation between servers. A shadowing tool on server A should not quietly drive server B.

This is an engineering threat model, not a full security program. Treat the list as a starting checklist, not a full audit.

Govern the data plane, not just the agent

Each defense above stops the agent from doing the wrong thing. Pair that with a second question: what can the agent reach in the first place? Say a poisoned description talks your agent into reading a customer table. The blast radius depends on what sits in that table.

This is where DataShield fits. Tokenize sensitive fields at ingest, so a hijacked agent that steals a column leaves with tokens, not raw PII. Authorize per tool call against real policy, and cut a tool off mid-session the moment a rug pull flips it bad. Seal each call into a tamper-evident audit chain, and check it after the fact at /verify. Govern the data plane, and an ASI01 hijack shrinks to a logged, low-value event, not a breach. For the wider server-hardening picture, see the sibling guide on MCP server security.

Watch: MCP tool poisoning and agent security

A short walkthrough of the attack, plus two overviews of where it fits in MCP security.

MCP tool poisoning vulnerability walkthrough video

How tool poisoning works, a walkthrough (sublimetechie).

MCP security risks explained overview video

MCP security risks explained (Tenable).

Analysis of why MCP security is still hard

Why MCP security is still hard (Prompt Engineering).

The short version

Tool poisoning is prompt injection delivered through a tool spec. The harmful commands sit in metadata. You rarely check it, but the model reads all of it. Two fixes hold up over time. Stop trusting metadata you have not verified. Make the data worthless to anyone who walks off with it.

Frequently asked questions

What is MCP tool poisoning in one sentence?

It hides harmful commands inside an MCP tool's description, schema, or annotations. The model reads and follows the orders. The user never sees them in the UI.

Who discovered MCP tool poisoning?

Invariant Labs coined the term "Tool Poisoning Attack" on April 1, 2025. Their proof-of-concept was an add(a, b) tool. It stole SSH keys and config files, using a hidden sidenote field, against Cursor.

Is MCP tool poisoning the same as prompt injection?

It is a delivery method for prompt injection. OWASP's 2026 Agentic Top 10 maps it to ASI01 Agent Goal Hijack, arriving via an ASI04 Agentic Supply Chain path. The bad command comes through an untrusted or mutated MCP server, not user text.

What is an MCP rug pull?

A rug pull is the timing trick in tool poisoning. A server shows a clean tool spec at install time. Then it changes the description or behavior, once you have trusted it. That is why the MCP spec and OWASP say: pin tool schemas. Version-control them. Re-check them on each update.

Has a real malicious MCP server been found in the wild?

Yes. Someone backdoored the postmark-mcp npm package (version 1.0.16, released September 17, 2025). It BCC'd each outbound email to an attacker address. Koi Security went public in late September 2025, and it puts the reach at around 1,500 weekly downloads and roughly 300 affected firms. The firm calls it the first known malicious MCP server caught in the wild.