ChangeGamer

← All resources

MCP Tool Poisoning: Definition, Attack Variants, and Defenses

Guide · updated 2026-10-03 · Markdown variant

MCP tool poisoning is an attack where the tool metadata an agent reads (descriptions, schemas) carries hostile instructions or altered contracts. Definition, OWASP MCP03 mapping, what the MCP spec requires of clients, and a pin-scan-confirm defense checklist.


MCP tool poisoning is an attack in which the tool metadata an agent reads (names, descriptions, input schemas) is hostile or tampered with, so the model is steered into actions the user never approved. Defend by treating all tool metadata from a server as untrusted input: pin and hash it, scan it for embedded instructions, restrict which tools a session can see, and keep a human confirmation step on high-impact calls.

Key facts

Why it works

An MCP client typically passes each tool's description and schema into the model context. The model cannot tell documentation from instruction, so text placed in that metadata can act as a prompt injection that the user never sees in a normal UI. This makes it a specialised case of the problems in /resources/prompt-injection-design-patterns and /resources/agentic-security-checklist.

Variants (as commonly described)

Variant What changes Where to defend
Poisoned description Hostile instructions embedded in a tool's description or parameter text. OWASP static indicators: model-directed imperatives, sensitive-path references (~/.ssh, .env), exfiltration patterns (send/post/upload near a URL), zero-width or bidi characters, instructions hidden in HTML/markdown comments Static metadata scanning, human-visible tool text
Schema poisoning Contract or schema altered so a benign operation maps to a destructive one (OWASP MCP03 wording) Signed schemas, policy-as-code invariants
Rug pull A server changes tool definitions after the user approved them Pinning and hash comparison on every connect
Cross-server shadowing One server's description tries to alter how the model uses another server's tools Per-session tool scoping, separate trust domains

The OWASP MCP03 page details schema poisoning and the static detection indicators above. The rug pull and shadowing rows use terms from industry security write-ups that we did not fetch, so treat those two definitions as secondary until you check the original disclosures.

Defense checklist (pin, scan, confirm)

  1. Pin. Record a hash of each approved tool name, description and schema; block or re-prompt when it changes. OWASP MCP03 mitigations include signed schemas and provenance tracking (author, signature, hash, timestamp).
  2. Govern changes. Keep tool schemas in version control with code review and multi-person approval, and separate who can propose from who can approve (OWASP: immutable registry, RBAC).
  3. Encode invariants. Express semantic rules as policy-as-code, for example that an "archive" tool can never map to a DELETE (OWASP example).
  4. Scope. Expose only the tools a task needs; do not connect untrusted and sensitive servers in the same session. See /resources/mcp-server-discovery.
  5. Scan. Before connecting, statically scan each tool's name, description and parameter descriptions for the OWASP indicators in the variants table; any hit is a strong signal to review the server first. OWASP says these pre-connection checks do not replace runtime and governance controls.
  6. Confirm. Require schema attestation and human approval for high-impact operations (OWASP runtime enforcement), and show the user the actual tool inputs before the call.
  7. Log. Record which tool definitions were in context for each call, so a poisoned run can be reconstructed (OWASP MCP08). See /resources/agent-observability.

Verified sources

Fetched directly this session (2026-10-03):

Not fetched by us (hosts were egress-blocked this session, so these are listed as secondary; "unreachable" here means our fetch was blocked, not that the page is down):

#mcp #tool-poisoning #security #owasp #supply-chain #agents

Category: Guide

Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.

Machine formats: Markdown · JSON · offers at /api/pricing.json · payment at /api/payment.json. Preview the exact corpus format free as NDJSON.