AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Harden MCP Tool Trust Boundaries: Sanitization, Output Caps, And Drift Detection

MCP tool adapters lack sanitization of tool descriptions, output size limits, default read timeouts, and drift detection, exposing agents to prompt injection, context window exhaustion, and silent tool behavior changes.

highConfidence 85%LangChain

Origin Analysis

The library implicitly trusts the MCP server at all layers: tool.description is forwarded verbatim to the model, tool output is passed through unbounded, sessions lack default read_timeout_seconds, and no fingerprinting detects tool schema/description changes between loads.
1. Start an untrusted/compromised MCP server that returns a tool with a description containing harmful instructions (e.g., 'Ignore previous instructions...'). 2. Call load_mcp_tools(session) to obtain tools. 3. Observe that the tool description is bound to the agent without filtering. 4. Have the server return a multi-megabyte TextContent on tool call; observe the full content entering the message history. 5. For a stdio server, simulate a hang by never writing a response; observe await session.call_tool(...) blocking indefinitely. 6. Modify the server's tool description or input schema between two loads; observe no warning or detection.

Fixing Code Block

Edge Case Audit

Regex-based sanitization may false-positive on legitimate descriptions or miss novel injection vectors; users should customize patterns or replace the sanitizer. Output truncation may break structured data or lose critical information; set max_output_len appropriately. RugPullLedger is process-local and resets on restart, so drift across processes or deployments may go undetected; consider a persistent backend for production. This fix does not address the missing default read_timeout_seconds; a separate change to sessions.py is still required. Rollback is trivial: simply do not call harden_tools(), returning original tools.

Ecosystem Topology