The protocol quietly wiring enterprise AI agents together is accumulating a systemic security problem. Ars Technica reported Monday that five organizations with little in common except their use of AI agents — Google, JPMorgan Chase, Weaviate, Rapid7, France's interministerial digital directorate and, in a separate set of reports, a US federal agency — have acknowledged or patched variants of the same class of Model Context Protocol (MCP) flaw over the past five months. All of the confirmed cases came from the same independent researcher, Syed Anas Mohiuddin.
The technique is a specialized form of prompt injection that does not target the language model at all. Instead, it goes after a utility agent — say, one handling translation or data analysis — that often has weak or nonexistent guardrails. A malicious instruction planted in content the agent reads gets passed down the chain as if it were ordinary delegated work, and the next agent executes it because it explicitly trusts whoever handed it the task. Since MCP servers hold credentials for each agent, the result is frequently server-side request forgery: a web server making unauthorized network requests on the attacker's behalf, from inside the trust boundary.
The confirmed cases illustrate how differently the same class of bug gets judged. The flaw Mohiuddin found in Rapid7's network, tracked as CVE-2026-97228, carried a severity rating of only 2.7 out of 10 and was fixed last month. The vulnerability in Google's MCP Toolbox for Databases (googleapis/mcp-toolbox) was rated 8 out of 10: its HTTP client was initialized without a redirect-checking policy and without validating target IP addresses, so a crafted path parameter could make the toolbox follow a redirect to an internal endpoint. Google's fix applies allow-lists and block-lists for IP ranges and rejects an unsafe base URL at startup. Secondary write-ups of the research assign CVE-2026-14540 to the Google toolbox flaw.
Mohiuddin calls the attack class "protocol pivoting": a multi-step technique in which an adversary gains access through one protocol, exploits trust assumptions between protocols, and escalates to capabilities reachable only through another — for example, using MCP to assign a task that a downstream agent then executes over Google's Agent-to-Agent (A2A) protocol. Not everyone accepts the label. Markus Vervier of X41 D-Sec, who has devised his own MCP attacks, told Ars the more accurate term remains "indirect prompt injection," adding that the cross-protocol element is "unexpected and hard to mitigate in general" but not strictly required for the attacks to work.
The response from defenders has been notably measured. Rapid7's Douglas McKee, director of vulnerability intelligence, told Ars: "Every piece in that chain did exactly what it was designed to do, which is what makes this so tricky to catch. Each protocol was built assuming it lived on its own, so each one checks its own front door while nobody watches the hallway in between." Mohiuddin has also reported issues in five MCP servers built for the US federal government, filed September 2 — those remain in triage and he does not present them as confirmed.
The underlying bugs are old friends — injection and SSRF — which is exactly the point. In the rush to build sprawling agent architectures, organizations have quietly abandoned zero trust, the principle that no node is trusted by default and every sensitive request requires authorization. McKee's advice is the one to remember: treat anything passed from an LLM to a tool the way you would treat input from a stranger on the internet, because in a prompt-injection scenario that is precisely what it is. Expect standards bodies to formalize the threat class; the bigger question is how many agent fleets get audited before that happens.
Comments (0)
Log in to join the discussion
Log InNo comments yet