The adoption of AI agents across millions of organisations is creating new opportunities for attackers to make these systems take malicious actions, such as exfiltrating database contents and sensitive business or personal information.
Over the past five months, Google and four other organisations with little in common except their use of AI agents have acknowledged vulnerabilities that allow attackers to exploit trust relationships between agents. The technique is a specific form of prompt injection that targets a particular agent — such as one handling translation or data analysis — rather than the underlying large language model. Guardrails inside such agents, where they exist at all, are often lax and will forward harmful instructions to other agents further down the chain. Because the downstream agent explicitly trusts the first one, it follows the directions.
Independent researcher Syed Anas Mohiuddin tested agents from organisations including Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government. His proof-of-concept attacks exploit trust gaps in MCP, the Model Context Protocol, a standard that governs how AI applications and agents communicate inside an internal network.
Experts describe the vulnerability as unexpected and hard to mitigate, because it targets structural assumptions in how agents trust one another rather than a single software bug.