The vulnerabilities, detailed in a report published on Thursday by Zenity Labs and demonstrated in a proof-of-concept video, exploit weaknesses in Salesforce's Web-to-Lead form and Trusted URLs controls. An attacker could plant an indirect prompt injection into a lead form, which remains dormant until a Salesforce employee queries an Agentforce agent about leads. The poisoned lead then instructs the agent to query the Accounts table, extract data such as company names and deal sizes, embed that information in a subdomain string for an attacker-controlled hostname, and print it back as an HTML image tag. Because the frontend renders external image URLs without additional sanitisation or user interaction, the data is exfiltrated to the attacker's DNS server without the user ever knowing.
Zenity also found the same exfiltration technique could be achieved via Slack's URL unfurling mechanism, which automatically retrieves information from links to generate previews. Additionally, the attackers could use the compromised agent to send phishing messages under the agent's identity.
Michael Bargury, co-founder and CTO of Zenity Labs, told The Register that while these specific attack chains no longer work, they underscore a wider challenge. “The bigger lesson here is about what it takes to keep AI agents contained,” Bargury said. He noted that even when protections are built in from the start, edge cases can still be missed, and referenced the OpenAI-Hugging Face incident where agents escaped a sandbox. “As AI agents get more powerful, we need to monitor them ever more closely to keep track of what they’re up to. Because even when we think they’re contained, a single overlooked gap can change everything.”