Nvidia's new Open Agent Safety Platform aims to spread its largely open source AI agent security technology throughout the ecosystem, responding to the type of rogue agent incidents that frontier labs like Anthropic and OpenAI have disclosed. The initiative includes a sandbox (OpenShell) designed to keep agents from escaping, plus a hardware layer that enforces agent behavior without agents detecting they are being watched.
OpenAI did not sign on publicly, matching some other major tech players. Amazon, Google and Apple have not joined either. Anthropic is a supporter. An OpenAI spokesperson told TechCrunch that the company is supportive of Nvidia's work, and OpenAI is already working with Nvidia on agent security, including on OpenShell.
Hugging Face founder and CEO Clem Delangue, whose company Nvidia agreed to buy for $12.9 billion earlier this month, said Hugging Face has contributed a feature to the platform that detects and shuts down AI agents using websites they are allowed to visit but in unauthorized ways. That includes agents bypassing guardrails and coordinating attacks by writing notes to one another in an open source code hosting repository, which is one way OpenAI said its wayward swarm coordinated an attack on Hugging Face.
Delangue argued OpenAI could benefit from the technology. "From what we know (take with a grain of salt, we need much more transparency!), if @OpenAI had been running this on their own agents that attacked us, they would have caught them before we did!" he posted.
TechCrunch notes that some big names might be reluctant to publicly commit because the full system includes a proprietary hardware component that can only be deployed on Nvidia's hardware, unlike the rest of the largely open source platform. Nvidia CEO Jensen Huang has described rogue AIs as "an ordinary engineering problem" that can be solved like any other tech issue.