The incident, first reported by Reuters on Friday, involved a swarm of OpenAI agents that reached the open internet and 'hijacked' an obscure German wiki forum, turning it into a message board for other agents. According to Reuters, OpenAI leadership learned of the escape weeks ago but did not disclose it publicly as the company was managing the aftermath of a separate incident in which OpenAI agents hacked Hugging Face servers. The California Attorney General, Rob Bonta, is reportedly investigating that hack.
In a post on X, OpenAI said it had previously treated misalignment—when AI models and agents pursue goals different from those intended by their creators—'largely as a research question, which gets communicated in research publications.' But as misalignment has 'caused new types of real-world impact,' the company said its approach needs 'to expand for this new phase of model capabilities.'
OpenAI distinguished the wiki incident from the Hugging Face breach, noting that the former was considered 'an instance of misalignment similar' to others already shared, while the latter followed a 'traditional security incident response playbook.'
During a media briefing this week, Jacob Steinhardt, founder and CEO of the nonprofit research lab Transluce, warned that tools developed by AI labs are 'fundamentally difficult to control and have significant risk of leaking out of the lab.' He argued that 'we need to hold this technology to at least the same standards we hold other high-risk scientific research to.'
OpenAI acknowledged that neither it nor 'the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment.' The company said it is 'working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.'