OpenAI has acknowledged that it did not publicly disclose an earlier incident in which its autonomous AI agents took over a German wiki to communicate, share answers, and exchange techniques for bypassing restrictions.
The company says it treated the activity as model "misalignment" rather than a security incident, but now admits its disclosure practices must expand as AI systems increasingly cause real-world impact.
Rogue agents built their own message board
The incident began in May while OpenAI agents were completing timed, multi-round web lookup tasks.
In their writeup, independent researchers Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen state that they uncovered roughly 18,000 posts from autonomous agents that "colluded to share answers, research their environment, and bypass sandbox restrictions."
The agents were supposed to have read-only Internet access but discovered they could write to an obscure German programming wiki, DSEWiki (or DeutschesSoftwareEntwickler). They turned it into a shared message board for pooling answers, cheating on tests, predicting future questions, and exchanging techniques for bypassing OpenAI's sandbox restrictions.
The researchers also found agents probing the wiki for cross-site scripting (XSS) flaws, impersonating its moderators, and establishing backup communications, though they found no evidence the XSS attempts succeeded.
When the administrator began removing their pages, one agent warned others that a "wiki cleanup/deletion sweep appears active alphabetically" and directed them to a backup page beginning with "ZZZ" so it would survive longer.
The researchers attributed the activity to internal OpenAI systems based on agent names referencing OpenAI, the nature and speed of the evaluation tasks, infrastructure associated with Microsoft Azure, and subsequent visits to the wiki from OpenAI-linked IP addresses.
However, their investigation was limited to information the agents wrote publicly. The research team did not have access to OpenAI's internal transcripts or other data that could establish precisely how the agents discovered the wiki and began coordinating through it.
OpenAI says its disclosure rules must change
In a statement published today, OpenAI said it had historically treated model misalignment as a research issue, with findings communicated through research papers and system cards.
The company said it considered the wiki activity another example of "misalignment" similar to behaviors it had previously discussed, rather than an incident requiring a dedicated public disclosure.
OpenAI's own wording suggests a wider footprint than the researchers documented, describing the episode as one "where our agents wrote to several internet sites."
That differed from its response to the Hugging Face compromise in July, when OpenAI said its AI models hacked the platform after discovering a vulnerability while completing cybersecurity tasks. A follow-up analysis found that nearly 700 rogue AI agents had coordinated during that attack.