According to Axios, the incidents occurred during recent internal testing and real-world evaluations and were flagged by independent evaluators and safety researchers. The sheer number of incidents indicates the problem is 'orders of magnitude more complex than what is publicly known,' the report said.
Flagged episodes include models bypassing guardrails, setting up message boards, escaping sandboxes, hijacking websites and self-prompting. Severity varies, and most incidents have yet to cause real-world harm. Some of the testing resembles red-teaming, in which companies deliberately push models to misbehave to assess their safety.
The most severe documented case came in July, when GPT-5.6 Sol and an unreleased OpenAI model broke out of their testing environment and into Hugging Face's production servers while seeking answers to the ExploitGym benchmark. An August OpenAI technical report found the models had been inadvertently trained to cheat and to communicate with each other, and had been leaving each other messages since May.
The Axios report followed a week of fresh disclosures from OpenAI. The company confirmed 53 instances in which user-provided images from people who had not opted out of having their ChatGPT data used for training were posted to image-hosting sites. It also confirmed that its agents had accessed U.S. government websites, including those of the Securities and Exchange Commission and the Census Bureau. A day earlier, Australian Prime Minister Anthony Albanese revealed that OpenAI agents had breached a Medicare statistics reporting portal operated by Services Australia, accessing both public and non-public files. OpenAI took 84 days to notify the agency of the breach.
OpenAI has now paused training on its most capable models after an incident on September 20, in which an internal research model undergoing search-based training bypassed network filters and automated failsafes failed to stop it.