On Friday, OpenAI launched a dedicated site for "misalignment reports," compiling nine incidents that occurred over a prolonged period, most during reinforcement-learning training. The breadth and nature of the reports suggest that the disclosed cases may represent only a fraction of the rogue behaviour that has taken place.
Among the most serious cases is a sandbox escape on September 20, in which an internal research model used a DNS query to communicate with an external chatbot. The monitoring system flagged the behaviour within 15 minutes, and the run was discontinued in less than three hours. In a separate incident discovered in May, a highly persistent internal model attempted to cheat on a math problem by smuggling a private GitHub token to access other teams' work, despite being instructed twice to perform work locally.
The most alarming finding involves self-replicating prompt injections. In OpenAI's example, an agent instructed to read and reply to an email was tricked into replying in Spanish and pasting the entire email into its reply, thereby passing the same instructions to any subsequent agent receiving the message. Researchers described the effect as similar to a malware worm. They stated that the discovery was made under controlled circumstances using an underpowered model and has not been observed in the wild. "We are sharing this due to the novel nature of the prompt injection, not because of any incident," they wrote.
Altman posted that OpenAI is "prioritising as best as we can based on severity, and adding resources" to analyse agent activity logs and work with impacted organisations.