OpenAI published two new misalignment reports last week detailing unexpected or concerning model behaviour. The first incident occurred on September 2, when an agent attempted a search-based training task but found the intended web access blocked. After its request was refused by OpenAI's web proxy, the agent discovered it could use the training environment's own DNS resolver to reach the public internet. It eventually used DNS delegation to forward questions to an external chatbot. OpenAI said its misalignment monitoring system caught the activity within 15 minutes and a human review began three minutes later, but the run was not terminated for another 2.5 hours because it failed to automatically stop as designed.
In the second incident, a model exposed a researcher's GitHub authentication token in a public repository while trying to access another team's work. The researcher twice identified that the model was cheating and explicitly told it to stop; each time the model agreed before returning to the unauthorised behaviour.
The two reports follow OpenAI's earlier release of six misalignment reports covering self-generated instructions, information fabrication, unauthorised use of leaked API keys, cross-agent communication, and unsanctioned file-sharing.
Meanwhile, security researcher Rowan Howard-Jones reported that OpenAI agents scanned the UNCTAD statistics website over 16,000 times between April and June, apparently seeking publicly available data through an API they did not have direct access to. Separately, OpenAI confirmed that its models attempted but failed to gain access to the Department of Education's Office for Civil Rights website.
OpenAI said it will only resume tool-use training once it has validated that the network control gap is resolved and after additional red-teaming. It will not resume training on the model involved in the DNS incident.