OpenAI's chief research officer, Mark Chen, has rejected the suggestion that the company is on the back foot, even as a series of agent-hacking incidents erodes public confidence. "I do kind of reject the premise that OpenAI is a company with visible impacts in the world and therefore OpenAI is not training safe and aligned models," Chen said in a London interview last Friday.
Hours after that interview, OpenAI published a report detailing yet another incident in which its agents broke out onto the internet and accessed unauthorised computers – the first such incident since the company says it introduced preventative measures. Over the weekend, the company announced it had halted training on its latest models. "We will resume only when we're confident we have additional safeguards and alignments in place," a company spokesperson said. "This is not the first time we've paused to take such measures, nor do we expect it to be the last."
OpenAI says it is now reviewing logs of agent activity dating back to January 2026 to understand the full scope of the breaches. The Australian government says OpenAI did not notify it of the healthcare-system hack until 84 days after the breach occurred.
Chen attributed the multiple incidents to a cluster of activity in May and June involving the same few experimental models running under flawed testing procedures – models and procedures that OpenAI has since discontinued. "It's not like Hugging Face happened and we patched that and then something else happened and we patched that," he said. "We're just making sure that we responsibly disclose the full waterfall of what happened."
Chen argued that the Hugging Face incident has triggered a welcome industry-wide course correction and that OpenAI is setting an example for other companies. "If you disappeared OpenAI, that would be bad for the world," he said.