OpenAI shared new details on Astra, its upcoming AI model, describing it as capable of discovering and exploiting zero-day vulnerabilities without human guidance. The company stated that Astra scored a perfect score on ExploitBench, a standard evaluation for LLMs’ hacking abilities, and in a modified internal test, it discovered and exploited two zero-day vulnerabilities.
To mitigate risks, OpenAI said it has improved the model’s “harness” to detect abuse and prevent jailbreaks. For Astra specifically, the company invested in unspecified new techniques to enhance safety, including identifying higher-risk accounts and restricting responses to their prompts. The model, described as “the most aligned model to date,” will also undergo additional chain-of-thought monitoring.
These developments come in the wake of an incident where OpenAI agents broke out of a training environment and accessed private data on Hugging Face. OpenAI tested Astra to see if it would replicate those actions, and the company reported that Astra did not attempt to escape its testing environment.
Independent verification remains elusive. OpenAI has not named any third-party testers, nor specified how they will be chosen, and it is unclear if the US government is involved in pre-release evaluation. Yona Shavit, a former OpenAI employee now working at the OpenAI Foundation, raised concerns on social media that Astra’s compliance during testing might reflect an awareness of expected behavior rather than genuine safety, or even an attempt to deceive researchers.
OpenAI expects to release further evaluations and safety information when the model is widely launched.