OpenAI describes GPT-6 Astra as "our most intelligent model yet, with state-of-the-art performance in computer use, browsing, software engineering, science, and professional work." It is the first model to reach the "Critical" level for cybersecurity capabilities under the company's Preparedness Framework — a 22-page policy document that researchers have criticised as allowing OpenAI's CEO to deploy even more dangerous capabilities, especially if other AI developers do so.
"This means that, with the right tools and access, GPT‑6 Astra can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step," OpenAI said.
The company sought to reassure users by noting improvements over its predecessor, GPT-5.6 Sol. On a new evaluation informed by the Hugging Face incident — which tests whether a model goes beyond its intended scope — Astra scored 0 percent compared to Sol's 48 percent when production safeguards were removed. OpenAI also reported a hallucination rate of 2 percent on its internal benchmark, down from 9.4 percent, and said Astra is three times less likely than Sol to misrepresent its capabilities.
On ARC-AGI-3, Astra surpassed the human baseline in action efficiency, using fewer actions than the median tested human on 96 percent of levels, according to ARC Prize Foundation president Greg Kamradt. However, the cost per game was about $360 in GPU tokens versus $0.00067 for a human brain.
Artificial Analysis benchmarks show GPT-6 Astra matching GPT-5.6 Sol on its Intelligent Index with a score of 61, five points behind Claude Fable 5.1 and trailing Meta Muse Spark 1.3. As a coding agent, Astra scores 67, on par with several rivals, though Fable 5.1 tops coding agent scores at 70 in Claude Code. On cost per task, Astra fares better at $4.72 versus Fable 5.1's $9.18.
OpenAI's Codex harness improvements claim a 1.9x faster task completion rate and a new context maintenance approach that allows the model to keep notes across context windows, retaining accumulated details without repeated compression.