OpenAI announced the release of GPT-6 Astra on Wednesday, positioning it as a direct competitor to Anthropic's Claude Fable series. According to the company's announcement, cited by Simon Willison, the model is "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS." A separate report indicates that companies participating in OpenAI's application-based cybersecurity program will receive early access.
Astra demonstrates exceptional performance on security-related tasks, scoring 100% on ExploitBench—well above GPT-5.6 Sol's 78.5%—and 42.4% on ExploitGym compared to Sol's 30.3%. On SRE-Bench binary reverse engineering, it achieved 99.2% within four attempts versus Sol's 68.7%. The model also shows improved long-context handling, achieving 100% accuracy on OpenAI's eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens.
However, independent analysis from Artificial Intelligence platform Artificial Analysis found that Astra scores equal to GPT-5.6 Sol on its Intelligence Index at 61 points, five points lower than Claude Fable 5.1 with fallback and also trailing Meta's Muse Spark 1.3. On the Coding Agent Index, Astra leads for cost efficiency, costing less than half of Claude Fable 5 per task for the same score.
Notably, Astra's 99.9% score on the ARC-AGI 3 benchmark was achieved using OpenAI's custom "Provider Adapter harness," which preserves opaque reasoning state between requests. Under the standard ARC-AGI harness, the model scored 62.7% at a higher cost of $26,000 compared to $19,000 with the custom setup. The ARC-AGI project noted that the custom harness allows the model to reuse prior work, potentially explaining the dramatic score difference.
OpenAI has not published benchmarks for Astra on several widely used evaluations, and the model's API label will be gpt-6-astra once fully deployed.