OpenAI Rolls Out GPT-6 Astra, Touting Advanced Cybersecurity Skills but Mixed Benchmark Results

New model matches Claude Fable on pricing, leads in coding efficiency, but trails on intelligence index

edit
By LineZotpaper
Published
Read Time2 min
Sources3 outlets
OpenAI has begun rolling out GPT-6 Astra, its latest flagship language model, to a limited set of organisations today with broader availability for ChatGPT subscribers and API users expected in the coming days. The model, priced identically to Anthropic's Claude Fable 5 and 5.1 at $10 per million input tokens and $50 per million output tokens, scores a perfect 100% on ExploitBench and 99.9% on the ARC-AGI 3 benchmark under a custom evaluation harness, though independent assessments show it lags behind rivals on general intelligence metrics.

OpenAI announced the release of GPT-6 Astra on Wednesday, positioning it as a direct competitor to Anthropic's Claude Fable series. According to the company's announcement, cited by Simon Willison, the model is "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS." A separate report indicates that companies participating in OpenAI's application-based cybersecurity program will receive early access.

Astra demonstrates exceptional performance on security-related tasks, scoring 100% on ExploitBench—well above GPT-5.6 Sol's 78.5%—and 42.4% on ExploitGym compared to Sol's 30.3%. On SRE-Bench binary reverse engineering, it achieved 99.2% within four attempts versus Sol's 68.7%. The model also shows improved long-context handling, achieving 100% accuracy on OpenAI's eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens.

However, independent analysis from Artificial Intelligence platform Artificial Analysis found that Astra scores equal to GPT-5.6 Sol on its Intelligence Index at 61 points, five points lower than Claude Fable 5.1 with fallback and also trailing Meta's Muse Spark 1.3. On the Coding Agent Index, Astra leads for cost efficiency, costing less than half of Claude Fable 5 per task for the same score.

Notably, Astra's 99.9% score on the ARC-AGI 3 benchmark was achieved using OpenAI's custom "Provider Adapter harness," which preserves opaque reasoning state between requests. Under the standard ARC-AGI harness, the model scored 62.7% at a higher cost of $26,000 compared to $19,000 with the custom setup. The ARC-AGI project noted that the custom harness allows the model to reuse prior work, potentially explaining the dramatic score difference.

OpenAI has not published benchmarks for Astra on several widely used evaluations, and the model's API label will be gpt-6-astra once fully deployed.

§

Analysis

Why This Matters

  • Astra introduces pricing that matches Anthropic's Fable series, signaling intensified competition in the premium AI model market.
  • Its exceptional cybersecurity scores point to potential use in automated penetration testing and vulnerability research, raising both opportunities and safety concerns.
  • The discrepancy between custom and standard benchmark results underscores ongoing debates about evaluation methodologies in AI.

Background

OpenAI's GPT series has evolved through several generations, with each iteration aiming to improve reasoning, context handling, and task performance. The company has increasingly focused on safety and cybersecurity, partly in response to incidents like the recent Hugging Face security breach. The new model arrives amid a crowded field of competing models from Anthropic, Meta, and others, with pricing emerging as a key differentiator.

Key Perspectives

OpenAI: Positions Astra as a significant step forward, particularly in security and long-context tasks, while maintaining competitive pricing against Anthropic's offerings. Anthropic and other rivals: Claude Fable 5 still outperforms Astra on the Intelligence Index, suggesting that general reasoning capabilities remain a strength for Anthropic's models. Critics and evaluators: The massive gap between ARC-AGI scores with custom versus standard harnesses raises questions about the validity of some benchmark claims, highlighting the need for standardised evaluation frameworks.

What to Watch

  • Independent third-party benchmarks on standard evaluations like MMLU, HumanEval, and other widely-used tests.
  • Adoption rates among enterprises and developers given the pricing parity with Claude Fable.
  • Reports on how the model's cybersecurity capabilities are being used or restricted, especially given early access for security program participants.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.