OpenAI Releases GPT-6 Astra, a Model That Uses Computers and Targets Advanced Coding

New AI agent can fill forms, write code, and browse the web; cybersecurity classification raises safety questions

edit
By LineZotpaper
Published
Read Time2 min
OpenAI has released GPT-6 Astra, a model that can interact directly with graphical software interfaces to perform multi-step tasks such as filling forms, updating records, and troubleshooting, while also achieving strong benchmark scores in coding and security-critical capabilities.

OpenAI has released GPT-6 Astra, a new model focused on computer use, coding, professional workflows, science, and cybersecurity. The model is initially available to a limited set of organizations and is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, as well as the OpenAI API, Microsoft Azure, and AWS Bedrock.

Astra extends OpenAI's models beyond generating responses toward performing multi-step tasks directly in software. It can interact with graphical interfaces to fill forms, update CRM records, conduct research, create websites, analyze data, install and test software, and troubleshoot problems visible on screen. On the OSWorld 2.0 benchmark, OpenAI reports a score of 72.6%, compared with 65.7% for its predecessor GPT-5.6 Sol.

Coding is another focus of the release. OpenAI reports 57.9% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1. Astra also introduces an experimental context mechanism in Codex that allows the agent to maintain notes across context windows instead of relying only on compaction. Previous context windows remain searchable, allowing the model to retrieve earlier requirements, test results, and tool outputs during long-running coding tasks.

The model supports long contexts of up to one million tokens in OpenAI's reported MRCR evaluations, scoring 96.3% in the 512K-to-1M range. OpenAI also reports improvements in professional tasks including database migrations, CAD generation, data science, browser research, and scientific workflows.

Cybersecurity represents a significant change from previous OpenAI models. Astra is the first OpenAI model classified at the critical cybersecurity capability level under the company's Preparedness Framework. In testing without production safeguards, OpenAI says the model discovered and used two previously unknown vulnerabilities and demonstrated the ability to develop exploits against hardened browsers and operating systems. The production version restricts advanced offensive tasks, while OpenAI plans to provide broader defensive capabilities through its security-focused offerings.

§

Analysis

Why This Matters

  • Astra moves AI from generating text to acting directly on software, potentially automating a wide range of professional tasks from data entry to software maintenance.
  • The critical cybersecurity classification signals that frontier AI models may soon pose risks or benefits in offensive security, raising the stakes for responsible deployment.
  • Long context and persistent memory across sessions could make AI agents more practical for complex, multi-hour workflows in coding and research.

Background

OpenAI has been progressively releasing GPT models with increasing reasoning and agentic capabilities. GPT-5 and its variants introduced multi-step planning and tool use. GPT-6 Astra represents a step further into autonomous computer interaction, building on research into agents that can control graphical user interfaces. The Preparedness Framework, which OpenAI introduced to evaluate catastrophic risks, classifies models by capability level in areas such as cybersecurity, biosecurity, and persuasion.

Key Perspectives

Developers and enterprises: Could benefit from automation of routine coding, debugging, and system administration tasks, potentially boosting productivity. Security researchers: May welcome defensive tools but also worry about misuse if safeguards are insufficient; the discovery of exploits in testing highlights dual-use potential. Policy advocates: Likely to call for transparency and third-party auditing given the critical cybersecurity classification and the model's demonstrated ability to find unknown vulnerabilities.

What to Watch

  • How quickly Astra is adopted in enterprise production environments and whether reported benchmark gains translate to real-world productivity.
  • Any safety incidents or exploits that emerge from attempts to bypass Astra's production safeguards.
  • Responses from competitors and regulators as AI agents with computer-use capabilities become more common.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.