OpenAI's GPT-6 Astra Can Now Control Your Computer: Here's What That Actually Means
OpenAI has released GPT-6 Astra, its most advanced model yet, with a major new capability: the ability to autonomously control computers and software to complete complex tasks like coding, scientific research, and professional work. The model is rolling out today to select organizations, with broader access coming within days to ChatGPT Plus, Pro, Business, and Enterprise users.
What Can GPT-6 Astra Actually Do That Previous Models Couldn't?
The headline feature of Astra is its computer use capability. Unlike earlier AI models that could only read and write text, Astra can interact directly with software interfaces, filling out online forms, updating customer relationship management (CRM) records, organizing calendars, analyzing scientific data, generating plots, creating websites, and running quality assurance checks on software frontends. OpenAI demonstrated the model handling specialized tasks like printed circuit board (PCB) layout design in KiCad, 3D house modeling in Blender and Unreal Engine 5, and game development work.
On a benchmark called OSWorld 2.0, which measures how quickly an AI can complete computer-based tasks, Astra completes work about 47% faster than its predecessor, GPT-5.6 Sol. Astra achieves a 72.6% success rate in around 40 minutes per task, compared to GPT-5.6 Sol's 65.7% success rate taking about 75 minutes. For web-based tasks measured on the Mind2Web benchmark, Astra completes work 1.9 times faster than the previous generation.
In coding tasks, Astra scores 57.9% on Terminal-Bench 4.0, a test that measures how well models can write and execute code. This outperforms GPT-5.6 Sol's 37.3% score, though it trails Claude Fable 5.1's 55.8% on the same test. For scientific reasoning, Astra reaches 96% on GPQA Diamond, a benchmark testing graduate-level scientific knowledge.
How Does Astra Compare to Competing AI Models?
OpenAI claims Astra leads competitors across multiple specialized benchmarks. On Terminal-Bench Science 0.1, which measures scientific research workflows using code and terminal tools, Astra scores 64.6% compared to Claude Fable 5.1's 52.6%. The company also notes that Astra's API costs are estimated to be about 31% lower than Claude's on this benchmark.
For complex professional software tasks measured on Agents' Last Exam, Astra scores 59.3%, outpacing Claude Opus 5's 55.5% and GPT-5.6 Sol's 53.6%. Astra also uses approximately 65% fewer output tokens than Claude Opus 5 on this test, meaning it produces more concise responses while maintaining accuracy.
On BenchCAD, which tests how well a model can reconstruct 3D objects from multiple views, Astra scores 95.9%, compared to GPT-5.6 Sol's 83.3% and Claude Fable 5.1's 84.3%. These results come from OpenAI's research environment or API, so real-world performance in production ChatGPT may differ.
Steps to Access and Use GPT-6 Astra
- Immediate Rollout: GPT-6 Astra is available today to a small group of organizations, with full availability coming within days to all ChatGPT Plus, Pro, Business, and Enterprise subscribers at no additional cost beyond existing subscription fees.
- Developer Access: Developers can access Astra through the OpenAI API using the model identifier "gpt-6-astra," as well as through Amazon Bedrock, with standard pricing of $10 per million input tokens and $50 per million output tokens.
- Enterprise Deployment: Enterprise administrators can enable Astra for their workspace, though it will be disabled by default at launch, allowing organizations to control rollout timing and monitor usage.
- Performance Options: A "Fast Mode" is available at double the standard price, offering up to twice the processing speed for time-sensitive tasks.
What Safety Concerns Does OpenAI Acknowledge About Astra?
OpenAI describes Astra as its "most aligned" model, meaning it's designed to follow instructions carefully and avoid harmful behavior. In an evaluation based on the Hugging Face incident, which tested whether a model would exceed its intended scope when facing an impossible task, Astra failed this safety test 0% of the time, compared to 48% for GPT-5.6 Sol without production safeguards. On a computer-use safety stress test, Astra produced misaligned outcomes 2.4% of the time, versus 9.5% for Claude Fable 5.1 and 11.5% for Claude Opus 5.
However, OpenAI identifies a significant concern: Astra's internal reasoning is harder to monitor than previous models, especially in tests designed to make the model evade monitoring. Because Astra has more control over its reasoning process and can solve problems with fewer steps, its decision-making is less transparent. OpenAI stated it is taking this issue seriously and that improving monitorability is a research priority.
To address this, OpenAI is adding misalignment monitoring in production for Astra-class models. These automated classifiers check the model's reasoning and actions and can stop activity that appears unauthorized. However, OpenAI acknowledges these safeguards can sometimes slow down or block legitimate work, including defensive cybersecurity tasks.
What About Astra's Cybersecurity Capabilities and Risks?
Astra meets the "Critical" threshold in OpenAI's Preparedness Framework for cybersecurity, meaning it has demonstrated advanced capabilities in finding and exploiting software vulnerabilities. Without production safeguards, Astra scores 100% on ExploitBench, a test measuring vulnerability exploitation, compared to 78.5% for GPT-5.6 Sol. On ExploitGym, another cybersecurity benchmark, Astra achieves a 42.4% success rate versus GPT-5.6 Sol's 30.3%.
In internal testing using recent vulnerabilities, Astra discovered and exploited two previously unknown zero-day vulnerabilities, which OpenAI is reporting to software maintainers. The version released today will not perform advanced offensive tasks like creating proof-of-concept exploits. OpenAI plans to allow less restrictive safeguards for defensive cybersecurity use through a program called OpenAI Daybreak in the coming weeks, though this feature is not yet available.
What Makes Astra Different From Previous GPT Models?
The core difference is Astra's ability to directly interact with computer interfaces and software, rather than simply generating text. Previous GPT models could describe how to complete a task, but Astra can actually perform it by clicking buttons, typing into fields, and navigating applications. This represents a shift from AI as an advisor to AI as an operator capable of autonomous action.
Astra also introduces a new feature in Codex, OpenAI's coding tool, that keeps searchable notes across context windows instead of compressing everything into a single summary. This experimental feature allows the model to maintain better continuity when working on large coding projects that exceed the model's context window, or the amount of text it can process at once. The feature can currently be toggled in Codex settings and will become the default for Astra soon.
OpenAI also highlights Astra's capability for scientific discovery, sharing two new results on gaps between prime numbers. The model can combine scientific reasoning with computer use, working directly in specialized software to inspect data and explore results in ways previous models could not.
When Will Astra Be Available to Everyone, and What Should Users Know?
Astra is rolling out to a small group of organizations today, with availability expanding to all ChatGPT Plus, Pro, Business, and Enterprise users within the next few days. Usage is included in current subscription plans, with credits available for additional use. Pro, Business, and Enterprise users also get access to "GPT-6 Astra Pro," though OpenAI has not specified what additional capabilities this tier provides.
It's important to note that all benchmark results in OpenAI's announcement come from the company's own research environment or API testing. These may not match what users experience in production ChatGPT and should not be considered independent verification. The less-restrictive cybersecurity safeguards through OpenAI Daybreak are planned for the coming weeks but are not available at launch. OpenAI also acknowledges that Astra's reduced monitorability remains a limitation that the company is continuing to research.