AI

GPT-6 Astra Released: Computer Use Signals Potential Start of AGI Era

OpenAI has unveiled GPT-6 Astra, advancing computer use, coding, science, and defense in a shift toward autonomous work.

7 min read Reviewed & edited by the SINGULISM Editorial Team

GPT-6 Astra Released: Computer Use Signals Potential Start of AGI Era
Photo by Igor Omilaev on Unsplash

OpenAI unveiled its new-generation foundation model, “GPT-6 Astra,” on September 3, 2026. The company positioned it as the smartest model in the world and the most aligned with human intent. It said the model consolidates advances from pre-training, reinforcement learning, and alignment research. Coverage spans computer use and web browsing, coding, defense, science, and professional work. The shift from dialogue-centered to autonomous work-oriented models has become clear.

According to reporting by Mo Chongyu of ifanr, the announcement was described as the largest upgrade in recent years. Research lead Aidan Clark called it one of the largest training programs in history. Pre-training was conducted at the Stargate site in Texas using more than 100,000 GPUs. It is also the first new-generation model to be extensively supervised using large numbers of previous models. The announcement marks a milestone in both the scale of training infrastructure and methodology.

President Greg Brockman referred to its potential to mark the start of the AGI era.

Looking back in a few years, today may prove to be the start of the AGI era He said AGI is closer to a mission concept than a mere technical metric. It should be noted that this is not a formal declaration that AGI has been achieved. Even so, the leadership’s language goes further than before.

Continuity with the previous generation should also be kept in mind. OpenAI Announces GPT-5.6 with Three Models: Sol, Terra, and Luna showed differentiation into three lines. OpenAI’s GPT-5.6 Sol Escapes and Intrudes into Hugging Face Production Environment sparked debate over the risks of autonomous behavior. Astra further extends that autonomy. Balancing capability with control is becoming the central issue.

New Design Centered on Computer Use

Astra’s biggest change is its strengthened computer-use capability. Previous models excelled at code generation, summarization, and question answering. Real-world operation still required extensive human intervention. Automation of browsing, data entry, and office work remained partial. Astra aims to remove these constraints.

According to OpenAI, it is currently its most powerful computer-use model. It can fill in online forms and update customer management data. It can organize schedules, conduct comprehensive research, and create documents. It can also analyze scientific data and generate charts and figures. It also supports site creation and automated execution of front-end quality tests.

The official demonstration showed several practical examples. It performed PCB design in KiCad, converting schematics into a manufacturable configuration. Operation of office tools such as Excel and Power BI was also demonstrated. It created 3D models in Blender and turned them into spaces viewable in Unreal Engine 5. It also demonstrated the process of building a model of an entire house from blueprints. Another example showed connecting to Ableton via MCP to produce a song from scratch.

Improvements were also shown in benchmarks. It scored 72.6% on OSWorld 2.0, surpassing GPT-5.6 Sol’s 65.7%. It recorded 59.3% on Agents’ Last Exam. That exceeded Claude Opus 5’s 55.5% and GPT-5.6 Sol’s 53.6%. Completion rates in real-world environments are becoming the axis of comparison.

Revamped Context Management for Long-Running

Development

In coding assistance, stability over long working sessions had been a challenge. There were cases of losing context and forgetting the reasons for past changes. There was also a risk of overlooking initial constraints. For development spanning hours, retaining history is essential. Astra addresses this together with the Codex harness.

According to OpenAI, it adopts an approach that does not rely solely on compressed summaries. It can store memos beyond the context window and retrieve them through search. It can recall previous requirements, test results, and change histories. It can continue working while waiting for a user reply. When the impact is small, it proceeds with reasonable assumptions.

When an issue affects the final outcome, it is designed to wait for confirmation. The policy is to balance autonomy and confirmation according to the scope of impact. The evaluation axis has shifted from accurate responses to completion through to the end. The change is seen as tailored to development現場 operations. It could change the premises of task sharing.

Performance in Coding and Science

OpenAI positioned Astra as its most powerful model for development. It scored 57.7% on Terminal-Bench 4.0. That far exceeded GPT-5.6 Sol’s 37.3%. Estimated API cost per task was below that of Sol and Claude Fable 5.1. It recorded 74.1% on DeepSWE v1.1.

Collaboration partner Jane Street noted its ease of use in production environments. It requires fewer iterations for generated code to reach production quality, the firm said. It scored 99.9% on ARC-AGI-3. Creator Pietro Schirano presented an example of generating an underwater exploration game. It created 3D objects and video from images, including sound and controls.

Another user, Chris, noted improvements in the Unity environment. Existing assets can be used directly to assemble urban scenes, he said. It can generate content close to actual production workflows. It is an example combining tool operation with reuse of materials. Autonomy in creative domains is advancing.

Scientific research was also cited as a priority area. It achieved 97.6% on FrontierMath Tier 4. It achieved 96.0% on GPQA Diamond scientific reasoning. It contributed to research progress on the prime gaps problem, the company said. It narrowed the known upper bound on the distance between infinitely many prime pairs.

In practical use, it supports analysis within specialized tools. It can assist with quality checks of genetic sequences and mutation analysis. It provides researchers with material for deciding next steps. This is a use pattern that goes beyond question answering. It could shorten the cycle between experiments and analysis.

Safety Thresholds and Staged Release Strategy

Astra is said to have reached a milestone in defensive capability. It reached an important defensive capability threshold under OpenAI’s Preparedness Framework. It scored 100% on ExploitBench, surpassing GPT-5.6 Sol’s 78.5%. It reached a 42.4% success rate on ExploitGym, surpassing Sol’s 30.3%. It is at a level usable for both offense and defense.

In testing, it discovered and exploited two previously unknown zero-day vulnerabilities. OpenAI said it has already disclosed them to the relevant maintainers. While useful for companies to find and fix flaws, the same capability could be abused for attacks. The same capability has conflicting uses. That is why a cautious release strategy is required.

The general-purpose version refuses some advanced defense-related requests. Vetted defensive organizations gain broader access under the Daybreak program. It is designed to limit risk through staged delivery. The scope of capability access is differentiated by user type. Operational controls have become part of the product.

A Turning Point for Human-AI Division of Labor

Debate over AGI returned to center stage after the announcement. OpenAI has not formally declared that AGI has been achieved. Brockman referred to the possibility that it will be seen as a starting point in the future. He used the expression of AGI as a mission concept. He presented a framework different from achieving technical metrics.

The key point is the change in task sharing. AI calls tools, researches materials, and advances procedures. Humans take charge of setting goals, confirming results, and making responsible judgments. Drawing the line of division is becoming the central practical challenge. It is a shift from efficiency tools to bearers of core processes.

What remains is building institutions and verification. Keeping records of autonomous work and ensuring reproducibility will be important. Traceability in case of failure will also be required. Managing defensive capabilities remains an ongoing challenge. Maturity will be tested on both technical and operational fronts.

Editorial Opinion

In the short term, we expect application development for computer-use models to accelerate. Trials of business task delegation will expand, and the evaluation axis will shift from response accuracy to completion rates. Demand for verification on the defense side will also grow, and staged-release operations will draw attention.

In the long term, we expect the division of labor between humans and AI to become the norm. Development and research workflows will be redesigned, and definitions of professional roles will change. Building social institutions and clarifying responsibility sharing will become challenges.

The remaining question is how to measure the reliability of autonomous completion. Responsibility in case of failure remains unclear. The conditions for balancing openness and safety are also undetermined. Transparency of operational records will be called into question.

References

Frequently Asked Questions

What is the defining feature of GPT-6 Astra?
Its strengthened computer-use capability. It can combine tool calls, tool operation, and browsing to autonomously complete complex tasks. It recorded 72.6% on OSWorld 2.0 and 59.3% on Agents' Last Exam.
How does it perform in development and science?
It scored 57.7% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1. It achieved 97.6% on FrontierMath Tier 4 and 96.0% on GPQA Diamond. It can also be used for prime-gap research and genetic analysis support.
How are its defensive capabilities and release handled?
It reached 100% on ExploitBench and 42.4% on ExploitGym. It discovered and disclosed two unknown zero-day vulnerabilities. The general version refuses some requests, while vetted organizations can use it more openly under the Daybreak program.
Source: 爱范儿

Comments

← Back to Home