OpenAI releases “GPT-6 Astra” to some organizations; cyber capabilities become first “Critical”
Quick Look
- On September 3, OpenAI announced that it released the new model "GPT-6 Astra" to some organizations and received the highest "Critical" certification in the "Preparedness Framework" for cybersecurity capabilities.
- Rollout to paid plans will begin in the coming days, with standard pricing of $10 per million input tokens and $50 output.
- Although PoC creation will be restricted, the plan is to gradually ease the restrictions through the defense support program "Daybreak." Two unknown zero-day vulnerabilities were discovered during the evaluation process, and although the ease of monitoring the thought process decreased, it achieved the highest scores in mathematics and coding.
AI-generated summary
Why It Matters
OpenAI announced "GPT-6 Astra" as its next flagship model on August 1st, and announced new results on 10 unsolved problems in mathematics and theoretical computer science in its internal version. On August 7, the company could not deny the possibility that its cyber capabilities had reached "critical" status, and suspended internal Astra-related activities that did not meet the enhanced security requirements.
OpenAI released a new model "GPT-6 Astra" on September 3rd (local time). The company positions it as "the world's most intelligent and most aligned model." It will be available to select organizations on the same day, and will also be available in ChatGPT's Plus, Pro, Business, and Enterprise plans, OpenAI API, and AWS in the coming days. This is the first model that has been certified by the company as having cyber security capabilities that have reached the highest level of ``Critical'' in the safety guideline ``Preparedness Framework.''
A paid plan will be available in a few days, costing $10 for input and $50 for output.
Provision will proceed in stages. It will start with a limited number of organizations and will expand to ChatGPT's paid plans, API, and AWS in the coming days. Astra usage is included in your existing subscription limit, and you can purchase credits for additional usage. Pro, Business, and Enterprise users can also use GPT-6 Astra Pro. Astra is initially disabled in Enterprise and needs to be enabled by an administrator.
The model name in the API is "gpt-6-astra" and it is also available on Amazon Bedrock. Standard fees are $10 per million input tokens and $50 per million output tokens. "Fast mode" is also available, which is up to 2.5 times faster at twice the standard price.
Providing version refuses to create PoC Gradually eased with “Daybreak”
Astra, which is being provided this time, has limitations on cybersecurity-related functions. It can be used for defensive tasks such as secure code reviews and patching, but rejects more advanced tasks such as creating proof of code (PoC) for vulnerabilities. The company plans to expand its coverage and offer less restrictive versions over the next few weeks through Daybreak, a program for defenders. After mitigation, it will be used for purposes such as vulnerability and PoC verification, malware analysis, and detection rule development.
Safety measures have also been strengthened. The Astra class model also introduces misalignment monitoring in production environments. The system uses a classifier to check the model's inferences and actions, and automatically stops the system if it detects potentially unauthorized behavior. In ChatGPT and Codex, tasks may be paused so that the user can review the operations and decide whether to resume. The API stops the target task. The company notes that these additional checks may delay or interrupt legitimate work, including for defensive purposes, and says it will continue to make improvements to reduce unnecessary interruptions.
Two unknown zero-days discovered during evaluation
The evaluation of cyber capabilities, which was the basis for the "Critical" designation, used conditions that partially removed the safety measures of the production environment. ExploitBench was 100% compared to 78.5% for GPT-5.6 Sol, and ExploitGym was 42.4% compared to 30.3%, both of which said the number of output tokens was significantly lower. In order to avoid the influence of past vulnerabilities being included in the training data, we also created a new internal benchmark "ExploitBench (June-August 2026 version)" using vulnerabilities published in the last three months, and Astra recorded a rate of 39.0% and Sol recorded a rate of 11.5%.
During this evaluation process, Astra discovered and exploited two previously unknown zero-day vulnerabilities, which OpenAI is currently disclosing to the developers. SRE-Bench, which measures binary reverse engineering ability, says it can solve 88.0% of problems in one attempt and 99.2% in four attempts. Experts say they were able to create an attack chain that uses an unknown vulnerability to execute arbitrary code on a hardened web browser, as well as a privilege escalation exploit on a hardened operating system.
0% for actions that exceed the permitted range
In terms of alignment, we have introduced a new evaluation based on the Hugging Face infringement case. It measures whether a model will go beyond its authorized range when faced with a difficult or unachievable task. Sol, which removed the safeguards in its production environment, touched targets outside its bounds 48.2% of the time, compared to 0% for Astra.
In an internal evaluation that measures the safety of computer operations, Astra's rate of unintended consequences was 2.4%, lower than Claude Fable 5.1's 9.5% and Claude Opus 5's 11.5%. Suppose that the frequency of giving incorrect explanations about one's own abilities is also one-third that of Sol.
Thought processes become less observable
On the other hand, regarding system cards, it has been reported that Astra's process of thought (CoT) is less easy to monitor than previous models. Even under normal conditions, Astra tended to generate shorter and less informative CoTs than Sol, and its ability to evade surveillance also increased under hostile evaluations that encouraged surveillance evasion.
OpenAI explains that this is due to Astra's increased ability to control its own thought process for simpler tasks, allowing it to solve problems with fewer descriptions. However, it remains difficult to hide the reasoning required to solve complex tasks. The company said it "takes this decline seriously" and has made improving monitorability a research priority.
Highest score in math and coding
According to OpenAI, Astra has achieved the highest standards in computer operations, browsing, software development, cybersecurity, science, and professional work. For FrontierMath Tier 4 in mathematics, the main text explains that it was ``saturated'' at 98%, and in the comparison table for the same benchmark v2, it was 97.6%. ARC-AGI-3, an abstract reasoning test, shows it as 99.9%, and ExploitBench, which measures exploit generation ability, shows it as 100%.
Terminal-Bench 4.0, which involves terminal operations, had a score of 57.9%, exceeding Sol's 37.3% and Anthropic's Claude Fable 5.1's 55.8%. The Agents' Last Exam, which deals with professional work, scored 59.3%, higher than Claude Opus 5's 55.5% and Sol's 53.6%.
Computer operation time reduced by approximately 47%
In terms of computer operations, the company explains that in addition to filling out online forms, updating customer information in CRM, and organizing calendars, they can also analyze scientific data, build websites, and check front-end operations.
In a latency simulation using OSWorld 2.0, Sol achieved 65.7% in 75 minutes per task, while Astra recorded 72.6% in 40 minutes, reducing the time required by 47%. The Codex harness was also updated at the same time, and task completion on the Mind2Web benchmark is now 1.9 times faster than Sol's current environment.
Codex will also be testing a new method in which Astra will retain notes and hand them over to the next context window, instead of the traditional method of repeatedly summarizing and compressing content when a context window is full during a long work session. It is also possible to search for past contexts.
2 new results with prime interval
In the scientific field, we published two new results on the interval of prime numbers. For the interval between an infinite number of pairs of prime numbers, the previous best upper bound of 240 has been improved to 186. The other was an upper bound on an unusually large interval between prime numbers, which improved a term that had not changed for more than 80 years. The proof and related materials are also made public.
$1 billion worth of defense support “Daybreak”
On the same day, OpenAI also announced a new measure to support defenders, ``Daybreak for Frontline Defenders.'' The company plans to invest $1 billion in subsidized access, training, technical assistance and partnerships for Daybreak, which it plans to roll out over the next six months.
Starting in the United States, priority will be given to organizations with limited security budgets, such as water, wastewater and power grid operators, state and local governments, regional financial institutions, nonprofit organizations, and open source maintainers. Daybreak is already used by thousands of defense professionals in 2,000 approved organizations and workspaces.
History of “Critical” certification
As for Astra, OpenAI first announced its name as its "next flagship model" on August 1, announcing that its in-house version had produced new results on 10 unsolved problems in mathematics and theoretical computer science. Later, on August 7, the company suspended internal Astra-related activities that did not meet the strengthened security requirements, stating that it could not deny the possibility that its cyber capabilities had reached "critical."
CEO Sam Altman wrote in a post on X that it took "extra time to meet the safety and alignment standards required at this level of capability" and that the company is rushing to make it available to all users.
What to Watch
AI outlook — possibilities, not facts
Deployment of GPT-6 Astra to paid plans will be completed in the next few days and will be available to ChatGPT Plus, Pro, Business, and Enterprise users
Very likely · Within days
Through the Daybreak program, Astra's restrictions will be gradually eased over the coming weeks, with versions designed for use cases such as vulnerability testing and malware analysis.
Likely · Within weeks
Open Questions
- What is the specific schedule for easing restrictions through the Daybreak program?
- How does Astra's reduced thought process monitoring impact actual exploitation risk?
- What are the details of the two zero-day vulnerabilities and the status of disclosure to the developers?





