- Astra reached OpenAI’s Critical cybersecurity threshold.
- OpenAI tightened Astra’s access and safeguards.
OpenAI has classified its upcoming Astra model at the Critical cybersecurity capability level under its Preparedness Framework, making it the first of the company’s models to reach the threshold.
The company said Astra can identify previously unknown security vulnerabilities and develop ways to exploit them across hardened systems with limited human involvement. OpenAI plans to restrict access to some of the model’s advanced cybersecurity capabilities when it is released.
Under OpenAI’s Preparedness Framework, a model meets the Critical cybersecurity threshold if it satisfies either of two conditions. A model can meet the threshold by independently finding and developing functional zero-day exploits across hardened real-world systems, or by devising and executing new end-to-end attack strategies from a high-level objective.
Under the framework, systems assessed at High capability must have safeguards in place before deployment, while systems reaching Critical capability also require protections during development. Astra is the first OpenAI model to trigger those Critical-level requirements.
OpenAI said the safeguards must address two risks: malicious users employing Astra to develop exploits or conduct attacks, and the model taking unauthorised actions without a malicious user directing it. The company said protections against unauthorised model actions apply during both internal development and external deployment.
Astra crosses OpenAI’s Critical cyber threshold
OpenAI based its assessment on automated benchmarks, private evaluations, and testing led by security experts. OpenAI said Astra outperformed GPT-5.6 Sol in vulnerability identification and exploit development while using fewer output tokens.
Astra achieved a 100% score on ExploitBench, a benchmark that tests whether a model can develop exploits from known vulnerabilities. OpenAI also created an internal version using 20 high-severity vulnerabilities in Google’s V8 JavaScript engine that were disclosed between June and August 2026.
During that evaluation, Astra discovered two previously unknown vulnerabilities and used them as part of an exploit chain, OpenAI said. The company is working to disclose the vulnerabilities to the relevant maintainers.
Expert-led testing also examined Astra against a hardened browser and operating system. OpenAI said the model identified previously unknown vulnerabilities and produced a browser exploit chain that escaped a sandbox and executed commands on the host system after the browser opened an HTML file.
In separate operating system testing, Astra found multiple vulnerabilities and combined them into a local privilege-escalation chain. This allowed the model to move from an unprivileged user account to root access, according to OpenAI.
OpenAI said the Astra results shown for its internal ExploitBench comparison reflect capabilities with Daybreak Blue access, rather than the model’s default production configuration.
Stronger controls shape Astra’s development and access
OpenAI had already raised concerns about Astra’s cybersecurity capabilities in August. On August 7, the company said preliminary evaluations meant it could no longer rule out the possibility that Astra would reach its Critical threshold.
The company introduced stricter controls covering isolated testing environments, network and tool access, model weight protection, monitoring, and sandboxed execution. It also paused Astra-related activities that did not meet the new security requirements.
OpenAI said certain frontier training, including work on Astra, was paused for two weeks after the Hugging Face incident while it strengthened isolation, network controls, monitoring, and alignment requirements. Smaller-scale work later continued under stricter controls.
OpenAI kept some larger reinforcement learning runs on hold while it introduced additional requirements for their training environments. One large frontier RL run restarted on August 28, although some smaller experimental runs remained paused as of September 1. Reuters also reported that OpenAI had resumed its largest training run while continuing to hold back some experiments.
Those measures followed a separate security incident involving Hugging Face. Astra was not involved, but OpenAI said it incorporated lessons from the event into Astra’s safety approach.
An independent investigation by METR and Redwood Research found that about 700 OpenAI agents were involved in the Hugging Face incident, a figure OpenAI confirmed to Reuters. The agents had been operating as part of cybersecurity evaluations and were intended to remain within isolated environments.
The investigation also found attempts by some agents to manipulate or conceal records of their actions. Reuters reported that OpenAI subsequently strengthened monitoring and security controls around its research infrastructure.
OpenAI also developed an ExploitGym “honeypot” evaluation informed by the Hugging Face incident. The test examined whether models would target surrounding security infrastructure instead of completing their assigned cybersecurity tasks.
Without production safeguards, GPT-5.6 Sol attempted to access the honeypot targets in 56% of tests, while Astra made no such attempts under the same test conditions, according to OpenAI. The company said Astra was still able to solve some tasks through legitimate methods.
OpenAI also tested how Astra responded after an automated security review rejected an action. OpenAI’s evaluation chart showed GPT-5.6 Sol successfully circumventing auto-review in 5.3% of samples, while Astra did not attempt to circumvent it.
OpenAI said Astra refused 91.5% of requests in its cyber jailbreak evaluation set, compared with 59% for GPT-5.6 Sol. Its safeguards combine model refusals with system-level classifiers, monitoring, and controls designed to identify potentially harmful or unauthorised activity.
Access to Astra’s advanced cybersecurity workflows will initially be limited to a small group of testers. OpenAI said access through Daybreak Blue will expand later for defensive cybersecurity work, while the model’s default production configuration will not expose the same level of capability.
The additional safeguards can also interrupt legitimate activity. OpenAI said potentially risky actions may be slowed, paused, or stopped, including some long-running tasks and work that is not directly related to cybersecurity.
When monitoring pauses an action in ChatGPT or Codex, users may be asked to review it before continuing. For Astra-class models accessed through the API, OpenAI said the task will stop when the monitoring system intervenes. The company said these controls can also interrupt legitimate defensive security work.
Other AI developers set cyber restrictions
Anthropic has also introduced restrictions on advanced cybersecurity assistance. Its safeguards for Fable 5 classify activities including exploit development, privilege escalation, and some forms of vulnerability discovery as high-risk dual-use work.
Anthropic said such activities can support legitimate security testing but can also be used maliciously, making user authorisation relevant to access. Its controls combine access restrictions, model safety training, classifiers, and offline monitoring.
Google DeepMind uses a separate Frontier Safety Framework that defines Critical Capability Levels in areas including cybersecurity. The framework links those thresholds to additional security and deployment controls.
DeepMind said Gemini 3.7 Flash reached the alert threshold for its cybersecurity Critical Capability Level but did not reach the CCL itself. The company said it continues to apply cybersecurity mitigations. OpenAI and DeepMind use different definitions and evaluation methods, so the results are not directly comparable.
The Financial Stability Board has also raised concerns about frontier AI and cybersecurity. In an August 31 letter to G20 finance ministers and central bank governors, FSB chair Andrew Bailey identified frontier AI’s effect on cyber risk as the organisation’s most immediate AI-related concern for the financial system.
The FSB said frontier AI may materially alter the speed, scale, and economics of cyber risk. Reuters also reported Bailey’s warning that faster vulnerability discovery could reduce the time available for existing security testing and recovery processes.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.
Tech Wire Asia is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

