Skip to content
AI.info

The Pulse

OpenAI Names Astra Its First ‘Critical’ Cybersecurity Model

OpenAI says its upcoming Astra model can find previously unknown vulnerabilities and build exploit chains without step-by-step human guidance. The company is limiting advanced cyber access to selected testers after adding new monitoring, is

OpenAI Names Astra Its First ‘Critical’ Cybersecurity Model

AI.info Team ·

OpenAI says its upcoming Astra model discovered and used two previously unknown vulnerabilities during testing, making it the first model the company has designated at its “Critical” cybersecurity capability threshold.

The designation does not mean OpenAI has classified Astra as a critical risk. It means the model meets the company’s highest current threshold for cyber capability: with suitable tools and access, Astra can identify unknown flaws and develop ways to exploit them across hardened systems without a person directing every step.

“With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step,” said Amelia Glaese, an OpenAI vice president overseeing its safety work, in comments reported by Reuters. Glaese also said the extra measures may “sometimes slow, pause, or stop legitimate work,” and that the company would work to reduce those interruptions.

Two zero-days moved Astra beyond benchmark scores

OpenAI’s evaluation combined public and private benchmarks with expert-led testing. Astra scored 100% on ExploitBench, which measures whether a model can turn known vulnerabilities into working exploits. OpenAI said its previous frontier cyber model, GPT-5.6 Sol, scored 78.5% on the same test.

Because older vulnerabilities may have appeared in training data, OpenAI created an internal benchmark using 20 high-severity vulnerabilities disclosed between June and August 2026. Astra achieved higher arbitrary code-execution rates than GPT-5.6 Sol while producing substantially fewer output tokens. During those tests, the model found and used two zero-day vulnerabilities in a single exploit chain; OpenAI said it is disclosing both flaws to their maintainers.

Expert assessments produced more direct demonstrations. Astra built a browser-compromise chain that escaped a sandbox and executed commands on the host after the browser opened an HTML file. In a separate hardened operating-system test, it combined several vulnerabilities to move from an unprivileged account to root.

OpenAI’s threshold covers autonomous attack paths

Under OpenAI’s Preparedness Framework, a model reaches the Critical cybersecurity level if it can find and develop functional zero-day exploits across many hardened real-world critical systems without human intervention, or devise and execute a novel end-to-end cyberattack from a high-level objective.

OpenAI says Astra crossed that bar in testing without its standard production safeguards. The company says the results reflect access through Daybreak Blue, a restricted testing environment, rather than the default configuration planned for ordinary users.

The distinction matters because the company is presenting Astra’s capability as useful for defenders as well as attackers. The same systems that locate weaknesses can help security teams review code, prioritize patches and test whether a vulnerability is exploitable. OpenAI also acknowledges that giving those abilities to a malicious user could compress the time needed to attack hardened software.

Two weeks of paused training preceded the designation

OpenAI says it delayed parts of Astra’s development and release while it strengthened protections against misuse and unauthorized actions. In an August 18 update, the company said it had paused reinforcement-learning training on its latest deployment-bound models for two weeks while hardening research environments and expanding monitoring.

The company added stricter workload and network isolation, reduced standing privileges, improved security logging and required higher-security environments for Astra-related work. OpenAI says its monitoring system can inspect model activity, tool actions and available reasoning, then alert safety, security and research teams when it detects possible unauthorized access, data theft or destructive behavior.

The monitoring system is designed to raise an alert within 30 minutes of concerning activity. If teams cannot establish within another 30 minutes that a warning is a false positive, OpenAI says they are expected to pause the activity. Monitoring adds about 20% to the inference compute for workloads under observation, according to the company.

Advanced cyber features will not launch broadly

OpenAI plans to make Astra available soon but says its most advanced cybersecurity functions will initially go to a small group of testers. Access through Daybreak Blue will follow, with the program intended to support defensive security work under more controlled conditions.

The company says Astra will refuse advanced requests such as creating proof-of-concept exploits for vulnerabilities in its initial configuration. OpenAI plans to provide more permissive safeguards through Daybreak Blue for authorized vulnerability validation, malware analysis and detection engineering.

OpenAI’s disclosure leaves two questions for the release: whether its monitoring can detect misuse when Astra acts through tools rather than text, and how much defensive work will be lost when access controls block legitimate security testing. For now, the company’s position is specific: Astra is the first OpenAI model to meet its Critical cyber capability threshold, and its strongest cyber functions will remain behind restricted access.

Source

OpenAI

Explore

More articles