Cybersecurity

OpenAI Reveals Astra as Its First Model to Reach the “Critical Cybersecurity Capability” Threshold

OpenAI said that the Astra model can discover previously unknown vulnerabilities and develop methods to exploit them, and that it achieved a perfect score on the ExploitBench test. The company will gradually roll out its cybersecurity capabilities to test users after delaying part of the development process to add safety controls, with no independent verification so far.

2026-09-02
4 min read
6 views
certi.news
OpenAI Reveals Astra as Its First Model to Reach the “Critical Cybersecurity Capability” Threshold

OpenAI revealed new details about the Astra artificial intelligence model, which it is preparing to make available soon, saying that it is the first of its models to reach the “critical cybersecurity capability” level under the Preparedness Framework. According to the company’s description, this means that, when provided with the appropriate tools and access rights, the model can discover previously unknown security vulnerabilities and develop ways to exploit them.

The development is not limited to the model’s ability to analyze code or suggest fixes; it also involves its ability to execute advanced stages of the exploitation chain. For this reason, OpenAI delayed part of Astra’s development and launch while it worked in recent weeks to reduce the likelihood of its use in harmful cyberattacks or unauthorized operations.

Strong Results in Exploitation Tests

OpenAI said that Astra achieved a score of 100% on the ExploitBench test, which measures the ability to develop exploit software for known vulnerabilities. The company also conducted an internal evaluation covering 20 high-severity vulnerabilities disclosed during the period from June to August 2026.

According to the results reported by OpenAI, Astra achieved higher rates of remote code execution in this evaluation than the GPT-5.6 Sol model, while using fewer token units. During the test, the company said that the model discovered and used two “zero-day” vulnerabilities as part of an exploitation chain, and that it had begun notifying the software developers concerned about these vulnerabilities.

Other tests conducted by experts examined Astra’s ability to exploit unknown vulnerabilities in a hardened browser, then escape the sandbox and execute commands on the host system. In operating-system tests, the model, according to OpenAI, was able to chain several vulnerabilities to move from an unauthorized user to root privileges.

Limited Availability and Additional Controls

OpenAI plans to make Astra’s most advanced cybersecurity capabilities available in stages. The first stage will be limited to a group of users participating in testing, before access is expanded to defensive uses through the Daybreak Blue program.

The company says it has increased Astra’s refusal rate for harmful cybersecurity requests and will impose stricter restrictions on accounts it considers high-risk. It will also add monitoring mechanisms to detect unauthorized behavior, along with layers of reasoning and oversight for actions performed by the model.

Why Does This News Matter?

What is changing in practical terms is the transition of a general-purpose model to a level at which it can, according to the company’s tests, contribute to discovering and exploiting vulnerabilities through a multistep chain. This could benefit defense teams and security researchers if use remains confined to authorized environments, but it simultaneously increases the importance of access control and monitoring, because the same capabilities could reduce the effort required to develop complex attacks.

The limits of the available information are particularly important here. The figures and results cited come from OpenAI, and there has so far been no independent verification of Astra’s security level or of the measures intended to limit its misuse. The full system and detailed evaluation results will not be published, according to the article, until the official launch. Therefore, practical assessments of the model’s capabilities and risks will depend on the published security system and the results of independent tests after availability, rather than on initial performance claims alone.

News source
c
Author

certi.news

In the same category

You may also like

View all news