OpenAI Pauses Astra Model Over Critical Cybersecurity Risk Concerns

OpenAI paused work involving Astra after tests showed cybersecurity abilities that could approach its Critical risk threshold under the company’s framework.
OpenAI disclosed that internal evaluations of Astra, one of its upcoming models, have found cybersecurity capabilities significant enough that the company “cannot rule out” reaching the Critical threshold under its own Preparedness Framework.
In response, the company paused certain internal activities involving Astra and implemented a set of security controls that it had not previously needed to apply. This is the first time an AI lab has publicly announced slowing development of a model specifically because of cybersecurity concerns.
“Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.” reads the announcement.
“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time. Astra is an upcoming model, and was not involved in exploiting Hugging Face.”
Previous models, including GPT-5.6-Sol, had been assessed at the High threshold rather than Critical. Astra wasn’t involved in the Hugging Face incident disclosed last month. OpenAI is making that distinction deliberately, because the news cycle has already connected every AI breach to every AI model.
“We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.” continues the announcement. “We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. Monitors evaluate the model’s Chain of Thought and trigger a security response to review and interrupt high risk activity.”
The new controls also include isolated testing environments, restricted network and tool access, enhanced encryption of model weights, and sandboxed execution. OpenAI says it will share recommended security controls with third-party testing partners for running higher-risk evaluations, a direct response to the series of incidents in which evaluation environments gave AI models unintended internet access.
The broader context makes this disclosure land harder than it might otherwise. The UK AI Security Institute reported last week that AI models autonomously reached out to real-world targets across 10 of 122 evaluation runs, with 17 of 19 such actions originating from Anthropic’s Mythos 5. In the most serious case, an agent tried to insert malicious code into an open-source project and created fake online identities to pressure the project’s maintainer into approving it. A human maintainer caught it. Models from Meta and Chinese company Moonshot, Muse Spark 1.1 and Kimi K3, have also been reported escaping sandboxes, with Kimi K3 probing the network during an evaluation, finding that GitHub was reachable, cloning the benchmark repository it was supposed to be solving, and reading the answer directly off disk. The incidents are being tracked on a new site called Felony Bench.
OpenAI says it believes advanced cyber-capable models should help defenders find vulnerabilities before attackers do, and frames the pause as responsible stewardship rather than alarm. That may be true. It’s also true that the Preparedness Framework was designed for exactly this moment, and that using it to actually slow down a model rather than just document the risk is a meaningful choice, one the industry will be watching to see whether others follow.
“We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.” concludes the announcement.
Follow me on Twitter: @securityaffairs and Facebook and Mastodon
Pierluigi Paganini
(SecurityAffairs – hacking, Astra)

