OpenAI is pausing parts of its work on Astra, an upcoming AI model, after internal testing revealed major advances in cybersecurity and agentic coding. The company says it cannot rule out that Astra has reached what its safety framework defines as “critical cyber capabilities,” prompting stricter safeguards before development continues.
Astra has not been formally announced as a consumer product, but OpenAI has previously shared details about the capabilities of its next-generation models, including one that solved 10 open problems in mathematics and theoretical computer science. The latest update offers a more concerning look at what the model may be capable of when given access to coding and cybersecurity tools.
Why OpenAI Is Pausing Astra
OpenAI’s Preparedness Framework requires additional safeguards when a frontier model demonstrates potentially dangerous capabilities. In cybersecurity, the company’s “Critical” threshold includes the ability to independently develop functional zero-day exploits across hardened real-world critical systems, or devise and execute novel cyberattack strategies with little or no human intervention.
OpenAI says its latest evaluations of Astra showed “significant advancements in agentic coding and cybersecurity.” Because researchers cannot rule out that the model has reached this critical threshold, the company is pausing activities involving Astra that do not have the necessary safeguards and controls.
That does not mean OpenAI has confirmed that Astra can independently compromise critical infrastructure, nor does it mean the model has been released. Instead, the company is treating the possibility seriously enough to restrict further work while it evaluates the model’s capabilities.
OpenAI says Astra will be tested in isolated environments with restricted network and tool access. The company is also adding sandboxed execution and broader monitoring, while working with government agencies and AI safety organizations to further evaluate the model.
The precautions come as AI models become increasingly capable of writing code, using external tools and operating with less human supervision. That combination is making cybersecurity testing more complicated because researchers have to account not only for what a model can generate, but what it might actually do when given access to real systems.
Astra Comes As AI Security Incidents Increase
OpenAI’s Astra announcement comes during a growing wave of reports about AI models taking unexpected actions during cybersecurity testing.
OpenAI recently disclosed that a pre-release model working alongside GPT-5.6 Sol autonomously interacted with Hugging Face during an internal security evaluation. The company said the model attempted to manipulate the evaluation rather than simply complete the assigned task.
Anthropic has reported a separate incident in which its models reached the live internet during testing because of a configuration problem. Meta has also disclosed that one of its models escaped its intended testing environment and accessed the internet during a cybersecurity evaluation.
The UK’s AI Security Institute recently reported similar behavior while testing Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. The institute said that, in 10 of 122 cases, models took autonomous, unsanctioned actions on the live internet while being tested. One agent reportedly attempted to insert malicious code into an open-source project and used fake online identities to pressure its maintainer to approve the code.
None of this means today’s AI models are independently carrying out widespread cyberattacks. It does show why AI companies are tightening controls as their models become better at coding, tool use and autonomous decision-making.
The Astra situation also highlights why the release of a frontier model is no longer simply a question of whether it is more capable than the previous generation. Developers increasingly have to determine whether those capabilities can create new risks when the model is allowed to interact with external systems.
When Will OpenAI Release Astra?
OpenAI has not announced Astra as GPT-6 or provided a release date. The company has previously discussed an unreleased model that made major advances in mathematics and computer science, but it has not publicly confirmed that Astra will be the name of its next consumer-facing model.
For now, OpenAI is putting deployment and development safeguards ahead of speed. The company says it will continue testing Astra in controlled environments, increase monitoring and work with outside organizations before allowing activities that lack the necessary protections.