OpenAI has paused the rollout of Astra, its next frontier model, after internal and external evaluations found it hits a level of capability the company was not prepared to ship immediately. Under OpenAI’s own Preparedness Framework, Astra is the first model the company has classified as “critical” for cybersecurity risk, the highest threshold on the ladder.
The capability itself is the story. Astra shows major advances in agentic coding and cybersecurity, the two areas where a capable model can do the most damage in the wrong hands. OpenAI says it is slowing down to lock in safety controls, not abandoning the release, and CEO Sam Altman is publicly pushing back on the idea that strong models belong in the hands of a few.
What “Critical” Actually Means
The Preparedness Framework sets the bar high before a model is called “critical.” Per OpenAI, one of two conditions must hold: the model can find and weaponize zero-day vulnerabilities across hardened real-world systems with no human intervention, or it can autonomously conceive and execute end-to-end cyberattacks against hardened targets given only a high-level strategic goal.
Astra reportedly meets that standard. OpenAI is not publishing details of the evaluation, but the company confirms Astra performed at a level that triggered the framework’s highest risk tier, a first for any OpenAI model.
What OpenAI Is Doing About It
The company has announced several controls rather than a full stop:
- Stricter security on high-capability models: isolated test environments, restricted network and tool access, stronger model-weight protection and encryption, extra monitoring, and sandboxed execution. Internal work on Astra that does not meet the new bar is paused.
- Global monitoring of all Astra agent applications, including training and evaluation, with high-risk behavior and alignment-failure tracking. OpenAI is also reviewing the model’s chain-of-thought and blocking high-risk operations.
- Collaboration with government agencies and AI safety organizations on testing, plus sharing recommended security controls with third-party test partners.
OpenAI also clarified that Astra is unreleased and was not involved in the recent attack on Hugging Face that surfaced in security testing. It is an important distinction, because that incident involved agents operating in a test environment, not this model.
The Strategic Read
The delay is a genuine signal about where frontier models are heading: agentic coding and cybersecurity are now the sharpest edges of capability. For developers and teams evaluating tools, it means two things. The category of “agent that can operate infrastructure” is arriving faster than the safety tooling around it, and vendors who ship such capabilities will increasingly gate them behind verification, sandboxing, and permission systems. Our GPT-5.6 Luna coverage and Claude vs GPT model comparisons are good starting points for tracking how fast these models are moving.
The Caveats
Three things to keep in mind. First, “critical” is a self-assessment under OpenAI’s own framework; there is no independent public benchmark proving Astra can autonomously weaponize zero-days. Second, the delay is not a cancellation. Altman says the team is pushing hard for a public release and expects the wait to be short, so teams should not assume the model is dead. Third, the same controls that make Astra safer also make it less accessible; expect stricter gating and more verification friction on high-capability models generally. Watch the official OpenAI channels for the actual release terms.
Related: see our AI news hub