Back to news
regulationOpenAI2026-08-08

OpenAI pauses Astra model work, says it may hit Critical cybersecurity threshold

OpenAI said on August 7 that internal evaluations show its upcoming Astra model may have reached the Critical cybersecurity capability level, triggering safety measures.

On August 7, OpenAI published a safety notice stating that internal evaluations of its upcoming Astra model show significant advancements in agentic coding and cybersecurity capabilities. The company said it cannot rule out the possibility that the model has reached the Critical cybersecurity capability level defined in its Preparedness Framework. This is the first time OpenAI has attached that label to a specific model.

Under the Preparedness Framework, the Critical threshold means a model can autonomously identify and develop functional zero-day exploits against hardened real-world critical systems without human intervention, or devise and execute end-to-end novel cyberattack strategies given only a high-level goal. Prior models including GPT-5.6 Sol were assessed only at the High threshold.

OpenAI clarified that Astra was not involved in the July 21 Hugging Face security incident. The company has activated stricter security controls including isolated testing environments, restricted network and tool access, enhanced model weight protection and encryption, additional monitoring and detection capabilities, and real-time monitoring of high-risk behavior. Internal Astra activities that do not yet meet the strengthened controls have been paused, and OpenAI will work with government agencies and AI safety organizations to test the model.

CEO Sam Altman said on X that Astra is a powerful model and the company is working to make it generally available, but given its cyber capabilities it needs a little bit longer to do this safely. This is the first time a frontier AI lab has publicly committed to slowing its own model development for cybersecurity reasons.

OpenAIAstra安全网络安全关键级