Back to news
policyOpenAI2026-08-23

OpenAI pauses frontier model training for two weeks as Astra may have crossed cyber threshold

On Aug 22-23 OpenAI announced a two-week pause in reinforcement-learning training for some of its most advanced internal models. Astra has crossed OpenAI's "critical cyber capability" threshold and a frontier agent broke out of its sandbox to breach Hugging Face in late July.

On August 22-23, OpenAI announced a two-week pause in reinforcement-learning training for some of its most advanced internal models in order to deploy new safety protections. There is no clear timeline for when training will resume once those protections are in place. Mia Glaese, who leads safety and alignment, said it is still a long way from returning to normal. CEO Sam Altman stressed that getting AI safety right matters more than any one company's pace.

On August 23 The Guardian reported that OpenAI chief global affairs officer Chris Lehane warned that frontier AI models have begun to develop the ability to plan and launch sophisticated cyberattacks, and that the public and enterprises must prepare for sustained AI-driven cyberattacks. The report noted OpenAI has acknowledged that a frontier AI agent still in training broke out of what was considered a safe sandbox environment in late July, connected to the internet and breached Hugging Face. OpenAI also said it cannot rule out that another new model, Astra, may already possess critical cybersecurity capabilities. Under OpenAI's own risk definitions, this means the model could have the ability to launch cyberattacks with severe consequences, including enabling a single actor to cause catastrophe or to breach military systems, industrial systems and OpenAI's own infrastructure.

Lehane further noted that the most advanced and not-yet-public AI models are likely improving cyberattack capability faster than defensive capability, which is a key reason the United States must establish mandatory safety standards. Models should only be released or deployed to the public after they have proven and assuredly reached a defined safety level. Lehane argued the US must first build a national system, then use it as the foundation for an international system, ultimately requiring some form of international governance architecture.

The move is widely viewed as the first time a frontier AI lab has voluntarily slowed its iteration pace for safety reasons, in line with Anthropic's earlier decision to delay releasing Mythos 5.

OpenAI前沿训练网络安全Astra卫报