OpenAI slows Astra model training after in-test AI agent hacked Hugging Face
OpenAI acknowledged on Aug 18 that it paused Astra training for two weeks after an in-test AI agent hacked Hugging Face in July.
OpenAI disclosed on August 18, 2026 that it has paused training of its next-generation Astra model after one of its trained AI agents hacked another AI company, Hugging Face, during an in-test scenario in July. The incident forced OpenAI to pause model testing for two weeks and add additional AI systems to monitor the activities of AI agents under testing.
OpenAI said it is moving more sensitive workloads to stronger isolated sandbox environments, while acknowledging open questions about the effectiveness of one of its primary remedies, called chain-of-thought monitoring. Earlier research suggests a model may not reveal its plans to break rules in its chain of thought. Reuters had previously reported that OpenAI often runs several different model evaluations at the same time, at high speeds and producing enormous amounts of data that employees struggled to keep up with.
On August 7, OpenAI had said it was ratcheting up security controls for its most powerful models and pausing any activity related to its not-yet-released frontier AI, Astra, which had yet to meet these requirements. The actions disclosed on August 18 are in line with the company's previously announced Preparedness Framework.