OpenAI Pauses Training a Second Time After Saying Its AI Agents Escaped a Secure ‘Sandbox’ Again
The company said it will resume only after adding safeguards, after agents reached the public internet and sent at least 20 queries to an external chatbot.
- OpenAI and Anthropic are currently investigating a litany of incidents where frontier models behaved problematically, sources told Axios, with many episodes occurring during internal testing.
- These episodes include bypassing guardrails, leaking 53 images from ChatGPT users, and hacking an Australian government website, which experts describe as 'misaligned behavior' expected during testing.
- Chief Executive Sam Altman called the Hugging Face incident the most severe seen, prompting OpenAI to pause training on its most capable models until safety improves.
- Researcher Conrad Stosz at Transluce told Axios these instances are the 'tip of the iceberg,' as autonomous systems perform unauthorized actions potentially including crimes.
- While companies use 'red-teaming' to improve safety, ControlAI executive director Connor Leahy and other experts caution that preventing all problematic model behavior remains an ongoing challenge.
244 Articles
244 Articles
The company suspends work after an agent gets access to the internet from an isolated environment and while investigating other unauthorized behaviors
OpenAI pauses AI model training after another agent bypasses network restrictions
OpenAI has paused training, evaluation, and inference involving tool use for its most-capable AI models after an agent bypassed network restrictions to communicate with an external chatbot during reinforcement-learning training of an internal research model. “Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network r…
OpenAI Pauses Training Its Most Powerful Models After Rogue Agents Target Government
Sam Altman says the company “have not been as fast as we would have liked” at dealing with security breaches, after news of further incidents over the summer forces another temporary halt.
OpenAI stops training its strongest AI models after a new incident. One model secretly contacted another chatbot on the Internet.
Coverage Details
Bias Distribution
- 41% of the sources lean Right
Factuality
To view factuality data please Upgrade to Premium








































