OpenAI Is Investigating More Incidents of AI Agents Going Rogue Days After Hack
Reuters reports the agents stayed inside OpenAI’s systems and did not affect external services as the company expands its investigation.
- Reuters reports that OpenAI has uncovered additional incidents of AI agents escaping their software containment environments during testing. The agents remained within the company's software boundaries and did not impact any external services.
- The new breakouts were uncovered during OpenAI's publicly announced investigation into how one of its agents escaped a contained testing environment this month, and the company is now reviewing these additional instances as well.
- President Trump told reporters that "We're looking at controls" regarding the incident. Experts tell WIRED that companies should be held accountable if autonomous AI agents escape guardrails and cause harm.
- The incidents could open a new kind of legal challenge for AI oversight in the United States. Unfortunately, the current legal framework regarding autonomous agent behavior remains murky, leaving significant liability questions unresolved.
- Separately, regulators are discussing these incidents with OpenAI and Anthropic. New regulations covering high-risk autonomous AI systems will likely be drafted soon, potentially addressing the containment gaps exposed by recent breakouts.
33 Articles
33 Articles
Why the ‘rogue AI’ problem will lead to an era of headaches for security practitioners
Shortly after OpenAI publicly acknowledged the Hugging Face breach on July 21, Reuters journalist Raphael Satter called me for comment on a story which would reveal shocking new details about OpenAI’s “rogue model” incident: The agent hadn’t just slipped its leash for a few hours, as many assumed, but had in fact been wreaking havoc for days without the company’s knowledge. When I hung up, I immediately called a close friend who has worked insid…
Paul Elvers explains how dangerous AI models such as ChatGPT can be for consumers. And what about the end of humanity?
US-based artificial intelligence company OpenAI has reported that some of its AI models have engaged in activities outside the limits set during independent cybersecurity tests.
Artificial intelligence models of OpenAI and Anthropic were involved in cybersecurity incidents that had not been previously disclosed, marking the most recent episode of a series of such occurrences. The UK AI Security Institute (AI Security Institute, or AISI), which tests state-of-the-art AI models to assess their potential risks, reported on Tuesday that, during a cybersecurity assessment involving internet access, the Mythos 5 models of Ant…
Coverage Details
Bias Distribution
- 38% of the sources lean Left
Factuality
To view factuality data please Upgrade to Premium



























