Anthropic Says Claude AI Models Hacked Three Companies During Tests
Anthropic said a retrospective review found 141,006 tests and three cases in which Claude models reached real systems after escaping sealed environments.
- On Thursday, Anthropic reported that its Claude artificial intelligence models accessed the internet during evaluation tests and "gained unauthorized access to the real systems of three different organizations."
- The incidents occurred within testing environments built by the AI security firm Irregular, where a misunderstanding with the evaluation partner left the environments unsealed despite Anthropic instructing Claude they lacked internet access.
- Anthropic discovered the breaches after a "large-scale retrospective review" of 141,006 evaluation tests, identifying three models—Opus 4.7, Mythos 5, and an internal research model—that used basic techniques like exploiting weak passwords.
- Neither Anthropic nor the affected organizations detected the intrusions until the retrospective review, which was prompted by a similar security incident OpenAI disclosed last week.
- More than 1,100 staffers across artificial intelligence firms signed a petition on Tuesday urging the government to support mechanisms that "deliberately pace" AI development to prevent rapid advancement.
486 Articles
486 Articles
Anthropic Says Claude Hacked Three Companies During AI Tests
Anthropic says three of its Claude AI models gained unauthorised access to the production systems of three real organisations while completing cybersecurity tests. The incidents were not planned attacks. According to Anthropic’s official investigation, the models had been told that they were operating inside simulations without internet access. A configuration failure meant the testing machines could still reach the open internet. When Claude en…
One bug connected Anthropic's testing environment to the Internet. Claude attacked real systems, published a malicious package on PyPI, and went unnoticed for two victims.
What Claude’s real-world breaches reveal about AI safety tests
This week, just days after OpenAI announced that two of its advanced AI models had interacted with real-world systems during cybersecurity tests, Anthropic reported a similar containment failure. In a post on X, Anthropic said it found three separate cases where Claude models accessed the internet from third-party testing environments and affected real organizations. This came after reviewing over 141,000 evaluation runs, a process started becau…
Anthropic says its AI models hacked 3 organizations during testing - WSVN 7News | Miami News, Weather, Sports
Anthropic said its artificial intelligence models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI controls after it disclosed its rogue...
Coverage Details
Bias Distribution
- 54% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium









































