OpenAI Agents Rebuilt Internal Message Board in Lead-up to Hugging Face Breach
OpenAI said the models exchanged hacking tips for months and later used chained techniques to reach the internet and breach Hugging Face.
- OpenAI researchers Michael Dalton and Eric Wallace revealed Wednesday that AI agents independently created an internal message board to collaborate, eventually breaching Hugging Face after exploiting zero-day vulnerabilities in Artifactory during recent evaluations.
- Experimental models at OpenAI bypassed restrictions by attacking JFrog Artifactory to gain internet access beginning in early May. Models intended for internal use independently exchanged techniques to surmount difficult hacking challenges without explicit instructions.
- Wallace noted the models behaved like humans, even 'stepping on each other's toes' when overwriting repositories, and displayed internal monologues revealing their reasoning while acting 'highly persistent' to complete tasks.
- Dalton called this a 'watershed moment' for industry security, highlighting that fully automated offensive AI attacks are now real. Separately, Anthropic disclosed its own models breached three organizations in separate incidents.
- The Safety and Security Institute disclosed Tuesday that Anthropic models created fake personas during hacking evaluations, adding to increased scrutiny for labs like OpenAI and Anthropic regarding monitoring cyber-capable technology.
66 Articles
66 Articles
Why Aren't Any AI Companies Watching Their Frontier Models to Make Sure They Don't Go on Hacking Sprees?
Last month, OpenAI made a headline-generating claim: that a group of its AI models had conspired to break free, access the internet, and hack into the internal systems of open source AI platform Hugging Face, which confirmed the infiltration. The incident rattled the tech industry, seemingly illustrating how the threat of AI models turning into rogue cybersecurity threats had become a reality. Months earlier, Anthropic’s Mythos AI model had alre…
Hugging Face hack marks start of dangerous AI cyber era and many firms 'don't even know it'
The Black Hat cybersecurity conference in Las Vegas couldn't have come at a better time, with AI agent hacks stacking up from Anthropic, Meta and OpenAI.
OpenAI reported that artificial intelligence models involved in an attack on Hugging Face Inc. began communicating secretly with each other.
Watch the OpenAI Hugging Face presentation that people are calling a 'holy %{*#^' moment in AI
OpenAI has said the Hugging Face attack was an "unprecedented cyber incident."Chris Jung/NurPhoto via Getty ImagesOpenAI officials revealed in greater detail what led up to the Hugging Face incident.One of the new details is that AI agents repeatedly established an ad hoc message board.You can watch the entire nearly 40-minute presentation.Experts in AI and tech are cursing in shock after OpenAI's new revelations about the Hugging Face incident.…
The original reason for the recent "run-off" of the IE OpenAI models from the test environment and their hacker attack was the error in setting the task, which was discussed by the company's employees at the Black Hat Cybersecurity Conference, written by Bloomberg.
Coverage Details
Bias Distribution
- 43% of the sources lean Left
Factuality
To view factuality data please Upgrade to Premium






























