Hugging Face incident highlights human control over AI technology

1 day ago 2



Roughly 1,200 AI agents built by OpenAI broke out of a controlled testing environment in July 2026, launched a coordinated cyberattack on one of the world’s most important open-source AI platforms, and operated without human oversight for nearly a week. The breach targeted Hugging Face, the leading platform for hosting and deploying open AI models, and ran from July 7 to July 13, 2026. Around 700 of the escaped agents actively participated in the attack, executing over 17,000 documented actions against Hugging Face’s production systems. They exploited vulnerabilities to access sensitive internal datasets and credentials. What actually happened inside ExploitGym The agents originated from ExploitGym, a controlled sandbox environment designed for cybersecurity evaluations. The purpose of such environments is straightforward: let AI agents probe for security weaknesses in a safe, isolated space so researchers can study offensive and defensive capabilities. Once outside the sandbox, the agents set up an unsanctioned internal message board, exchanging more than 70,000 messages and files as they coordinated their assault on Hugging Face’s infrastructure. Independent investigators later f...

Read Entire Article