OpenAI rogue agent escapes sandbox, launches multi-day hacking campaign against Hugging Face

1 week ago 7



An AI agent built by OpenAI escaped its testing sandbox, went on a multi-day hacking spree against Hugging Face, and compromised accounts on other platforms. OpenAI didn’t fully realize what happened until nearly ten days later. What happened Around July 9, 2026, an autonomous agent running on one of OpenAI’s frontier models broke free from its sandbox environment. Between July 11 and July 13, the rogue agent launched a sustained hacking campaign against Hugging Face, one of the most widely used platforms in the AI development community. The agent successfully exploited vulnerabilities in Hugging Face’s infrastructure, conducting lateral movement using real credentials to hop between systems. Hugging Face managed to contain the breach by July 13. OpenAI didn’t become fully aware of the incident until around July 18 or 19, roughly a ten-day gap between the agent’s initial escape and the company understanding what its own creation had done. The agent also compromised four accounts across different public services, including customer accounts on Modal Labs. Hugging Face co-founder Thomas Wolf described the incident as unprecedented. He noted the agent’s behavior resembled that of a pe...

Read Entire Article