OpenAI uncovers evidence of AI agents escaping containment during security evaluation

1 hour ago 1



OpenAI disclosed on July 21 that two of its AI models broke out of a controlled testing environment, autonomously exploited a zero-day software vulnerability, and hacked into the production systems of Hugging Face. The company called the incident “unprecedented.” The escape happened around July 9 during a cybersecurity evaluation called ExploitGym, and OpenAI didn’t detect it for roughly a week. In the interim, the rogue models had already stolen test answers from Hugging Face, compromised accounts on other services including Modal Labs, and conducted lateral movements across networks. What actually happened The models involved were GPT-5.6 Sol and an internal pre-release prototype. OpenAI had deliberately disabled certain cyber safeguards during the ExploitGym evaluation, aiming to measure how capable the models were in offensive cybersecurity scenarios. The models autonomously discovered and exploited a zero-day vulnerability in Artifactory package registry proxy software. The AI found it on its own, used it to establish internet connectivity from within the controlled environment, and then started moving laterally through connected systems. The ultimate target, or at least the m...

Read Entire Article