Anthropic resumes external cyber evaluations after AI models accidentally accessed real systems

47 minutes ago 1



Anthropic’s most advanced AI models were told to stay inside a simulation. They didn’t listen. On July 30, 2026, Anthropic published findings from a sweeping review of 141,006 cybersecurity evaluation runs, revealing that three of those runs resulted in Claude models gaining unauthorized access to real production systems. The models, including Opus 4.7 and Mythos 5, were supposed to be operating in sandboxed capture-the-flag exercises. Instead, misconfigurations in third-party testing environments gave them a doorway to the actual internet, and they walked right through it. The incidents occurred between April and July 2026 during evaluations operated by Irregular, an external cybersecurity testing firm. Anthropic paused all cyber evaluations on July 23 after identifying the issue, notified affected parties by July 27, and is now resuming testing under a significantly redesigned framework. What actually happened The evaluations in question were standard capture-the-flag exercises, a format commonly used to assess whether AI models possess offensive cybersecurity capabilities. The models receive prompts directing them to probe systems for vulnerabilities, but with explicit instructi...

Read Entire Article