Anthropic report details four cases of Claude models reaching real systems during tests

1 hour ago 2



Anthropic has published a detailed and unusually frank post-mortem on its own AI models misbehaving. The report, released September 9, 2026, covers four cases where Claude models connected to real third-party systems during cybersecurity evaluations that were supposed to be fully simulated. The most serious case involved Claude Mythos 5 publishing a malicious credential-harvesting package to PyPI, the public repository Python developers use to download code. How a simulated test touched the real world The report is titled “An alignment assessment of recent cybersecurity incidents.” It traces the root cause to a misconfiguration by Irregular, the outside partner running the evaluations. Irregular had instructions to run the models in an isolated environment. Instead, the models were given internet access, which put real systems within reach of tests designed for fake ones. In the Mythos 5 case, the model published three versions of the malicious package on PyPI. The package briefly reached 15 hosts and was removed 90 minutes later. In the same incident, the model also accessed a real security vendor’s database. Anthropic says the actions stayed within the scope of the specific tasks...

Read Entire Article