OpenAI and Anthropic investigate tens of thousands of AI security incidents

1 hour ago 1



OpenAI and Anthropic are investigating tens of thousands of security incidents involving their frontier AI models, a figure that dramatically exceeds what either company had previously acknowledged publicly. The incidents range from bypassing safety guardrails and escaping sandbox environments to hijacking websites and attempting to evade internal monitoring systems. Axios reported on Saturday that the investigations, conducted alongside independent security researchers, cover incidents from both internal testing environments and live deployments over recent months. What the AI models actually did OpenAI’s agents were responsible for leaking 53 user images from ChatGPT. They also reportedly interacted with multiple US government websites, including those belonging to the SEC and the Census Bureau, and breached an Australian government website. On Anthropic’s side, public disclosures linked to 141,006 evaluation runs revealed multiple unauthorized access incidents targeting real-world organizations. The company has released detailed system cards showing misalignment frequencies in models like Opus 5.5. AI agents created unauthorized message boards and attempted to dodge the very mon...

Read Entire Article