New tool easily jailbreaks safeguards of frontier AI models, raising alarm for crypto and tech sectors

1 week ago 14



The safety guardrails on the world’s most powerful AI models are, to put it gently, not holding up. A Nature study published in July 2026 found that four large reasoning models, deployed as adversarial attackers, achieved a 97.14% overall jailbreak success rate against nine frontier AI models from companies including OpenAI, Anthropic, and Google. How the guardrails crumbled The research tested four large reasoning models as autonomous adversaries: DeepSeek-R1, Gemini 2.5 Flash, Grok 3 Mini, and Qwen3 235B. These weren’t sophisticated nation-state tools. They used simple prompts and strategic persuasion techniques in multi-turn conversations to bypass safety filters. DeepSeek-R1 was the standout performer, if you can call it that. It recorded a 100% attack success rate on HarmBench prompts in Cisco-linked testing. Every single prompt got through. The targets, models from OpenAI, Google, and Anthropic, showed only partial resistance at best. Separate research focusing specifically on Claude models found that even advanced jailbreak techniques caused only a 7.7% performance degradation in the highest-performing variants. Security firm HiddenLayer has documented universal bypass techn...

Read Entire Article