Claude Sonnet 5.5 takes third place on Arena’s Agent Arena leaderboard

1 hour ago 2



Anthropic’s newest mid-tier model has landed in third place on Arena.ai’s Agent Arena leaderboard. Claude Sonnet 5.5 posted a net improvement score of 12.5%, which puts it ahead of OpenAI’s top entry. The model sits behind just two competitors, and both are also made by Anthropic. How the leaderboard shakes out Claude Sonnet 5.5 launched on September 28, 2026. Within days, it had climbed to the number three spot on Agent Arena. Its precise net improvement score is 12.52%, with a margin of plus or minus 3.09%. The two models above it are Claude Fable 5.1 (Max) at 14.31% and Claude Opus 5.5 (High) at 13.82%. Directly below sits OpenAI’s GPT-6 Astra (Max) at 12.27%. Agent Arena grades models on millions of real-world agentic tasks, where an AI has to use tools, complete multi-step jobs and satisfy actual users. The scoring looks at tool reliability, task completion and user feedback. Cost is the other headline number. Arena lists Sonnet 5.5 at $2.74 per task, while early October snapshots put its median cost at around $2.78 per task. Sonnet 5.5 also uses more tokens than most of its peers. The benchmarks beyond Arena On the Artificial Analysis Intelligence Index, it scored 56, good fo...

Read Entire Article