Grok 4.7 ranks first on the Artificial Analysis Cyber Index

2 hours ago 2



Grok 4.7 now sits at the top of the Artificial Analysis Cyber Index, a new leaderboard built to measure how well AI agents defend software. It posted a composite score of 56. It shares the top spot with MiMo-V2.6-Pro, so this is less a coronation and more a co-headlining tour. The timing matters. The index launched on or around September 28, 2026, which makes Grok 4.7 one of the first models to claim a top position on a benchmark aimed squarely at enterprise security work. How the scoreboard shakes out The ranking is tight at the top. Grok 4.7 and MiMo-V2.6-Pro each scored 56 to tie for first place. GPT-6 Luna landed in third with a score of 53. The Cyber Index tests AI agents on a set of enterprise cyber defense tasks. Models must find vulnerabilities, reproduce them, and then ship working patches inside real codebases. Grok 4.7’s composite score rests on strong results in specific sub-benchmarks. On CWE-Bench-AA, it recorded a 68% pass@1 rate, which tied for the lead on that test. Pass@1 measures whether a model gets the task right on its first attempt. No retries, no do-overs, which is roughly how a security team would want an automated tool to behave. On the CyberGym-E2E-AA pat...

Read Entire Article