Computer-use agents hit 85% on OSWorld benchmark, up from 12% in two years

1 hour ago 2



Two years ago, AI agents that operate computers the way humans do, by looking at a screen and deciding where to click, were completing roughly one in eight assigned tasks correctly. Today, the best systems are clearing 85% on the same test. The benchmark in question is OSWorld, currently the most widely cited evaluation platform for multimodal computer-use agents (CUAs). These are AI systems that navigate a desktop environment using screenshots, mouse clicks, and keystrokes rather than direct API calls, essentially watching a screen and acting on what they see. The numbers behind the leap In April 2024, leading agents were scoring around 12% on OSWorld. By mid-2025 that figure had climbed into the mid-30s. By June 2026, three Anthropic models sit at the top of the OSWorld-Verified leaderboard: Claude Mythos Preview at 85.4%, Fable 5 at 85.0%, and Opus 4.8 at 83.4%. For context, the human baseline on OSWorld sits at approximately 72%. When Simular’s Agent S3 crossed that threshold in December 2025, reaching 72.6%, it was the first widely noted instance of a CUA surpassing average human performance on the test. Anthropic’s current crop has now pushed well past it. Anthropic is not al...

Read Entire Article