Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job

1 hour ago 2



Epoch AI spends much of its time measuring how capable AI models are becoming. With its latest release, the organization turned the question inward: could these models actually do Epoch’s own work? According to the new Epoch Automation Reports, not yet. Frontier models get close on clearly defined tasks, but they still fall short of fully automating the work, especially when assignments turn open-ended or call for judgment. A job trial, not a pop quiz Epoch AI introduced the Automation Reports on October 8, 2026. Instead of relying on abstract puzzles, the benchmark pulls its tasks from the organization’s internal research. The evaluation covers 11 distinct tasks spread across five categories. Human graders score each model’s output against the same internal quality rubrics Epoch uses for its own work. The scoreboard Two models share the top spot. Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra each posted an average score of approximately 65% across the tasks. Grok 4.6 followed at 59%. Qwen 3.8 Max came in at 53%, edging out Kimi K3 at 52%. Gemini 3.8 Flash rounded out the listed results at 42%. That leaves a spread of more than 20 points between the leaders and the bottom o...

Read Entire Article