Epoch classifies first AI solution as major advance in FrontierMath benchmark

1 day ago 4



When Epoch AI launched Tier 4 of its FrontierMath benchmark in July 2025, the best AI models on the planet could solve roughly 5% of its problems. Fourteen months later, OpenAI’s GPT-6 Astra has cracked nearly all of them, earning the first-ever “Major Advance” classification from Epoch for an AI-generated mathematical solution. The model achieved between 97.6% and 98% accuracy across all 43 problems in Tier 4, the benchmark’s most difficult category. The problem that was supposed to be unsolvable The last Tier 4 problem to fall was authored by combinatorialist Jay Pantone. It was specifically designed to resist the kinds of shortcuts that earlier AI models had exploited to inflate their scores on other FrontierMath problems. GPT-6 Astra solved it anyway, which is precisely why Epoch elevated the result to “Major Advance” status. The classification isn’t just about getting the right answer. It reflects Epoch’s assessment that the solution demonstrates genuinely novel mathematical reasoning capability rather than pattern-matching tricks. This is the first time any AI system has received this designation from Epoch’s FrontierMath program. What FrontierMath actually measures FrontierM...

Read Entire Article