Epoch AI confirmed today that FrontierMath — dubbed the toughest math benchmark around — has a massive problem. An AI-assisted review found that roughly one-third of the problems contain "fatal errors," and the team believes most of the bug reports are valid.
FrontierMath Tiers 1-4 is a collection of hundreds of unpublished problems maintained by Epoch AI. Tiers 1-3 cover undergrad to early postdoc difficulty, while Tier 4 is research-level math. When the benchmark debuted, models like GPT-4o, Claude 3.5 Sonnet, and o1-preview couldn't solve more than 2% of the problems, making it a go-to yardstick for how far AI is from advanced mathematical research.
This mess undermines the interpretability of past FrontierMath scores. If a problem is unsolvable, has a wrong answer, or is missing conditions, a model's inability to solve it isn't a fair measure of its capability. Even more awkward: the problems were caught by AI-assisted review — as models' reasoning improves, they're starting to correct the human experts who wrote the benchmark.
It's too early to say that model scores will suddenly jump. Epoch has only confirmed that it will re-release revised scores for a corrected dataset. Exactly how much scores will rise, how rankings will change, and which problems will be removed all depend on a pending manual review.