Menu

Categories

Tags

DeepMind’s AI math assistant nearly triples score on toughest benchmark

May 9, 2026 | alex | AI, Google | 158 views 0 comments

Google DeepMind just released something it's calling an "AI co-mathematician" — a multi-agent research workspace for mathematicians. On the current hardest research-level math benchmark, FrontierMath Tier 4, it scored 47.9% (23 out of 48 problems). That beats the previous best, GPT-5.5 Pro, which managed 39.6%.

What's wild? The system doesn't use a new foundation model. It's built on Gemini 3.1 Pro — the same model that powers Gemini API's File Search. That model alone scores a measly 19% on Tier 4. Add the agent framework, and the score more than doubles.

DeepMind built a layered architecture: a top-level "project coordinator" breaks research tasks into workstreams, dispatching them to sub-agents for literature search, coding, and reasoning. Written proofs then go before a committee of "reviewer agents" — no passing without their sign-off. The heavy scaffolding proves a point: in top-tier math reasoning, orchestration can squeeze out more capability gain than a model upgrade.

The blind test was run by Epoch AI. The DeepMind team never saw the problems, and each problem got 48 hours of compute. The system not only topped the leaderboard — it solved three problems that had defeated every other model.

Despite the name "co-mathematician," it acts more like a colleague who occasionally has brilliant half-baked ideas. Group theory expert Marc Lackenby used it in real research to crack an open conjecture from the Kourovka Notebook. Here's the twist: The system's initial strategy was flagged as "flawed" by its own review agent. But Lackenby spotted the clever logic buried in the rejected approach, filled in the gaps, and completed the proof.

Right now, the AI co-mathematician is only available in a closed beta to a small group of mathematicians.

Tags: #Gemini

Leave a Reply

Your email address will not be published. Required fields are marked *