A consortium of Shanghai AI Lab, Tsinghua University, Peking University, Shanghai Jiao Tong University, and Chinese University of Hong Kong has open-sourced an olympiad-level reasoning model called SU-01. Based on the P1-30B-A3B architecture and fine-tuned without any external tools, the model scored 35 points on official International Mathematical Olympiad 2025 problems—hitting the gold medal threshold. It doesn't call code executors, theorem provers, or symbolic solvers; it just relies on an internal generate-verify-revise loop.
SU-01 proves that solving top-tier math problems doesn't require bolting on complex code tools. Give a 30-billion-parameter model enough test-time compute scaling (TTS), and it can chew through the toughest problems through sheer iterative deliberation.
The team trained the model using an inverse perplexity curriculum: feed samples in descending order of perplexity, forcing the model to tackle the hardest-to-imitate reasoning trajectories first. Then reinforcement learning stabilizes the answer-finding ability, and finally, relentless self-correction fills in rigorous proofs.
That mechanism is brutally compute-intensive. When facing a tough problem, SU-01 writes a first draft, spots holes, and rewrites. On a 2026 USAMO problem, the median token count for just the first complete solution is 106,000 tokens; even a single revision eats another 83,000 tokens.
Without TTS, SU-01 scores only 21 on IMO 2025. Its gold performance comes almost entirely from multiple rounds of deep search, self-checking, and revision under a high token budget.