Mistral AI has open-sourced Leanstral 1.5, a model optimized for formal theorem proving in Lean 4. The model packs 119 billion total parameters, with about 6.5 billion activated per inference, and is released under the Apache-2.0 license with a free API tier.
According to official benchmarks, Leanstral 1.5 solves 587 out of 672 problems in PutnamBench, a notoriously difficult math competition dataset. On the abstract algebra benchmarks FATE-H and FATE-X, it scores 87% and 34% respectively, setting new records for models in its class.
The real kicker? The average cost per PutnamBench problem is around $4 — a fraction of the tens or hundreds of dollars some earlier systems required. As the per-problem token budget increases, the model's success rate keeps climbing. In one extreme case, proving the complexity of an AVL tree took over 2.7 million tokens and 22 rounds of context compression, but it eventually nailed the proof.
Beyond math, Mistral's team put Leanstral 1.5 to work on code verification. Scanning 57 open-source Rust repositories turned up 11 real bugs — 5 of which had never been reported before.