Muon's Hidden Bug Starves 25% of Neurons; Aurora Fix Boosts Efficiency 100x
Tilde Research found that Muon, the optimizer behind top models like DeepSeek V4, Kimi K2.5, and GLM-5, has a nasty hidden defect: it quietly kills more than a quarter of neurons in the MLP layers during early training. The team designed a replacement optimizer called Aurora and open-sourced it. A 1.1B parameter model trained on roughly 100 billion tokens matched Qwen3-1.7B (which used 36 trillion tokens) on language understanding benchmarks like HellaSwag and Winogrande.
The problem lies in a mathematical quirk in how Muon processes MLP weight matrices. Early on, some neurons happen to receive weaker gradient signals. Traditional optimizers like AdamW normalize per parameter, naturally smoothing out the difference. But Muon's orthogonalization step passes those weak signals through unchanged. Weak neurons get weaker updates, fall further behind, and enter a "rich get richer" death spiral. By step 500, over a quarter of neurons are effectively dead, wasting parameter capacity.
Previous attempts to fix this, like NorMuon, forced uniform row updates but broke orthogonality—the very property that makes Muon efficient. Aurora treats "update uniformity" and "orthogonality" as joint constraints, satisfying both with an alternating iteration. Every neuron gets a fair learning opportunity without sacrificing update precision.
Out of the box, Aurora adds only 6% overhead over Muon and can be swapped in directly. On the modded-nanoGPT optimization benchmark, Aurora set a new record at 3,175 steps. The advantage scales with MLP width: the wider the layer, the bigger the improvement.
Code and the 1.1B pretrained model are now open source.