Menu

Categories

Tags

Stop Throwing Compute at AI: Models Get Dumber the More You Train Them

June 28, 2026 | Source: zyphra | AI | 102 views 0 comments

As AI models train longer, they gradually lose the ability to absorb new knowledge — a phenomenon called plasticity loss — and end up getting dumber with more training. If we can't crack plasticity loss, large models will never be able to learn continuously at low cost; every update would require retraining on all historical data plus new data, burning through massive amounts of compute.

Research from AI startup Zyphra shows for the first time that scaling up models delays the degradation, but with diminishing returns. Throwing more parameters at the problem alone can't fix plasticity loss. Extrapolations indicate that a 1-billion-parameter model goes stupid after training on 1.8 trillion tokens, while a 7-billion-parameter model shows signs after 9 trillion. Even more striking: plasticity loss still occurs when models train on a stable, mixed dataset without ever switching tasks.

The study identifies three direct causes: parameter volumes grow during training, which under LayerNorm hinders gradient flow; neurons in MLP layers go dormant on a massive scale (up to 95% in some models); and attention heads either become paralyzed (fixating on a single token) or lazy (smoothing attention uniformly across all context). Potential fixes include capping parameter growth, periodically force-activating dormant neurons (a "neural reset"), and injecting random noise into the attention mechanism to break bad patterns.

Leave a Reply

Your email address will not be published. Required fields are marked *