Google has open-sourced a new set of Gemma 4 models that use multi-token prediction (MTP) to dramatically speed up inference. These lightweight draft models, built on a speculative decoding architecture, can triple the speed of the main model without any loss in output quality or reasoning ability. The main model still retains final validation, so you're not sacrificing accuracy for speed.
Standard large language models generate one token at a time, often bottlenecked by memory bandwidth and leaving compute idle. The MTP approach lets a small draft model use that idle compute to predict multiple future tokens in one shot. Those predictions are then passed to the big model (like the 31B dense model) for parallel validation. If the draft passes muster, the whole sequence is accepted at once. To boost efficiency further, the draft model shares activation states and KV cache with the main model — and for on-device variants (E2B and E4B), the team introduced clustering in the embedding layer.
The MTP models are now fully open source under the Apache 2.0 license, just like the rest of the Gemma 4 family. They work out of the box with popular inference frameworks such as vLLM, SGLang, and Ollama. This speedup lowers the barrier to entry: developers can run the 26B Mixture-of-Experts and 31B dense models smoothly on consumer-grade GPUs, and mobile devices can handle real-time AI interactions with lower power draw.