Menu

Categories

Tags

Google's zero-copy MTP boosts Gemini Nano speeds over 50% on Pixel

June 29, 2026 | Source: research | AI, Google | 362 views 0 comments

Google has quietly deployed a multi-token prediction (MTP) architecture on its Pixel 9 and Pixel 10 devices, giving the on-device Gemini Nano v3 model a serious speed bump. By tacking lightweight Transformer prediction heads onto a frozen main model, the company says it’s seeing inference speeds increase by over 50% — while keeping safety alignment and output quality intact.

Traditional speculative decoding requires a separate draft model to predict candidate tokens. That eats up precious RAM on a phone, and because the draft model can’t peek inside the main model’s hidden states, accuracy suffers. Google’s approach embeds MTP heads at the tail end of a frozen main model, reusing already-computed feature activations to boost candidate token prediction accuracy.

To avoid the redundant memory overhead of draft computation during autoregressive generation, Google designed a zero-copy mechanism. In conventional setups, the draft model must maintain its own key-value (KV) cache. The zero-copy trick lets the external prediction heads read directly from the main model’s existing cache via cross-attention. That not only eliminates startup latency for draft prediction but also frees up about 130MB of RAM on the phone.

In real-world Pixel tasks like notification summaries and text proofreading, the MTP architecture allows the model to successfully predict nearly two extra tokens per inference on average, reducing how often the main processor needs to wake up for verification. That saves system power. On highly structured text generation tasks like smart replies, the token acceptance rate jumped by 55%.

Tags: #Gemini

Leave a Reply

Your email address will not be published. Required fields are marked *