Menu

Categories

Tags

Google's zero-copy MTP speeds up Pixel's Gemini Nano by 50% and saves 130MB

June 28, 2026 | Source: research | AI, Google | 278 views 0 comments

Google has quietly rolled out a multi-token prediction (MTP) architecture on its Pixel 9 and Pixel 10 series, giving the built-in Gemini Nano v3 model a major speed boost. By attaching lightweight Transformer prediction heads to the frozen main model, the new setup boosts inference speed by over 50% — while preserving safety alignment and output quality.

Traditional speculative decoding requires running a separate draft model to predict candidate tokens. That eats up memory and accuracy, since the draft model can't access the main model's hidden states. Google's MTP heads reuse the main model's already-computed feature activations, dramatically improving prediction accuracy.

To avoid redundant memory overhead from draft calculations, Google designed a zero-copy mechanism. Normally, the draft model maintains its own key-value (KV) cache memory. With zero-copy, the prediction heads read the main model's existing cache directly via cross-attention. This eliminates startup latency and saves about 130MB of RAM on the phone.

In real-world Pixel tasks like notification summaries and text proofreading, MTP allows the model to successfully predict nearly two extra tokens per inference run, reducing how often the main processor wakes up for validation — saving power. For highly structured text generation like smart replies, token acceptance rates jumped by 55%.

Tags: #Gemini

Leave a Reply

Your email address will not be published. Required fields are marked *