Nvidia finds a way to reuse AI caches across different models

For the past two years, the biggest optimization target in LLM inference has been KV Cache. It's the intermediate computation a model leaves behind after reading context — and the longer the context, the more VRAM and bandwidth it devours. DeepSeek-V2's MLA approach cut KV Cache by 93.3% compared to DeepSeek 67B, and Kimi Linear trims it by another 75% on top of full MLA. Everyone in the industry is trying to make long context cheaper.
But these approaches all hit a wall: the cache can usually only be reused inside the same model. When an AI agent switches from a small model to a big one, the target model has to recompute the first tens or even hundreds of thousands of tokens from scratch. The more often you route between models, the harder it is to ignore that repeated computation.
Nvidia's latest paper starts chipping away at that wall. The team found that across differently sized models from the same family with matching KV structures, the caches have a clear linear relationship. Calibrate once with 500 passages of 1,024 tokens each, and you can fit a mapping that transfers a KV Cache computed by one model to another model so it can just keep going.
When switching from Qwen3-14B to Qwen3-32B, recomputing a 32K-token context takes about 7 seconds, while converting the cache takes only 0.28 seconds — roughly 25x faster. In other words, you could soon let a small model do the "reading" by turning the long context into a KV Cache, then hand that cache to a larger model when deeper reasoning is needed, letting the big model jump straight to "thinking" and generation.
That could make model routing significantly cheaper. Right now, running simple tasks on a small model mainly saves generation cost; when you do switch to a big model, it typically has to recompute the long context anyway. If cross-model KV Cache matures, even that prefill work can be done by the small model. An agent could maintain context on a cheap model indefinitely, only waking the big model when it hits a hard problem — no more reloading from the beginning every time.