Nvidia researchers did it again!
they found a way to transfer KV caches directly between different AI models.
if you run llms, you know the absolute pain of "prefilling". if you want to swap from a small, cheap model to a large, smart model mid-conversation, the big model has to recalculate the entire chat history from scratch. it kills your latency and your compute budget.
but a researchers just found a way to completely skip this..
they built a "closed-form linear mapping" that literally transfers the "memory" (kv cache) from a smaller model directly into a larger model.. without ever re-reading the prompt.
here is exactly why this is actual magic..
the discovery: they proved that kv caches across different sized models in the same family (like qwen3 14b to 32b) are linearly connected. a single layer in the small model can predict 56% of the variance in the big model's keys. if you bundle a few layers together, it jumps to 79%.
the rope trick: to make this work for any prompt length, they strip the "rope" (rotary positional embedding) off the keys before mapping them. they transfer the pure, position-free semantic data, and then re-apply the rotation on the target model.
the insane speed: this mathematical bridge transfers the context 2.7x to 25x faster than forcing the big model to re-prefill the text.
zero fine-tuning: there is no expensive training loop here. they just pass 500 texts (1,024 tokens each) through both models and solve a simple ridge regression equation per head. the calibration takes less than 90 minutes on an 8xh100 node.
the accuracy: across four out of six model pairs tested, this simple linear mapper retained 73-98% of the big model's native accuracy. for the two pairs that struggled, they use a nonlinear mlp fallback that recovers up to +37 points on hellaswag.
this means dynamic routing just got fully unlocked..
you can now use a cheap 8b model for the basic parts of a chat, and instantly hand off its exact brain state to a massive 70b model the second a user asks a complex coding question. zero latency penalty.
we are officially entering the era of fluid, cost-cascading ai architectures..