Register and share your invite link to earn from video plays and referrals.

Superman
@thesupermannx
AI & Neuroscience
Joined May 2026
17 Following    14.7K Followers
Nvidia researchers did it again! they found a way to transfer KV caches directly between different AI models. if you run llms, you know the absolute pain of "prefilling". if you want to swap from a small, cheap model to a large, smart model mid-conversation, the big model has to recalculate the entire chat history from scratch. it kills your latency and your compute budget. but a researchers just found a way to completely skip this.. they built a "closed-form linear mapping" that literally transfers the "memory" (kv cache) from a smaller model directly into a larger model.. without ever re-reading the prompt. here is exactly why this is actual magic.. the discovery: they proved that kv caches across different sized models in the same family (like qwen3 14b to 32b) are linearly connected. a single layer in the small model can predict 56% of the variance in the big model's keys. if you bundle a few layers together, it jumps to 79%. the rope trick: to make this work for any prompt length, they strip the "rope" (rotary positional embedding) off the keys before mapping them. they transfer the pure, position-free semantic data, and then re-apply the rotation on the target model. the insane speed: this mathematical bridge transfers the context 2.7x to 25x faster than forcing the big model to re-prefill the text. zero fine-tuning: there is no expensive training loop here. they just pass 500 texts (1,024 tokens each) through both models and solve a simple ridge regression equation per head. the calibration takes less than 90 minutes on an 8xh100 node. the accuracy: across four out of six model pairs tested, this simple linear mapper retained 73-98% of the big model's native accuracy. for the two pairs that struggled, they use a nonlinear mlp fallback that recovers up to +37 points on hellaswag. this means dynamic routing just got fully unlocked.. you can now use a cheap 8b model for the basic parts of a chat, and instantly hand off its exact brain state to a massive 70b model the second a user asks a complex coding question. zero latency penalty. we are officially entering the era of fluid, cost-cascading ai architectures..
Show more