Incredible writeup! Some notable 💎s:
Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens.
They also use separate attention modes for disaggregated prefill vs decode (is this the first public account of arch difference between the two? 👀).
1/🧵
显示更多