Incredible writeup! Some notable ๐s:
Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens.
They also use separate attention modes for disaggregated prefill vs decode (is this the first public account of arch difference between the two? ๐).
1/๐งต