Register and share your invite link to earn from video plays and referrals.

Susan Zhang
@suchenzang
Always hungry for intelligence.
Joined April 2014
1.4K Following    49.8K Followers
Incredible writeup! Some notable ๐Ÿ’Žs: Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens. They also use separate attention modes for disaggregated prefill vs decode (is this the first public account of arch difference between the two? ๐Ÿ‘€). 1/๐Ÿงต
Show more
0
16
890
103
Forward to community