注册并分享邀请链接,可获得视频播放与邀请奖励。

Susan Zhang
@suchenzang
Always hungry for intelligence.
加入 April 2014
1.4K 正在关注    49.8K 粉丝
Incredible writeup! Some notable 💎s: Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens. They also use separate attention modes for disaggregated prefill vs decode (is this the first public account of arch difference between the two? 👀). 1/🧵
显示更多
0
16
890
103
转发到社区