登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Susan Zhang
@suchenzang
Always hungry for intelligence.
参加 April 2014
1.4K フォロー中    49.8K ファン
Incredible writeup! Some notable 💎s: Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens. They also use separate attention modes for disaggregated prefill vs decode (is this the first public account of arch difference between the two? 👀). 1/🧵
もっと見る