๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Susan Zhang
@suchenzang
Always hungry for intelligence.
๊ฐ€์ž… April 2014
1.4K ํŒ”๋กœ์ž‰ ์ค‘    49.8K ํŒฌ
Incredible writeup! Some notable ๐Ÿ’Žs: Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens. They also use separate attention modes for disaggregated prefill vs decode (is this the first public account of arch difference between the two? ๐Ÿ‘€). 1/๐Ÿงต
๋” ๋ณด๊ธฐ