๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Yifan Qiao
@yifandotqiao
MTS @inferact | ex-Postdoc at Sky Computing Lab @UCBerkeley | PhD @UCLA | Building efficient systems for AI
๊ฐ€์ž… January 2023
420 ํŒ”๋กœ์ž‰ ์ค‘    632 ํŒฌ
I did not see TPUs beating GPUs on decode any time soon, until @woosuk_k and the cracked team pulled it off with a single Pallas kernel running all of Kimi K3 at 709 tok/s ๐Ÿคฏ First TPU inference megakernel I know of, and it is open source today! The numbers speak for themselves, and the blog is definitely worth your time ๐Ÿš€
๋” ๋ณด๊ธฐ