註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    50.3K 粉絲
The @inferact team open sourced a TPU megakernel for Kimi K3 achieving 709 tokens/s, against 450 tokens/s on GB200. All 92 of K3's MoE layers run in a single Pallas kernel, with weight prefetching that reaches across layer boundaries so transfers for one layer overlap with computation in the previous one. Shoutout to the team! Writeup and repo link in the thread.
顯示更多
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
顯示更多
0
8
129
11
轉發到社區