註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Inferact
@inferact
Building the future of inference through @vllm_project
加入 December 2025
5 正在關注    7K 粉絲
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
顯示更多
0
22
827
91
轉發到社區