注册并分享邀请链接,可获得视频播放与邀请奖励。

Inferact
@inferact
Building the future of inference through @vllm_project
加入 December 2025
5 正在关注    7K 粉丝
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
显示更多
0
22
827
91
转发到社区