가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Inferact
@inferact
Building the future of inference through @vllm_project
가입 December 2025
5 팔로잉 중    7K 팬
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
더 보기