登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Inferact
@inferact
Building the future of inference through @vllm_project
参加 December 2025
5 フォロー中    7K ファン
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
もっと見る