註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Inferact
@inferact
Building the future of inference through @vllm_project
加入 December 2025
5 正在關注    7K 粉絲
Kimi K3 on @vllm_project is now 2.2–2.8× faster 🚀 Inferact is proud to have co-led this optimization effort with @RedHat_AI, @NVIDIAAI, and @Huawei, spanning scheduling, KDA state handling, and custom MoE kernels. Read the technical deep dive:
顯示更多
Kimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1. We break down the work across scheduling, KDA state handling, and MoE kernels, with benchmarks and commands to reproduce the results. Thanks to the vLLM community for pushing Kimi K3 performance forward! Read the deep dive:
顯示更多
0
3
66
15
轉發到社區