注册并分享邀请链接,可获得视频播放与邀请奖励。

Inferact
@inferact
Building the future of inference through @vllm_project
加入 December 2025
5 正在关注    7K 粉丝
Thanks for the shoutout @SemiAnalysis_ ! Full breakdown linked here:
ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700 tok/s/user,  56% BETTER PERFORMANCE THAN NVIDIA GB200 NVL72 THROUGH MEGAKERNEL OPTIMIZATION ON KIMI K3. As we said awhile ago, the TPU externalization of software is full steam ahead. This is ultra important to follow the progress of this.
显示更多