注册并分享邀请链接,可获得视频播放与邀请奖励。

Vincent Zhong
@vincentzed_cuda
GPU inference at @liquidai.
加入 January 2026
235 正在关注    661 粉丝
On KDA prefill (Kimi K3's attention), llms independently writes kernel that gets a 2.05x speedup over official FlashKDA on B200; and it's already merged and running in sglang. this kernel has 23 shape-specialized dispatch options and is the best performing kernel open source'd (in this workload). But how? We studied flashinfer-ai/flashinfer to figure out exactly how these kernels are produced. Inside: Many thousands of lines of straightline frozen cuda + 1/3 inline ptx Explanining why in 🧵, and code explainers (1/n)
显示更多
0
5
122
13
转发到社区