註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Vincent Zhong
@vincentzed_cuda
GPU inference at @liquidai.
加入 January 2026
235 正在關注    661 粉絲
On KDA prefill (Kimi K3's attention), llms independently writes kernel that gets a 2.05x speedup over official FlashKDA on B200; and it's already merged and running in sglang. this kernel has 23 shape-specialized dispatch options and is the best performing kernel open source'd (in this workload). But how? We studied flashinfer-ai/flashinfer to figure out exactly how these kernels are produced. Inside: Many thousands of lines of straightline frozen cuda + 1/3 inline ptx Explanining why in 🧵, and code explainers (1/n)
顯示更多
0
5
122
13
轉發到社區