Register and share your invite link to earn from video plays and referrals.

Vincent Zhong
@vincentzed_cuda
GPU inference at @liquidai.
Joined January 2026
235 Following    661 Followers
On KDA prefill (Kimi K3's attention), llms independently writes kernel that gets a 2.05x speedup over official FlashKDA on B200; and it's already merged and running in sglang. this kernel has 23 shape-specialized dispatch options and is the best performing kernel open source'd (in this workload). But how? We studied flashinfer-ai/flashinfer to figure out exactly how these kernels are produced. Inside: Many thousands of lines of straightline frozen cuda + 1/3 inline ptx Explanining why in 🧵, and code explainers (1/n)
Show more