On KDA prefill (Kimi K3's attention), llms independently writes kernel that gets a 2.05x speedup over official FlashKDA on B200; and it's already merged and running in sglang. this kernel has 23 shape-specialized dispatch options and is the best performing kernel open source'd (in this workload).
But how?
We studied flashinfer-ai/flashinfer to figure out exactly how these kernels are produced.
Inside: Many thousands of lines of straightline frozen cuda + 1/3 inline ptx
Explanining why in 🧵, and code explainers (1/n)