๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Vincent Zhong
@vincentzed_cuda
GPU inference at @liquidai.
๊ฐ€์ž… January 2026
235 ํŒ”๋กœ์ž‰ ์ค‘    661 ํŒฌ
On KDA prefill (Kimi K3's attention), llms independently writes kernel that gets a 2.05x speedup over official FlashKDA on B200; and it's already merged and running in sglang. this kernel has 23 shape-specialized dispatch options and is the best performing kernel open source'd (in this workload). But how? We studied flashinfer-ai/flashinfer to figure out exactly how these kernels are produced. Inside: Many thousands of lines of straightline frozen cuda + 1/3 inline ptx Explanining why in ๐Ÿงต, and code explainers (1/n)
๋” ๋ณด๊ธฐ