Thank you @pranjalssh for the shoutout 🥰
Some thoughts:
- Reading lane_id is interesting... Usually I do warp_uniform() on warp_id (broadcast from lane0) + elect.sync. Wonder if there's perf advantage (at the mercy of ptxas)
- ftw
Sequel for my H100 blog is out now.
We implement Blackwell matmul for NVFP4 from scratch, and outperform cuBLAS by 4.7% at N=8192.
This time we have something better than Hilbert curves - that exact-fits for the Nvidia hardware.
Full blog post with all details: