๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Thien Tran
@gaunernst
๊ฐ€์ž… October 2016
276 ํŒ”๋กœ์ž‰ ์ค‘    3.1K ํŒฌ
Thank you @pranjalssh for the shoutout ๐Ÿฅฐ Some thoughts: - Reading lane_id is interesting... Usually I do warp_uniform() on warp_id (broadcast from lane0) + elect.sync. Wonder if there's perf advantage (at the mercy of ptxas) - ftw
๋” ๋ณด๊ธฐ
Sequel for my H100 blog is out now. We implement Blackwell matmul for NVFP4 from scratch, and outperform cuBLAS by 4.7% at N=8192. This time we have something better than Hilbert curves - that exact-fits for the Nvidia hardware. Full blog post with all details:
๋” ๋ณด๊ธฐ