๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
๊ฐ€์ž… June 2015
180 ํŒ”๋กœ์ž‰ ์ค‘    4.1K ํŒฌ
Gated DeltaNet-2 (GDN-2) is now a lot faster on NVIDIA GPUs ๐Ÿš€ Thanks to an amazing effort by our cuDNN team, we now have full cuDNN support for GDN-2. This covers not just prefill (forward), but also a highly optimized backward pass built on CUTLASS primitives, available now as part of the CUTLASS 4.7.0 release. On GB300 (BF16, batch 4, 64 heads, d=128), compared to the FLA Triton implementation: โšก Forward: up to 6.4x faster โšก Backward: up to 2.8x faster โšก End-to-end: ~3x faster training Throughput stays flat across sequence lengths from 2K to 32K, meaning the kernels are compute-bound where they should be. Benchmarks below. ๐Ÿ‘‡ If you are interested to know more about GDN-2, see our manuscript below:
๋” ๋ณด๊ธฐ