GDN-2 scales nicely with Muon 🚀
We trained hybrid MoE models at three scales (2B-7B) with Muon.
GDN-2 matches Mamba2's loss with ~10% less training compute!
Gated DeltaNet-2 (GDN-2) is now a lot faster on NVIDIA GPUs 🚀
Thanks to an amazing effort by our cuDNN team, we now have full cuDNN support for GDN-2. This covers not just prefill (forward), but also a highly optimized backward pass built on CUTLASS primitives, available now as part of the CUTLASS 4.7.0 release.
On GB300 (BF16, batch 4, 64 heads, d=128), compared to the FLA Triton implementation:
⚡ Forward: up to 6.4x faster
⚡ Backward: up to 2.8x faster
⚡ End-to-end: ~3x faster training
Throughput stays flat across sequence lengths from 2K to 32K, meaning the kernels are compute-bound where they should be.
Benchmarks below. 👇
If you are interested to know more about GDN-2, see our manuscript below:
I think we should literally stop calling every linear model as an SSM.
Mamba2: Sₜ = αₜSₜ₋₁ + kₜvₜᵀ
GDN: Sₜ = αₜ(I − βₜkₜkₜᵀ)Sₜ₋₁ + βₜkₜvₜᵀ
GDN-2: Sₜ = (I − kₜ(bₜ⊙kₜ)ᵀ)DₜSₜ₋₁ + kₜ(wₜ⊙vₜ)ᵀ
GDN family is a gradient step on a local regression loss, not a discretized ODE.
Umbrella term should be "linear RNNs", with SSMs as one sub-family.