Register and share your invite link to earn from video plays and referrals.

Search results for gdn2
gdn2 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including gdn2
a better recurrence is all you need: S_t = (I - k_t (b_t . k_t)^T) Diag(exp(g_t)) S_{t-1} + k_t (w_t . v_t)^T #gdn2#
GDN-2 scales nicely with Muon 🚀 We trained hybrid MoE models at three scales (2B-7B) with Muon. GDN-2 matches Mamba2's loss with ~10% less training compute!
GDN-2 is fast, memory-efficient, and scalable. All you need to build a frontier hybrid LLM.
Gated DeltaNet-2 (GDN-2) is now a lot faster on NVIDIA GPUs 🚀 Thanks to an amazing effort by our cuDNN team, we now have full cuDNN support for GDN-2. This covers not just prefill (forward), but also a highly optimized backward pass built on CUTLASS primitives, available now as part of the CUTLASS 4.7.0 release. On GB300 (BF16, batch 4, 64 heads, d=128), compared to the FLA Triton implementation: ⚡ Forward: up to 6.4x faster ⚡ Backward: up to 2.8x faster ⚡ End-to-end: ~3x faster training Throughput stays flat across sequence lengths from 2K to 32K, meaning the kernels are compute-bound where they should be. Benchmarks below. 👇 If you are interested to know more about GDN-2, see our manuscript below:
Show more
I think we should literally stop calling every linear model as an SSM. Mamba2: Sₜ = αₜSₜ₋₁ + kₜvₜᵀ GDN: Sₜ = αₜ(I − βₜkₜkₜᵀ)Sₜ₋₁ + βₜkₜvₜᵀ GDN-2: Sₜ = (I − kₜ(bₜ⊙kₜ)ᵀ)DₜSₜ₋₁ + kₜ(wₜ⊙vₜ)ᵀ GDN family is a gradient step on a local regression loss, not a discretized ODE. Umbrella term should be "linear RNNs", with SSMs as one sub-family.
Show more