Register and share your invite link to earn from video plays and referrals.

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
180 Following    4.1K Followers
We’ve already trained larger variants of this model, and they outperform competing approaches, including Mamba-2, GDP, and KDA, by a substantial margin. One near-term release candidate is the hybrid GDN-2-3B latent MoE, which follows the Nemotron-3 Nano architecture but replaces the Mamba-2 layers with GDN-2. A public release requires several approvals, and we’re working hard to secure them!
Show more
I am happy to announce that Gated DeltaNet-2 has been accepted to NeurIPS 2026 ! 🎉 🎊🥂 Our models are the new SOTA for hybrid/linear models ! Paper: Code:
Show more
OPD for long-horizon agentic tasks is broken. But we have a solution. Please stay tuned !
In post-training, OPD can sit between SFT and final RL. RL + OPD with fused advantages is also an option, but requires careful annealing as OPD saturates quickly. Either way, the live RL run could have started from an OPD-trained checkpoint.
Show more
Hoping someone can explain to me what's going on here. The report says they trained multiple mixRL teachers, then used MOPD to combine those teachers into a student model. But..the team literally live-streamed their RL run, and then released model + report just one day after it finished. How does this timeline work out? I would assume the streamed RL run is just a final climb after MOPD, but the paper doesn't mention anything about a final climb. Maybe the livestreamed run is only for one of their mixRL teachers, and they did MOPD quickly in the 24 hours before launch?
Show more
GDN for diffusion LLMs 🚀💚
dQwen3.5: Hybrid-Attention Diffusion Language Models Adapting hybrid attention-RNN language models into diffusion models achieves a given training loss in about half the tokens of a full-attention control, while supporting strong parallel decoding.
Show more
I think we should literally stop calling every linear model as an SSM. Mamba2: Sₜ = αₜSₜ₋₁ + kₜvₜᵀ GDN: Sₜ = αₜ(I − βₜkₜkₜᵀ)Sₜ₋₁ + βₜkₜvₜᵀ GDN-2: Sₜ = (I − kₜ(bₜ⊙kₜ)ᵀ)DₜSₜ₋₁ + kₜ(wₜ⊙vₜ)ᵀ GDN family is a gradient step on a local regression loss, not a discretized ODE. Umbrella term should be "linear RNNs", with SSMs as one sub-family.
Show more
A carefully controlled look at looped transformers (arXiv 2609.19107): 1. Weight sharing is not compute-optimal on fresh data. It costs a constant ~1.06–1.16x compute at every scale. 2. The looped model's gains over vanilla come from elsewhere: the norm + input re-injection boundary operator, and growing depth mid-training. Both improve the scaling exponent. Neither needs weight sharing. 3. Looping is still a parameter-free way to get growth. Tied growth keeps the exponent gain at vanilla param count. 4. Weight sharing wins on multi-epoch, data-constrained training.
Show more
Personal experience: Muon is very strong for post-training (SFT+RLVR). p.s: always read the details and check code when a paper claims "x did not work". The devil is always in the details.
I don't believe ( establishes anything about Muon for RLVR. They run GRPO with Qwen3-1.7B/4B over GSM8K, and report that Muon consistently fails. But as always, the devil is in the details. Just quickly checking their setup in their repo, Muon's per-element update RMS comes out at 5×lr. AdamW's is typically 0.2 to 0.4×lr. So they are effectively running Muon at roughly 20x AdamW's step size at the same LR. And of course they did not sweep the optimal LR for Muon in this recipe. We have seen massive success in both pre-training and post-training experiment for Muon and there's no reason to believe otherwise.
Show more
I don't believe ( establishes anything about Muon for RLVR. They run GRPO with Qwen3-1.7B/4B over GSM8K, and report that Muon consistently fails. But as always, the devil is in the details. Just quickly checking their setup in their repo, Muon's per-element update RMS comes out at 5×lr. AdamW's is typically 0.2 to 0.4×lr. So they are effectively running Muon at roughly 20x AdamW's step size at the same LR. And of course they did not sweep the optimal LR for Muon in this recipe. We have seen massive success in both pre-training and post-training experiment for Muon and there's no reason to believe otherwise.
Show more
Contrary to the claims in this tweet, ML community did not stop working on alignment and safety. The proposals laid out for future research directions are already what post-training has been doing. Let's take a closer look: (1) "How do we take a gradient w.r.t. alignment?" That's what RLHF did ! train a reward model on human comparisons, take the policy gradient against it. DPO [1] drops the RL loop entirely and gives you a closed-form loss on preference pairs, a direct gradient of the policy w.r.t. preference data. (2) "Three laws as the objective." That's exactly what Constitutional AI [2] proposed. A written list of natural-language principles, the model critiques and revises against them, and a preference model is trained from that feedback. Along the same lines, Deliberative alignment [3] goes further and trains the model to read the spec and reason over it before answering . (3) "RL needs unsuccessful rollouts, so we'd harm humans." This is probably true for robotics not language modeling. A rollout is text generated in a sandbox and scored by a judge. The unsuccessful rollout is the model writing something harmful and the reward model scoring it low. Nobody is harmed. That's the point of a learned reward. Even for agents, training and evals run in sandboxed environments with simulated users and honeypots. So the "simulated environment" branch is not something we are going to do in the future. It's already the default option. (4) "Pretraining isn't the alignment objective." Mostly, but true but not entirely. Conditioning pretraining on preference labels [4] beat post-hoc alignment in their setup, and filtering hazardous domains out of pretraining data is done in production. You can ask any pretraining team about this and they will tell you ! (5) "Expensive environments, so hackable proxies." This is probably the only part of the tweet I agree with, and it has been well documented. Reward hacking and preference models that reward sycophancy are real issues. But there are also several effective mitigations, and they attack the problem at different stages. To make the proxy harder to game, rule-based rewards [5] grade outputs against explicit written rules with an LLM grader, and reward-model ensembles [6] stop the policy from exploiting the blind spots of any single reward model. To catch hacking when it happens, chain-of-thought monitoring [7] has a separate model read the policy's reasoning, which catches hacks the outcome reward misses. Although, [7] in my opinion needs to be used with caution as training against the monitor teaches the model to hide its intent, so it should flag, not grade. And to contain the damage, inoculation prompting [8] stops reward hacking that does get learned from generalizing into broader misalignment. My understanding is that adding and output gating with safety classifiers at inference may just reduce the rate of policy mistakes that reach a user. In my opinion, the open direction is actually not how to take the gradient. It's whether the policy generalizes the principles off-distribution (alignment faking, emergent misalignment from narrow finetuning), whether models behave the same when they think they're being evaluated. References [1] Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D. and Finn, C. (2023) 'Direct Preference Optimization: Your Language Model is Secretly a Reward Model', Advances in Neural Information Processing Systems, 36. [2] Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C. et al. (2022) 'Constitutional AI: Harmlessness from AI Feedback'. arXiv preprint arXiv:2212.08073. [3] Guan, M.Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., Helyar, A., Dias, R., Vallone, A., Ren, H., Wei, J., Chung, H.W., Toyer, S., Heidecke, J., Beutel, A. and Glaese, A. (2024) 'Deliberative Alignment: Reasoning Enables Safer Language Models'. arXiv preprint arXiv:2412.16339. [4] Korbak, T., Shi, K., Chen, A., Bhalerao, R., Buckley, C.L., Phang, J., Bowman, S.R. and Perez, E. (2023) 'Pretraining Language Models with Human Preferences', Proceedings of the 40th International Conference on Machine Learning, PMLR 202, pp. 17506-17533 [5] Mu, T., Helyar, A., Heidecke, J., Achiam, J., Vallone, A., Kivlichan, I., Lin, M., Beutel, A., Schulman, J. and Weng, L. (2024) 'Rule Based Rewards for Language Model Safety', Advances in Neural Information Processing Systems, 37. [6] Coste, T., Anwar, U., Kirk, R. and Krueger, D. (2024) 'Reward Model Ensembles Help Mitigate Overoptimization', The Twelfth International Conference on Learning Representations (ICLR). [7] Baker, B., Huizinga, J., Gao, L., Dou, Z., Guan, M.Y., Madry, A., Zaremba, W., Pachocki, J. and Farhi, D. (2025) 'Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation'. arXiv preprint arXiv:2503.11926 [8] MacDiarmid, M., Wright, B., Uesato, J., Benton, J., Kutasov, J., Price, S., Bouscal, N., Bowman, S., Bricken, T., Cloud, A., Denison, C., Gasteiger, J., Greenblatt, R., Leike, J., Lindsey, J., Mikulik, V., Perez, E., Rodrigues, A., Thomas, D., Webson, A., Ziegler, D. and Hubinger, E. (2025) 'Natural Emergent Misalignment from Reward Hacking in Production RL'. arXiv preprint arXiv:2511.18397.
Show more
Alignment is not as hard to solve as many claim, but it is in the end an algorithmic problem which a lot of ML community mostly stopped working on. Formulation of alignment stated by three (or four) laws of robotics can take us very far, so we roughly know the objective. The tricky part is, how do we take gradient with respect to alignment? We have two algorithms right now at our disposal: pretraining and RL. Pretraining takes gradient with respect to next token prediction - that's def not the alignment objective. We could create RL environments that embody the alignment objective, but: - those are expensive to create, so often cheaper, hackable proxies are used in practice - RL as an objective needs successful and unsuccessful rollouts to happen to take the gradient step. We DO NOT want to harm any humans in the process of aligning our modes - this is a pretty big problem. Therefore there are two solutions forward for the alignment problem: - either we RL models in a simulated environments with simulated alignments and decreasing the likelihood of harming simulated humans, which will never be perfect - or we create a new algorithm that can teach our models to not harm humans without harming any humans in the process Science is the process how we solve the hardest problems ahead and that is one of them
Show more
Jensen calls Jacob Coxon comments outlandish & “deeply untrue.” Said the labs are great, but his comments were wrong, arrogant, and ignorant of all the work being done around the industry to drive safety. 🧐🧐
Show more
0
160
2.7K
305
Forward to community
Impressive, but why is DS not using linear attention yet ?
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
Show more
Why the recurrent half of a hybrid LLM is actually easy to quantize Minima quantized all 496 linear layers of Qwen3.8-27B — including Gated DeltaNet — to NVFP4 W4A4, matching BF16 performance while cutting size ~2.9×
Show more
this is a crazy response to very reasonable feedback on the paper LOL why must we optimize papers for clicks and engagement? isn’t the point of writing papers to disseminate knowledge into the community? this is seriously a bad look for all coauthors… fwiw, why not use SoTA linear layers like KDA or GDN2 for ablations?
Show more
GDN-2 scales nicely with Muon 🚀 We trained hybrid MoE models at three scales (2B-7B) with Muon. GDN-2 matches Mamba2's loss with ~10% less training compute!
Gated DeltaNet-2 (GDN-2) is now a lot faster on NVIDIA GPUs 🚀 Thanks to an amazing effort by our cuDNN team, we now have full cuDNN support for GDN-2. This covers not just prefill (forward), but also a highly optimized backward pass built on CUTLASS primitives, available now as part of the CUTLASS 4.7.0 release. On GB300 (BF16, batch 4, 64 heads, d=128), compared to the FLA Triton implementation: ⚡ Forward: up to 6.4x faster ⚡ Backward: up to 2.8x faster ⚡ End-to-end: ~3x faster training Throughput stays flat across sequence lengths from 2K to 32K, meaning the kernels are compute-bound where they should be. Benchmarks below. 👇 If you are interested to know more about GDN-2, see our manuscript below:
Show more
GDN-2 is fast, memory-efficient, and scalable. All you need to build a frontier hybrid LLM.
a better recurrence is all you need: S_t = (I - k_t (b_t . k_t)^T) Diag(exp(g_t)) S_{t-1} + k_t (w_t . v_t)^T #gdn2#