We’ve already trained larger variants of this model, and they outperform competing approaches, including Mamba-2, GDP, and KDA, by a substantial margin.
One near-term release candidate is the hybrid GDN-2-3B latent MoE, which follows the Nemotron-3 Nano architecture but replaces the Mamba-2 layers with GDN-2.
A public release requires several approvals, and we’re working hard to secure them!