First time in my life seeing a benchmark scramble overnight to fit a model. It used to be benchmaxxing; now it’s model-maxxing 🤣 Build a great model, and the leaderboard will chase after you 😉
🚀 Introducing Audex, a unified audio-text LLM for text, speech, sound, and music 🚀
🏆 Audex delivers best-in-class performance among open models across:
🎧 Audio understanding
🗣️ Speech recognition and translation
🔊 Text-to-speech
🎵 General audio generation
🔄 Speech-to-speech generation
🥇 Audex also achieves best-in-class results in math and code reasoning, alignment, and instruction following, even outperforming the text-only Qwen3.5-35B-A3B.
🧩 Minimalist architecture: a single 30B-A3B MoE model. Audio inputs are projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation.
🧠 Strong text intelligence, preserved
• Built on the Nemotron-Cascade-2-30B-A3B text backbone
• Trained with multi-stage SFT on blended audio-text data, Cascade RL, and multi-domain on-policy distillation
• The result: broad and SOTA audio capabilities with no regression in text intelligence.
🤗 Model:
👉
📄 Technical report:
👉