đ Introducing Audex, a unified audio-text LLM for text, speech, sound, and music đ
đ Audex delivers best-in-class performance among open models across:
đ§ Audio understanding
đŖī¸ Speech recognition and translation
đ Text-to-speech
đĩ General audio generation
đ Speech-to-speech generation
đĨ Audex also achieves best-in-class results in math and code reasoning, alignment, and instruction following, even outperforming the text-only Qwen3.5-35B-A3B.
đ§Š Minimalist architecture: a single 30B-A3B MoE model. Audio inputs are projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation.
đ§ Strong text intelligence, preserved
âĸ Built on the Nemotron-Cascade-2-30B-A3B text backbone
âĸ Trained with multi-stage SFT on blended audio-text data, Cascade RL, and multi-domain on-policy distillation
âĸ The result: broad and SOTA audio capabilities with no regression in text intelligence.
đ¤ Model:
đ
đ Technical report:
đ