๐ Introducing Audex, a unified audio-text LLM for text, speech, sound, and music ๐
๐ Audex delivers best-in-class performance among open models across:
๐ง Audio understanding
๐ฃ๏ธ Speech recognition and translation
๐ Text-to-speech
๐ต General audio generation
๐ Speech-to-speech generation
๐ฅ Audex also achieves best-in-class results in math and code reasoning, alignment, and instruction following, even outperforming the text-only Qwen3.5-35B-A3B.
๐งฉ Minimalist architecture: a single 30B-A3B MoE model. Audio inputs are projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation.
๐ง Strong text intelligence, preserved
โข Built on the Nemotron-Cascade-2-30B-A3B text backbone
โข Trained with multi-stage SFT on blended audio-text data, Cascade RL, and multi-domain on-policy distillation
โข The result: broad and SOTA audio capabilities with no regression in text intelligence.
๐ค Model:
๐
๐ Technical report:
๐