็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

Wei Ping
@_weiping
Distinguished Research Scientist @NVIDIA | LLM post-training, reasoning, agentic coding, and multimodality
ๅ‚ๅŠ  June 2020
405 ใƒ•ใ‚ฉใƒญใƒผไธญ    3.7K ใƒ•ใ‚กใƒณ
๐Ÿš€ Introducing Audex, a unified audio-text LLM for text, speech, sound, and music ๐Ÿš€ ๐Ÿ† Audex delivers best-in-class performance among open models across: ๐ŸŽง Audio understanding ๐Ÿ—ฃ๏ธ Speech recognition and translation ๐Ÿ”Š Text-to-speech ๐ŸŽต General audio generation ๐Ÿ”„ Speech-to-speech generation ๐Ÿฅ‡ Audex also achieves best-in-class results in math and code reasoning, alignment, and instruction following, even outperforming the text-only Qwen3.5-35B-A3B. ๐Ÿงฉ Minimalist architecture: a single 30B-A3B MoE model. Audio inputs are projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. ๐Ÿง  Strong text intelligence, preserved โ€ข Built on the Nemotron-Cascade-2-30B-A3B text backbone โ€ข Trained with multi-stage SFT on blended audio-text data, Cascade RL, and multi-domain on-policy distillation โ€ข The result: broad and SOTA audio capabilities with no regression in text intelligence. ๐Ÿค— Model: ๐Ÿ‘‰ ๐Ÿ“„ Technical report: ๐Ÿ‘‰
ใ‚‚ใฃใจ่ฆ‹ใ‚‹