ๆณจๅ†Œๅนถๅˆ†ไบซ้‚€่ฏท้“พๆŽฅ๏ผŒๅฏ่Žทๅพ—่ง†้ข‘ๆ’ญๆ”พไธŽ้‚€่ฏทๅฅ–ๅŠฑใ€‚

Wei Ping
@_weiping
Distinguished Research Scientist @NVIDIA | LLM post-training, reasoning, agentic coding, and multimodality
ๅŠ ๅ…ฅ June 2020
405 ๆญฃๅœจๅ…ณๆณจ    3.7K ็ฒ‰ไธ
๐Ÿš€ Introducing Audex, a unified audio-text LLM for text, speech, sound, and music ๐Ÿš€ ๐Ÿ† Audex delivers best-in-class performance among open models across: ๐ŸŽง Audio understanding ๐Ÿ—ฃ๏ธ Speech recognition and translation ๐Ÿ”Š Text-to-speech ๐ŸŽต General audio generation ๐Ÿ”„ Speech-to-speech generation ๐Ÿฅ‡ Audex also achieves best-in-class results in math and code reasoning, alignment, and instruction following, even outperforming the text-only Qwen3.5-35B-A3B. ๐Ÿงฉ Minimalist architecture: a single 30B-A3B MoE model. Audio inputs are projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. ๐Ÿง  Strong text intelligence, preserved โ€ข Built on the Nemotron-Cascade-2-30B-A3B text backbone โ€ข Trained with multi-stage SFT on blended audio-text data, Cascade RL, and multi-domain on-policy distillation โ€ข The result: broad and SOTA audio capabilities with no regression in text intelligence. ๐Ÿค— Model: ๐Ÿ‘‰ ๐Ÿ“„ Technical report: ๐Ÿ‘‰
ๆ˜พ็คบๆ›ดๅคš