่จปๅ†Šไธฆๅˆ†ไบซ้‚€่ซ‹้€ฃ็ต๏ผŒๅฏ็ฒๅพ—ๅฝฑ็‰‡ๆ’ญๆ”พ่ˆ‡้‚€่ซ‹็Žๅ‹ตใ€‚

Wei Ping
@_weiping
Distinguished Research Scientist @NVIDIA | LLM post-training, reasoning, agentic coding, and multimodality
ๅŠ ๅ…ฅ June 2020
405 ๆญฃๅœจ้—œๆณจ    3.7K ็ฒ‰็ตฒ
๐Ÿš€ Introducing Audex, a unified audio-text LLM for text, speech, sound, and music ๐Ÿš€ ๐Ÿ† Audex delivers best-in-class performance among open models across: ๐ŸŽง Audio understanding ๐Ÿ—ฃ๏ธ Speech recognition and translation ๐Ÿ”Š Text-to-speech ๐ŸŽต General audio generation ๐Ÿ”„ Speech-to-speech generation ๐Ÿฅ‡ Audex also achieves best-in-class results in math and code reasoning, alignment, and instruction following, even outperforming the text-only Qwen3.5-35B-A3B. ๐Ÿงฉ Minimalist architecture: a single 30B-A3B MoE model. Audio inputs are projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. ๐Ÿง  Strong text intelligence, preserved โ€ข Built on the Nemotron-Cascade-2-30B-A3B text backbone โ€ข Trained with multi-stage SFT on blended audio-text data, Cascade RL, and multi-domain on-policy distillation โ€ข The result: broad and SOTA audio capabilities with no regression in text intelligence. ๐Ÿค— Model: ๐Ÿ‘‰ ๐Ÿ“„ Technical report: ๐Ÿ‘‰
้กฏ็คบๆ›ดๅคš