Register and share your invite link to earn from video plays and referrals.

Wei Ping
@_weiping
Distinguished Research Scientist @NVIDIA | LLM post-training, reasoning, agentic coding, and multimodality
Joined June 2020
405 Following    3.7K Followers
🚀 Introducing Audex, a unified audio-text LLM for text, speech, sound, and music 🚀 🏆 Audex delivers best-in-class performance among open models across: 🎧 Audio understanding đŸ—Ŗī¸ Speech recognition and translation 🔊 Text-to-speech đŸŽĩ General audio generation 🔄 Speech-to-speech generation đŸĨ‡ Audex also achieves best-in-class results in math and code reasoning, alignment, and instruction following, even outperforming the text-only Qwen3.5-35B-A3B. 🧩 Minimalist architecture: a single 30B-A3B MoE model. Audio inputs are projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. 🧠 Strong text intelligence, preserved â€ĸ Built on the Nemotron-Cascade-2-30B-A3B text backbone â€ĸ Trained with multi-stage SFT on blended audio-text data, Cascade RL, and multi-domain on-policy distillation â€ĸ The result: broad and SOTA audio capabilities with no regression in text intelligence. 🤗 Model: 👉 📄 Technical report: 👉
Show more