Register and share your invite link to earn from video plays and referrals.

Bowen Cheng
@bowenc0221
Research @Meta Superintelligence Lab; previously multimodal @OpenAI, research: "thinking with images", advance voice mode, perception
Joined September 2014
449 Following    4.5K Followers
Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL which does ASR, diarization and endpointing all in real-time. It supports hour-long audio, 20+ speakers, multilingual, seamless code-switching, and contextual biasing. We shared detailed design and what makes Muse Voice Transcribe different from other real-time voice models in our blog post. We give full control on when and how long to listen to audio back to the model. And with RL, it learns “adaptive delay” to wait longer only on a few hard words. This is one of the key contributions to achieving the new pareto frontier on speed/accuracy trade-off. Muse Voice Transcribe is available today via Meta Model API, Meta AI for Mac, and Muse Code. Please give it a try and let us know how you like it:
Show more