注册并分享邀请链接,可获得视频播放与邀请奖励。

Bowen Cheng
@bowenc0221
Research @Meta Superintelligence Lab; previously multimodal @OpenAI, research: "thinking with images", advance voice mode, perception
加入 September 2014
449 正在关注    4.5K 粉丝
Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL which does ASR, diarization and endpointing all in real-time. It supports hour-long audio, 20+ speakers, multilingual, seamless code-switching, and contextual biasing. We shared detailed design and what makes Muse Voice Transcribe different from other real-time voice models in our blog post. We give full control on when and how long to listen to audio back to the model. And with RL, it learns “adaptive delay” to wait longer only on a few hard words. This is one of the key contributions to achieving the new pareto frontier on speed/accuracy trade-off. Muse Voice Transcribe is available today via Meta Model API, Meta AI for Mac, and Muse Code. Please give it a try and let us know how you like it:
显示更多
0
7
110
14
转发到社区