가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Bowen Cheng
@bowenc0221
Research @Meta Superintelligence Lab; previously multimodal @OpenAI, research: "thinking with images", advance voice mode, perception
가입 September 2014
449 팔로잉 중    4.5K 팬
Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL which does ASR, diarization and endpointing all in real-time. It supports hour-long audio, 20+ speakers, multilingual, seamless code-switching, and contextual biasing. We shared detailed design and what makes Muse Voice Transcribe different from other real-time voice models in our blog post. We give full control on when and how long to listen to audio back to the model. And with RL, it learns “adaptive delay” to wait longer only on a few hard words. This is one of the key contributions to achieving the new pareto frontier on speed/accuracy trade-off. Muse Voice Transcribe is available today via Meta Model API, Meta AI for Mac, and Muse Code. Please give it a try and let us know how you like it:
더 보기