登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Bowen Cheng
@bowenc0221
Research @Meta Superintelligence Lab; previously multimodal @OpenAI, research: "thinking with images", advance voice mode, perception
参加 September 2014
449 フォロー中    4.5K ファン
Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL which does ASR, diarization and endpointing all in real-time. It supports hour-long audio, 20+ speakers, multilingual, seamless code-switching, and contextual biasing. We shared detailed design and what makes Muse Voice Transcribe different from other real-time voice models in our blog post. We give full control on when and how long to listen to audio back to the model. And with RL, it learns “adaptive delay” to wait longer only on a few hard words. This is one of the key contributions to achieving the new pareto frontier on speed/accuracy trade-off. Muse Voice Transcribe is available today via Meta Model API, Meta AI for Mac, and Muse Code. Please give it a try and let us know how you like it:
もっと見る