登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

kwindla
@kwindla
Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
参加 September 2008
3.9K フォロー中    15.9K ファン
Pipecat v1.9.0 is out today, with support for @AIatMeta's new Muse Voice Transcribe model. Muse Voice Transcribe is a streaming speech-to-text model. It has the best (lowest) semantic word error rate of any model we've tested. We maintain a public, open source test suite called the Pipecat STT benchmark. This benchmark measures how well an STT model performs in a voice agent pipeline. We calculate a "semantic word error rate" across 1,000 speech input fragments. Semantic WER ignores small differences in transcription that don't impact an LLM's understanding of user speech, and in our experience is a better proxy for STT model accuracy than the simpler WER algorithms used in most speech recognition benchmarks. The other critical metric for a voice agent is latency. We measure latency as "time to final segment" of the transcription. You can think of this as how long it takes for the STT model to deliver a complete text transcription after the user finishes speaking. Muse Voice Transcribe's TTFS is very good, though its long tail TTFS is a bit higher than we'd like to see. (The Meta team says they are optimizing for natural, model-driven endpointing, which certainly makes sense.) The P50 TTFS is 392 milliseconds and the P95 is 1,292 milliseconds. We'd like to see a P50 under 300ms and a P95 as close to that as possible! Note, though, that this P50 number is as good as we've seen for any initial release of a new ASR model. The model also supports speaker labels (diarization) and keyword biasing, has built-in endpointing, and can transcribe mixed-language speech. Kudos to the Meta team. We love seeing new models that are strong performers in realtime voice agent pipelines.
もっと見る