가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

kwindla
@kwindla
Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
가입 September 2008
3.9K 팔로잉 중    15.9K 팬
I spend a lot of time on benchmarks, these days. I help maintain several public benchmarks, work with customers on private benchmarks, and talk to partners about how we can all build benchmarks that are useful for evaluating all of the components of our voice agents. This benchmark by @_josh_meyer_ and team is carefully constructed to measure the things that are critical to voice agent success, in a scenario that closely matches the real-world use case we see in enterprise deployments. A few notes, from my perspective, in praise of things I really appreciate about this benchmark work: - A good voice agent benchmark tests multi-turn conversation, with multiple defined tools and several tool calls during the conversation. - Users care about both overall conversation success and the turn-to-turn naturalness and latency during a conversation. It's important to do a good job quantitatively scoring both of these. You can aggregate them into a single score, or break out the components (overall success, voice-to-voice latency, unwanted interruptions, etc). - Teams building agents want to see comparisons between cascaded pipelines (TTS -> LLM -> STT) and speech-to-speech models. A good benchmark can judge both architectures on the same tasks. - Cost matters. Most teams think about cost as cost-per-minute. Translating token counts to cost-per minute is note easy. And even forecasting token counts is hard to do in the abstract. A realistic benchmark can generate token and caching data that helps pin down the cost of different models and architectures. - Tool call failures (and other types of agent failures, too) can be subtle. When you're working on a benchmark, it usually takes a fair amount of iteration and scrutinizing the transcripts to get to the point where your LLM-as-a-judge is doing a good job judging. - Open source benchmarks and detailed write-ups allow other people to learn from the hard work that goes into benchmarks, and force you to stand behind the choices and implementation quality of your benchmarking work. Benchmarks are hard! It's good to build them in the open.
더 보기
Introducing VAmoS Bench from @veris_ai 🔥 3,300 calls, 11 agents, 100 scenarios We compared what agents said with what they did @pipecat_ai @trydaily @livekit @Vapi_AI @elevenlabs @cartesia @OpenAI @GoogleAI @retellai @NVIDIAAI @DeepgramAI 🧵
더 보기