Register and share your invite link to earn from video plays and referrals.

kwindla
@kwindla
Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
Joined September 2008
3.9K Following    15.9K Followers
I joined the Pipecat TV crew to talk about PhoneLLM: a super low-latency LLM trained on voice agent scenarios. We talked about building agents with a small-ish LLM like this, compared to using a big model. PhoneLLM is based on Nemotron 3 Nano, so it's a 30B-A3B MoE. Prompting a 30B-A3B model reminds me of prompting GPT-4. The model will try very, very hard to do what you ask it to do. But it will follow your instructions more literally than you maybe expect. And if you aren't careful about specifying "guardrails" in your prompt, it will hallucinate. You need precise prompting and unambiguous tool definitions. Which brought us around to talking about evals. Every agent that's in production needs evals. And if you have good evals, then you can use a smaller model that requires more careful prompting. (Because you know what's working and what's not, in both test and real conversations.) We also talked about our latency targets. PhoneLLM's server-side TTFT is sub-100ms. A B200 can host 80 real-world agent sessions with an end-to-end p95 TTFT, measured client-side, under 600ms. These are latencies and per-agent cost numbers that aren't achievable if you're using LLMs via public APIs. Having said that, I always tell people to prototype voice agents in the easiest possible way: use hosted services rather than running your own infra, use big models until your scenarios are well-defined and you have evals in place, don't optimize *anything* until you have working agents that do most of what you want them to do. Finally, we talked about training your own models. PhoneLLM is a product of the low-latency model factory we've been working on for a long time. Low-latency in two senses: we can quickly and repeatably do custom training runs, and the use cases we're focused on are low-latency applications like voice agents. Small, customized models are taking over more and more agent workloads. Open weights base models are very good across a wide range of sizes. We know how to build and maintain good evals. And, across the industry, our training/fine-tuning tooling is maturing.
Show more