This project is a really clean architecture example of software built from the ground up around LLMs.
The UI is WebView running the Pipecat JavaScript client, connected to a native macOS audio transport with echo cancellation, microphone device management, etc.
Audio, vision, screen, history, ui, and shell workers communicate via a bus.
Workers pull messages off the bus, do LLM-ish things and drive other systems, and put messages back on the bus.
I'm pretty convinced that this is more or less how all software is going to look, fairly soon.
Show more
.
@aconchillo has been hacking on an experimental macOS voice assistant. Audio transcription and wake word monitoring run locally. LLM and screen vision use cloud models by default, but of course you can swap in a local model/endpoint.
"Peekaboo, tell me when the build finishes."
"Peekaboo, somebody invited me to a meetup about inference optimization tonight, but I can't find the message. Was that in Slack, email, WhatsApp, or iMessage?"
"Peekaboo, did we do the evals last week for the new Nemotron model on boule, pancake, flatbread, or the dgx spark? I think we forgot to push to the repo."
Show more
Pipecat v1.9.0 is out today, with support for
@AIatMeta's new Muse Voice Transcribe model.
Muse Voice Transcribe is a streaming speech-to-text model. It has the best (lowest) semantic word error rate of any model we've tested.
We maintain a public, open source test suite called the Pipecat STT benchmark. This benchmark measures how well an STT model performs in a voice agent pipeline. We calculate a "semantic word error rate" across 1,000 speech input fragments. Semantic WER ignores small differences in transcription that don't impact an LLM's understanding of user speech, and in our experience is a better proxy for STT model accuracy than the simpler WER algorithms used in most speech recognition benchmarks.
The other critical metric for a voice agent is latency. We measure latency as "time to final segment" of the transcription. You can think of this as how long it takes for the STT model to deliver a complete text transcription after the user finishes speaking. Muse Voice Transcribe's TTFS is very good, though its long tail TTFS is a bit higher than we'd like to see. (The Meta team says they are optimizing for natural, model-driven endpointing, which certainly makes sense.) The P50 TTFS is 392 milliseconds and the P95 is 1,292 milliseconds. We'd like to see a P50 under 300ms and a P95 as close to that as possible! Note, though, that this P50 number is as good as we've seen for any initial release of a new ASR model.
The model also supports speaker labels (diarization) and keyword biasing, has built-in endpointing, and can transcribe mixed-language speech.
Kudos to the Meta team. We love seeing new models that are strong performers in realtime voice agent pipelines.
Show more
Really cool to see the new GPT-Live-1 model and API in production, today.
For me, the most exciting thing about the new model is that it's trained to be used together with other models. This is important because voice agents need to be continuously responsive. You can't block the voice conversation loop. But agents often need to make tool calls, interact with backend systems, and leverage test-time-compute.
So GPT-Live-1 natively supports working cooperatively with a backend model. You prompt the frontend model and the backend model separately.
Here's part of an example from the GPT-Live-1 docs:
```
Delegation policy:
Backend tools:
- Appointments: check available times and create, change, or cancel bookings.
Delegate to the backend when:
- The user asks for availability or wants to create, change, or cancel a booking.
- A correction changes a booking task already in progress.
- The answer needs careful reasoning beyond a simple reply.
```
Multi-model systems are the future of software.
Almost everything I build for myself, now, and most of the work I do with customers and partners, involves multiple models, multiple inference loops, multiple prompts, and multiple layers of context management/sharing.
That's true whether I'm building voice agents, desktop software, or duct taping together management tooling for all the coding agents I have running all the time across various machines in my house and in the cloud.
We're just starting to explore the abstractions that make it easy and productive to distribute work between multiple LLMs and coordinate their output.
OpenAI's "delegation" pattern is one such abstraction. It's great to see this pattern trained into a production model, so that we can leverage it natively at the prompt level.
GPT-Live-1 is also the first production speech-to-speech model that can talk and listen at the same time. That ability unlocks better turn detection and more natural voice behaviors (like making small filler sounds that let people know you're listening to them, called "backchanneling").
Several research models have implemented full-duplex audio streaming, most notably the Moshi model from Kyutai. If you are interested in LLM architectures and you haven't spent much time with the 2024 paper about Moshi, stop reading this and go load that paper up in your web browser. I promise it's worth it.
But Moshi was a much smaller model than GPT-Live-1, and we've had two years of AI research progress in general since Moshi was released. Full duplex is a hard problem, and it's wonderful to see a model that pushes the frontiers of audio capabilities, released as a general availability product, served by a production API.
And speaking of the API, the OpenAI team has evolved the realtime API to support a complete set of context engineering and session management capabilities. We can now build things like state machine conversation graphs, threaded/backtracking modes, and structured data input that were difficult with the first generation of speech-to-speech APIs.
Congratulations to everyone at OpenAI who worked on this model and API. It's great to see. Er, to hear. :-)
Show more
Very excited to announce we're launching GPT-Live-1 in the API today -- this is a state-of-the-art voice model with unprecedented turn-taking naturalness and the ability to delegate work to the text model of your choice.
Show more
One of the lessons from our first couple of years building voice agents is that AI product teams choose a voice entirely based on vibes. This was a bit surprising, given how quantitative so many other agent metrics are.
Voice is emotional. People know the right voice when they hear it.
People often say to us, "I listened to all the voices in the catalog and none of them are quite right."
Gradium's new voice designer makes it easy to iteratively explore the design space, try wildly different things, narrow in on and refine ideas that are promising, and do all of this conversationally by talking to the designer interface, if you want to. (Also, this is really fun. Go try it.)
Show more
Today, we’re launching Voice Design.
Write a prompt, create a voice.
Describe the accent, age, gender, and pace your use case needs, and get new voices in seconds, ready to use.
Live and free in the Gradium API and Studio.
Show more
I've been meaning to write up how/why I use tmux to manage my agent herds. Alexis explains it better than I could have.
I’m seeing lots of tweet replies asking basically “but how do I get two models to talk to each other? Without me manually copy/pasting results between apps?”
And I’ve heard this same question repeatedly from friends who use AIs mainly through GUI apps.
Yes, there are various apps starting to be released for just this problem. But it’s worth knowing that in a pinch you can already do all this yourself, on the command line, just by using tmux.
Here’s how.
On the command line start a named tmux session with the command “tmux new -s chat”.
Create a second window within that session. (The default keystroke for this is doing “ctrl-b c”. You can switch between windows by doing “ctrl-b n” for the next window.)
Launch the codex CLI in one window, and the claude CLI in the other.
Then you can literally say to the AI in window 1, “please read the analysis of the AI in window 2 of the tmux session ‘chat’. Wdyt?” And vice versa. Now you have cross-harness review.
Or you can instruct one AI to delegate and manage the other AI over multiple turns, or monitor it over time, or whatever. This works because each AI can use the tmux CLI to read and write to the other’s window.
This does not on its own solve coordination problems like deciding who is in charge, or who talks first, or who waits for whom, but it establishes the basic communication primitives of “read the other AI’s output” and “write to the other AI”. You can build more complex workflows on top of that mostly by prompting.
The fundamental reason this all works is that the Unix-era command line environment is more interoperable than the modern GUI environment, so a decades-old tool like tmux is still one of the best ways to compose modern AIs. Obviously the command line tools are a PITA in other ways. But it’s great that they are so flexible, they already exist, and they work right now, so you don’t need to wait for the vendors to get their apps to play nicely with each other (which might never happen) or for someone else to catch up and write a new app just to connect things.
Show more
If you're interested in voice UI, or music, or both, it's worth watching all of this video that Jon just dropped.
Voice-controlled, generative music.
The interface will be familiar to anyone who has used a digital audio workstation. But also ... deep integration of voice makes this feel like a very new kind of software.
I think it might be that when we've fully figured out the transition to voice-first UIs, we'll feel like the transition from GUI to AI-native, voice-first software was as big as the transition from command-line to GUI software.
Jon has shown me videos of various stages of development on this project. We've been talking a little bit about voice UI patterns for consumer software and voice UI patterns for professional software.
Both categories are clearly going to become voice-centric. Talking is just so efficient and flexible.
Show more
Over a year ago, I posted a video where Gemini and I collaborated on an Ableton Live Session. Music AI is fun! Creating voice-native music software is something I'm really passionate about, so, here is Jamcat.
Jamcat's concept is a voice controlled, generative music performance tool, designed to inspire live, in the moment jam sessions. It uses LLMs, with both text-to-audio and audio-to-audio models to create music in realtime, entirely hands-free.
➡️ Pipecat PhoneLLM Alpha 1 for session control, running on
@modal.
➡️ "StemGen" routing and workload subagent.
➡️ Foundation 1, for melodic stems.
➡️ Various audio models on
@fal (like the excellent
@ElevenLabs Music) for vocals, drums, textures and one-shots.
➡️ ... and of course
@pipecat_ai!
Voice commands can be toggled via push-to-talk (or a foot pedal, if you're wielding an axe), or it can listen the entire time (diarization.)
It makes some mistakes, or sometimes misses a beat, but most of the mistakes in this video were me not paying attention to the bar counts, like a true fake musician.
Show more
I joined the Pipecat TV crew to talk about PhoneLLM: a super low-latency LLM trained on voice agent scenarios.
We talked about building agents with a small-ish LLM like this, compared to using a big model. PhoneLLM is based on Nemotron 3 Nano, so it's a 30B-A3B MoE.
Prompting a 30B-A3B model reminds me of prompting GPT-4.
The model will try very, very hard to do what you ask it to do. But it will follow your instructions more literally than you maybe expect. And if you aren't careful about specifying "guardrails" in your prompt, it will hallucinate. You need precise prompting and unambiguous tool definitions.
Which brought us around to talking about evals. Every agent that's in production needs evals. And if you have good evals, then you can use a smaller model that requires more careful prompting. (Because you know what's working and what's not, in both test and real conversations.)
We also talked about our latency targets. PhoneLLM's server-side TTFT is sub-100ms. A B200 can host 80 real-world agent sessions with an end-to-end p95 TTFT, measured client-side, under 600ms. These are latencies and per-agent cost numbers that aren't achievable if you're using LLMs via public APIs.
Having said that, I always tell people to prototype voice agents in the easiest possible way: use hosted services rather than running your own infra, use big models until your scenarios are well-defined and you have evals in place, don't optimize *anything* until you have working agents that do most of what you want them to do.
Finally, we talked about training your own models. PhoneLLM is a product of the low-latency model factory we've been working on for a long time. Low-latency in two senses: we can quickly and repeatably do custom training runs, and the use cases we're focused on are low-latency applications like voice agents.
Small, customized models are taking over more and more agent workloads. Open weights base models are very good across a wide range of sizes. We know how to build and maintain good evals. And, across the industry, our training/fine-tuning tooling is maturing.
Show more
Sample code for a customer support agent built with the new PhoneLLM Alpha 1 model.
(Including a very clean web UI. Despite the name, there's nothing stopping you from connecting web apps to PhoneLLM. 😄)
Show more
Here is an example Pipecat project for trying PhoneLLM Alpha 1. Deploy the model to a Modal endpoint with one click. Client from the video in repo too (with all that sweet sweet terminal-ish aura.)
Next up, a TTS that can pronounce
@bmervetan's name correctly? 😅
Show more
Come hang it with me on ThursdAI to talk about the PhoneLLM model launch!
Kwindla from Daily (
@kwindla) is on now.
Open-weights support model with NVIDIA and Modal.
Introducing PhoneLLM, an open model for voice agents.
GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost.
For voice agents, we need models that are both very low latency and very good at tool calling and instruction following.
There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem.
For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model.
But if you need your agent to respond at voice conversation speed, you can't use thinking models.
PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled.
The results are really good: accurate tool calling and concise, on-topic responses in long conversations.
And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-)
But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines.
You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today.
More details about this model, including weights on
@huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...
Show more
Some weekend profiling and visualization experiments.
I spend a lot of time on benchmarks, these days. I help maintain several public benchmarks, work with customers on private benchmarks, and talk to partners about how we can all build benchmarks that are useful for evaluating all of the components of our voice agents.
This benchmark by
@_josh_meyer_ and team is carefully constructed to measure the things that are critical to voice agent success, in a scenario that closely matches the real-world use case we see in enterprise deployments.
A few notes, from my perspective, in praise of things I really appreciate about this benchmark work:
- A good voice agent benchmark tests multi-turn conversation, with multiple defined tools and several tool calls during the conversation.
- Users care about both overall conversation success and the turn-to-turn naturalness and latency during a conversation. It's important to do a good job quantitatively scoring both of these. You can aggregate them into a single score, or break out the components (overall success, voice-to-voice latency, unwanted interruptions, etc).
- Teams building agents want to see comparisons between cascaded pipelines (TTS -> LLM -> STT) and speech-to-speech models. A good benchmark can judge both architectures on the same tasks.
- Cost matters. Most teams think about cost as cost-per-minute. Translating token counts to cost-per minute is note easy. And even forecasting token counts is hard to do in the abstract. A realistic benchmark can generate token and caching data that helps pin down the cost of different models and architectures.
- Tool call failures (and other types of agent failures, too) can be subtle. When you're working on a benchmark, it usually takes a fair amount of iteration and scrutinizing the transcripts to get to the point where your LLM-as-a-judge is doing a good job judging.
- Open source benchmarks and detailed write-ups allow other people to learn from the hard work that goes into benchmarks, and force you to stand behind the choices and implementation quality of your benchmarking work. Benchmarks are hard! It's good to build them in the open.
Show more
Jon's much-beloved Pipecat logo goes Disney.
Oh, are we animating our 'little guys'? The
@pipecat_ai mascot was made for this moment...
ThursdAI is live, covering the big week in open source model releases:
Deepgram has been shipping realtime speech models that deliver both very high accuracy and very low latency since before we had LLMs we could use to build realtime voice agents!
Scott played me clips of their new Flux TTS model while they were training it, and I've talked a lot to the Deepgram team about what we want a next-generation text-to-speech model to be good at.
For enterprise use cases, we need our models to vocalize text very, very reliably. The model can't skip numbers in an alpha-numeric string, can't mispronounce common words when they appear in an unusual sequence like a brand name or a medical term, and can't stumble over acronyms. (All of these are widespread issues with many text-to-speech models.)
We also want voice models to be highly expressive and steerable. And voice agents sessions are long, multi-turn conversations. We need both pronunciation choices and expressiveness to be consistent across conversation turns.
Deepgram delivered on all of these must-haves with their new Flux TTS model. I'm really impressed.
Show more
Congrats to
@DeepgramAI on the GA launch of Flux TTS, now live in Pipecat 🎉
@JonPTaylor looks at how Flux TTS delivers a more consistent voice experience. Flux reads the whole conversation, not just the next line — adaptive tone, consistent pronunciation, clean interruption handling natively (no SSML markup, no style tags!) all at sub-200ms latency.
➡️
Show more
This is living in the future.
Developing a robot by chatting with a robot
I've been thinking lately about how the new software we're writing fits into the 80-year history of modern computing.
As
@stevesi wrote earlier today, it feels like computing is new again. We don't know what we don't know about building software that uses LLMs to their full potential.
Lots of lessons from the past are relevant to the building blocks, abstractions, and tools we're building for the LLM era. But there's also lots of brand new stuff to figure out. LLMs are qualitatively different from the components we've built with before.
I wrote up some notes about this for
@aiDotEngineer World's Fair last month ... what the big challenges were in each decade from the 1950s to now.
Also, some stories about working on the UI and spatial computing tech stack made famous by Minority Report. And about hanging out with Robert Downey, Jr during Iron Man production, talking about Jarvis (Iron Man's snarky, capable, AI assistant).
We can now actually build Jarvis.
Show more