Great work at @baseten running vLLM-Omni in production — open-source, production-grade, cost-efficient omni-modal serving 🎙️
Multi-stage audio, streaming multi-modal, real-time TTS — workloads where closed-source APIs have been the default.
→
We serve Qwen3-TTS on vLLM-Omni at $3 per 1M characters. That's 90% lower in cost than comparable closed-source TTS APIs.
Our engineers optimized a single-replica serving stack to get there. Details on the optimized stack and cost per concurrent stream here.