multimodal inference feels pretty underexplored, especially when you move beyond standard transformer architectures. fun to share what we learned when scaling our text to speech model to 10x throughput
顯示更多
Fast text serving ≠ fast speech.
Here's how we got our time to first audio under 30ms with nearly 10x more audio throughput.