multimodal inference doesn't fit the mold of standard text autoregressive generation. text-to-speech models feature different architectures, different states, and different batch shapes. we rebuilt our tts serving around that and simultaneously reduced our time to first audio while improving throughput by several fold.