tts naturalness is hard. sft has its limits, but rl is tricky for flow-matching models since they lack the log-probs dpo and grpo rely on.
excited to share work on adapting both methods for TTS with @cyrusasg at @DecagonAI!
Modern TTS models can sound great — and still fail badly on pacing, pauses, and prosody.
We adapted DPO + GRPO to flow-matching models to tackle the tail end of TTS behavior: