註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Rohan Siva
@_rsiva
research @decagonai ml & robotics research @utaustin
加入 August 2025
29 正在關注    46 粉絲
tts naturalness is hard. sft has its limits, but rl is tricky for flow-matching models since they lack the log-probs dpo and grpo rely on. excited to share work on adapting both methods for TTS with @cyrusasg at @DecagonAI!
顯示更多
Modern TTS models can sound great — and still fail badly on pacing, pauses, and prosody. We adapted DPO + GRPO to flow-matching models to tackle the tail end of TTS behavior:
顯示更多