Modern TTS models can sound great — and still fail badly on pacing, pauses, and prosody.
We adapted DPO + GRPO to flow-matching models to tackle the tail end of TTS behavior:

Article
Teaching Flow-Matching Text-to-Speech Models with RL
Written by Rohan Siva (@_rsiva) and Cyrus Asgari (@cyrusasg)
Good demos, unreliable distributions
Modern text-to-speech (TTS) systems can sound remarkably natural. But average quality hides the...



