Evaluations are difficult and vague for all generative models, and benchmarks only capture a small slice. Our blog post dives into the nuances for TTS
"Is this TTS model good?" gets harder to answer as models improve.
"Good" is at least five axes: correctness, naturalness, contextual correctness, robustness, and most evals only capture the first.
We wrote up the failure modes that make TTS eval hard: cartesia.ai/blog/is-this-t…





