Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
The evaluation of text-to-speech models involves multiple dimensions including correctness, audio quality, naturalness, and robustness, and metrics like word error rate saturate as models improve.
From the source
This multidimensionality is why TTS evaluation is tricky and as models improve, the evaluation shifts away from obvious defects like mispronounced words, garbled audio, or background noise towards more nuanced aspects.
cartesia.ai