From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
TTS models and voices are separate components; voice shapes model output characteristics.
Cartesia publishes a guide explaining the distinction between text-to-speech models and voices, covering fixed-speaker models, multi-speaker models, zero-shot voice cloning, and professional voice cloning.
The post describes how voices function as inputs that shape model output, and how Cartesia's Sonic model and voice-cloning services handle speaker characteristics.
From the source
In TTS, voice is an input often separate from the model, and it shapes the model's response. And the model's response - audio - is much more complex than "mere" text. Models that generate speech had to know what should be said, how it should be said, at what pace, how loud, what sort of intonation, prosody, emphasis…all the many characteristics and metrics for a "good voice".
cartesia.ai