# Cartesia — TTS Model vs. Voice: They're Not the Same

- Company: Cartesia (cartesia.ai)
- Announced: 2026-09-10
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://www.cartesia.ai/blog/tts-model-vs-voice
- Record: https://forck.live/items/10465-tts-model-vs-voice-they-re-not-the-same
- Subject: Sonic / Ink

Cartesia publishes a guide explaining the distinction between text-to-speech models and voices, covering fixed-speaker models, multi-speaker models, zero-shot voice cloning, and professional voice cloning. The post describes how voices function as inputs that shape model output, and how Cartesia's Sonic model and voice-cloning services handle speaker characteristics.

## Evidence

Verbatim from https://www.cartesia.ai/blog/tts-model-vs-voice:

> In TTS, voice is an input often separate from the model, and it shapes the model's response. And the model's response - audio - is much more complex than "mere" text. Models that generate speech had to know what should be said, how it should be said, at what pace, how loud, what sort of intonation, prosody, emphasis…all the many characteristics and metrics for a "good voice".

---

Record: https://forck.live/items/10465-tts-model-vs-voice-they-re-not-the-same
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
