# Sesame — Crossing the uncanny valley of conversational voice

- Company: Sesame (sesame.com)
- Announced: 2025-02-27
- Category: new-model
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://www.sesame.com/blog/crossing-the-uncanny-valley-of-voice
- Record: https://forck.live/items/8046-crossing-the-uncanny-valley-of-conversational-voice
- Subject: Maya / CSM
- Models affected: CSM

Sesame introduced the Conversational Speech Model (CSM), a multimodal text and speech model that uses two autoregressive transformers. Key contributions include a single-stage architecture and an evaluation suite for contextual capabilities.

## Evidence

Verbatim from https://www.sesame.com/blog/crossing-the-uncanny-valley-of-voice:

> we introduce the Conversational Speech Model (CSM), which frames the problem as an end-to-end multimodal learning task using transformers. It leverages the history of the conversation to produce more natural and coherent speech. There are two key takeaways from our work. The first is that CSM operates as a single-stage model , thereby improving efficiency and expressivity. The second is our evaluation suite , which is necessary for evaluating progress on contextual capabilities

---

Record: https://forck.live/items/8046-crossing-the-uncanny-valley-of-conversational-voice
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
