Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
Sesame introduced the Conversational Speech Model (CSM), a multimodal text and speech model that uses two autoregressive transformers. Key contributions include a single-stage architecture and an evaluation suite for contextual capabilities.
From the source
we introduce the Conversational Speech Model (CSM), which frames the problem as an end-to-end multimodal learning task using transformers. It leverages the history of the conversation to produce more natural and coherent speech. There are two key takeaways from our work. The first is that CSM operates as a single-stage model , thereby improving efficiency and expressivity. The second is our evaluation suite , which is necessary for evaluating progress on contextual capabilities
sesame.com