From the source
Thank you for the valuable feedback on the drafts: Chung-Ming Chien, Moritz Boehle, Richard Hladík, Eugene Kharitonov, Patrick Perez, and Tom Sláma.
I’d also like to thank the rest of the Kyutai team for the the research discussions without which this article could not exist.
The plan: sandwich a language model in an audio encoder/decoder pair (=neural audio codec), allowing it to predict audio continuations.
As of October 2025, speech LLMs suck.
Many LLMs have voice interfaces, but they usually work by transcribing your speech, generating the answer in text, and using text-to-speech to read the response out loud.
That’s perfectly fine in many cases (see Unmute ), but it’s a wrapper, not real speech understanding.
The model can’t hear the frustration in your voice and respond with empathy, it can’t emphasize important words in its answer, it cannot sense sarcasm, and so on.
Yes, there are LLMs ( Gemini , ChatGPT ’s Advanced Voice Mode, Qwen , Moshi ) that understand and generate speech natively.
But in practice, they’re either not as smart, or they behave like text model wrappers.
…