From the source
Update — May 2026: Phonon now reaches 1.00% WER on the same benchmark (down from 1.48%) and 0.83% WER with a fixed voice, beating Kokoro and Magpie.
See the latest results in Phonon update: 1.00% WER on Seed-TTS .
Phonon is our on-device text-to-speech model: building on continuous audio language models Phonon packages high-fidelity TTS into ~100M parameters, 6x real-time on a single MacBook CPU core, small enough to run in a browser and can reproduce any voice, style and accent.
In this post we put it through a standard evaluation and report the results.
On the Seed-TTS English benchmark, Phonon achieves a 1.48% word error rate and 56.37% speaker similarity, outperforming models 2x to 5x its size.
How we evaluate on-device text-to-speech quality Phonon supports on-device voice cloning: given a 10-second sample of someone's voice, it generates new speech that matches that speaker's tone, accent, and cadence.
This is what makes a lightweight, on-device TTS model useful in practice.
A navigation app can speak in the user's preferred voice.
A language learning tool can maintain a consistent teacher voice across sessions.
…





