From the source
Voice cloning is disabled for this comparison.
Evaluated May 2026.
Phonon, our 100M-parameter on-device Text-To-Speech model, now reaches a 1.00% word error rate on the Seed-TTS English benchmark, outperforming NeuTTS Air (552M), KaniTTS2 (450M), and NeuTTS Nano (229M).
With voice cloning disabled and a fixed high-quality voice, Phonon drops to 0.83% WER, ahead of Kokoro and Magpie.
Phonon is based on Continuous Audio Language Models with flow-matching for waveform generation, first announced in April .
It is in private beta.
You can request access here .
On-device Text-To-Speech enables deployment scenarios that cloud APIs cannot serve: offline voice agents in vehicles and remote equipment, latency-sensitive products where a network round trip is unacceptable, privacy-sensitive applications in healthcare and consumer hardware where audio cannot leave the device, and high-volume applications where per-request API costs are prohibitive.
Phonon runs entirely on the edge, removing the network from the voice pipeline.
This post compares the updated Phonon against the version from our previous benchmark post .
…





