From the source
Phonon, our 100M-parameter on-device Text-To-Speech model, produces up to 3.5x fewer word errors than NVIDIA's Magpie TTS in French, German, and Spanish, at 3.6x fewer parameters, and it leads the other on-device voice-cloning models, NeuTTS Nano and the equally compact 100M-parameter Pocket TTS, on both word error rate and speaker similarity in every language we tested.
Phonon clones a reference voice across all five languages it supports: English, French, German, Spanish, and Portuguese.
Phonon is built on Continuous Audio Language Models [1] with flow-matching for waveform generation, first announced in April .
Our previous benchmark posts established Phonon as the leader on English.
A model that tops the English charts can still mangle French liaison, German compound nouns, or Spanish stress patterns, so this post extends the evaluation to the other four languages: French, German, Spanish, and Portuguese, comparing Phonon against Pocket TTS, Magpie, and NeuTTS Nano wherever a baseline exists.
Phonon at a glance How we evaluate multilingual Text-To-Speech quality …






