# Gradium — The most accurate multilingual text-to-speech, by the numbers

- Company: Gradium (gradium.ai)
- Announced: 2026-04-29
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://gradium.ai/blog/word-error-rate-evaluations
- Record: https://forck.live/items/18321-the-most-accurate-multilingual-text-to-speech-by-the-numbers
- Subject: Gradium TTS / Phonon

As we ship and improve our text-to-speech models, it is important to measure one key aspect of such models: pronunciation accuracy. Our system should pronounce exactly what is input by the user. How this is typically done in research work is by first transcribing the produced audio and comparing this transcript to the original text input. Normalization There exists no one-to-one mapping between a text and the corresponding speech. Many texts exist that could map to the same speech audio. One way to alleviate the problem is to try and normalize both the input and transcribed text. For instance, converting to lowercase, removing punctuation, converting numbers to an all-digit form... A popular option is to use the Whisper English normalizer . Extending it to other languages can be challenging though, requiring fine linguistic knowledge. The parsing itself can be daunting. Parsing tools such as parser combinators can come to the rescue, such as with Kyutai's tts_longeval French normalizer . That said, even a clean parser-combinator design hits a ceiling. Some of the things that are hard or genuinely impossible to handle robustly: Homophones. …

---

Record: https://forck.live/items/18321-the-most-accurate-multilingual-text-to-speech-by-the-numbers
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
