From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
The pace of open-source text-to-speech (TTS) model releases has been incredible.
On the Hugging Face Hub (as of Sep 30, 2026) there are more than 8K TTS models available 🚀 Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized.
The gold standard is human preference scores such as MOS or MUSHRA (more on metrics ).
To this end, several arena-based leaderboards have established themselves as useful reference points for the community: These arenas compare models by presenting users with TTS outputs from two models, and asking them to choose one over the other.
After collecting a sufficient number of votes, an Elo score is computed to rank models, typically with the Bradley–Terry model (see Voice Arena methodology ).
While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases .
This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena .
…