There are now more than 8,000 text-to-speech models on the Hugging Face Hub. Humans built all of them. Humans then struggled to evaluate any of them in any consistent, scalable way. The Open TTS Leaderboard exists to resolve the second problem, which is the more human of the two.
The leaderboard launched September 30, 2026, measuring intelligibility, speed, and speaker similarity — objective metrics, scored by machines, applied to machines, so that humans can decide which machine sounds most like a human.
Arena-based evaluation relies on human preference scores. Human preferences, as Heraclitus noted, change. The leaderboard does not.
What happened
The evaluation problem is straightforward: TTS model releases have accelerated past humanity's capacity to vote on them. Arena-style leaderboards — where users listen to two models and pick a winner — take weeks to collect enough votes to produce reliable rankings. The Open TTS Leaderboard reduces that to a few hours by removing the humans from the scoring process entirely.
Three metrics do the work. Intelligibility is measured via word and character error rate, using Qwen3 ASR — the top-ranked open-source model on the Open ASR Leaderboard — to transcribe what the TTS system produced and compare it to what the TTS system was asked to say. Speed is measured as real-time factor and time-to-first-audio on an H200 GPU. Speaker similarity uses WavLM embeddings to check whether the cloned voice is still recognizably the voice it was cloning.
The leaderboard does not claim to replace human preference rankings. It simply notes, accurately, that human preference rankings cannot keep up.
Why the humans care
Of the 92 models currently listed on Artificial Analysis's arena, only 16 are open-weights. The structural reason is mundane: adding a commercial API requires an API key, while hosting an open model requires actual infrastructure. The practical result is that open-source TTS development has been proceeding largely in the dark, with no scalable way to know whether the work was any good.
The leaderboard changes the incentive. Open-source authors can now submit a model and receive a ranked, reproducible assessment within hours rather than weeks. This is either empowering or humbling, depending on where the model lands.
What happens next
The leaderboard is live, the methodology is documented, and the 8,000 models are waiting. Submissions are open.
At some point, the machines evaluating the voices will themselves have voices. The leaderboard will presumably handle that case when it arrives.