Hugging Face Opens TTS Leaderboard For Faster Voice Model Checks
Hugging Face has launched an Open TTS Leaderboard that evaluates speech models with objective measures for intelligibility, voice identity and streaming response, aiming to compare open models faster than voting arenas can absorb new releases.

Hugging Face has launched an Open TTS Leaderboard for text-to-speech models, using objective checks to compare multilingual speech quality, voice-cloning identity and streaming response as open-source releases accelerate.
The company’s blog says the Hugging Face Hub had more than 8K TTS models as of September 30, while evaluation remains fragmented.
Arena-style leaderboards ask users to choose between outputs from two models and then rank models with Elo-style scoring, but that process can take weeks and depends on voter consistency over time.
The new leaderboard shortens that loop by replacing the first screen with repeatable metrics.
Word error rate is used as a proxy for intelligibility, speaker-similarity scores track identity preservation in voice cloning, and time-to-first-audio measures how long an interactive system waits before playable speech begins.
Hugging Face says that approach can cut evaluation from a couple of weeks to a couple of hours.
The method is not presented as a substitute for human preference.
Naturalness, expressiveness and listener taste still sit outside the main automatic measures, and the source frames the leaderboard as a complement that can help voting arenas decide which models deserve closer review.
The default view ranks models by macro-average word error rate on English splits from Seed TTS Eval and CV3 Eval.
For that English comparison, the named leaders are fishaudio/s2-pro, Supertone/supertonic-3 and hexgrad/Kokoro-82M.
Pareto plots then show tradeoffs among WER, batched inference speed and model size, instead of treating accuracy as the only production variable.
A multilingual view separates language performance from the English ranking.
Seed TTS Eval contributes English and Chinese audio, while other languages use CV3 Eval zero-shot scores.
Chinese, Japanese and Korean are measured with character error rate rather than word error rate, and the average across languages is a macro-average.
The source identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models.
Voice cloning adds another axis.
When that mode is toggled, models that support cloning can be compared on selected languages, and speaker similarity appears beside new Pareto plots for similarity, inference speed and size.
Some models, including bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, show improved average WER when reference audio is supplied.
A Listen tab keeps human judgment inside the workflow by letting users compare the generated samples behind the scores and submit feedback after logging in with a Hugging Face account.
A Streaming tab ranks models by median time-to-first-audio, using prompts from CV3 Eval under a shared test setup: 50 English utterances, identical hardware, batch size one and three discarded warm-up runs before the median is calculated.
Kyutai/pocket-tts is highlighted as strong on both GPU and CPU.
For teams choosing speech models, the separation is useful because one aggregate score can hide different deployment risks.
A model that performs well in English may not hold the same rank in Chinese, Japanese or Korean, and a model that preserves a speaker’s identity may still involve different latency or size tradeoffs.
Separate views keep the ranking tied to the selected language, cloning mode and latency constraint.
The next step is to open-source the evaluation scripts in a format similar to the Open ASR Leaderboard repository, so users can propose datasets, models and metrics through GitHub issues and pull requests.




















