🔍 Read the full analysis: Open TTS Leaderboard: Scalable Evaluation For Multilingual Text-to-Speech And Voice Cloning on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has introduced the Open TTS Leaderboard, which compares text-to-speech models using automated measures of speech accuracy, generation speed and speaker similarity. The project says results can be produced in hours, but the scores do not assess naturalness or replace human listening and preference judgments.
Hugging Face has launched the Open TTS Leaderboard, a system for comparing text-to-speech models on speech accuracy, generation speed and speaker similarity. It is intended to make comparisons across a growing field of models faster to produce, while the project cautions that its automated scores do not measure naturalness or listener preference.
The leaderboard checks speech accuracy by comparing transcripts of generated audio with the text prompts, using Qwen3 automatic speech recognition. It reports word error rate and character error rate. Its speed measures include offline generation performance, reported as inverse real-time factor on an H200 GPU, and streaming responsiveness, measured as time to first audio on both an H200 GPU and a CPU.
For voice cloning, the system reports a speaker-similarity score based on WavLM embeddings from generated speech and reference audio. The default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval. Users can select other languages and switch to a voice-cloning view, according to the project description.
Hugging Face identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leading models. These are results on the listed measures, not an overall judgment of which systems sound best.
Faster Comparisons Across TTS Models
The new leaderboard offers developers and researchers a quicker way to screen text-to-speech systems before investing time in more detailed tests. Hugging Face says an evaluation using the automated metrics can take a couple of hours, compared with weeks for arena voting. That speed may help users check multiple dimensions, including accuracy and delay, as new models appear.
Those dimensions matter for different uses. For a voice agent, time to first audio can affect how responsive an interaction feels. For an application that must preserve a particular voice, speaker similarity may be relevant. Meanwhile, speech error rates offer one way to compare whether generated words can be recognized by an ASR system. The measures can expose tradeoffs rather than provide a single winner: performance on one score does not establish that a model is fastest, most natural or best suited to every task.
The leaderboard’s value will depend on how well its automated measures reflect real-world performance across languages, voices and speaking styles. Hugging Face presents it as a source of structured comparisons, with a separate listening feature for people to judge outputs directly.
As an affiliate, we earn on qualifying purchases.
A Different Route to TTS Rankings
Text-to-speech models turn written text into spoken audio, and the number of available systems makes consistent comparisons difficult. Hugging Face says its Hub contained more than 8,000 TTS models as of September 30, 2026, while comparisons were fragmented and could take time to update.
Existing reference points cited by Hugging Face include TTS Arena v2, Artificial Analysis and Voice Arena. Arena-style rankings ask listeners to compare two outputs and vote; rankings may use Elo scores calculated with a Bradley–Terry model. That approach captures subjective preferences, but collecting votes takes time and depends on enough participants. Hugging Face argues that hosting the models being tested also creates practical barriers for such comparisons.
The company says open-weight models were underrepresented in some existing rankings. As of September 30, it counted 16 open-weight models among 92 on Artificial Analysis and reported a similar skew on Voice Arena. That count is Hugging Face’s characterization of those rankings; it does not by itself establish why particular models were absent.
“The Open TTS Leaderboard does not replace human preference ranking.”
— Hugging Face
multilingual text-to-speech speaker
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Published Scores
The supplied announcement does not provide the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not establish how closely the automatic scores track human judgments across languages, accents and speaking styles.
Error rates estimate intelligibility through an automatic speech-recognition system; they are not direct ratings of how natural or expressive a voice sounds. Speaker similarity estimates how closely generated audio matches a reference voice, but does not settle whether listeners prefer the result. The project’s own description says those qualities require listening and human preference judgments.
Coverage also varies by language. The announcement says Chinese, Japanese and Korean use character error rate. For languages beyond English and Chinese, it says Seed TTS Eval has no audio and scores come from CV3 Eval alone. The announcement does not specify how often rankings will be refreshed, how updated models or evaluation data will be handled, or when community votes might affect rankings.
As an affiliate, we earn on qualifying purchases.
Listening Tests and Community Votes
Users can open the leaderboard’s Listen tab to compare generated audio, select a language and dataset, and choose whether to compare voice-cloning outputs. They can submit preferences after signing in with a Hugging Face account; the project says sign-in is intended to limit spam and bot submissions.
Hugging Face says community votes may be incorporated into the leaderboard as feedback accumulates, but has not announced a schedule or threshold for doing so. For now, the published automated metrics and the listening feature serve distinct roles. The next useful signals will be fuller score and methodology details, regularity of updates, and whether community preferences are eventually added to the rankings.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the Open TTS Leaderboard measure?
It reports speech error rates, offline generation speed, streaming time to first audio and speaker similarity for voice cloning. Users can also listen to generated outputs and submit preferences.
Does the leaderboard identify the best-sounding TTS model?
No. Its automated measures do not directly rate naturalness, expressiveness or listener preference. Hugging Face says the scores should be complemented by listening tests.
Which models does Hugging Face list as strong multilingual options?
The project names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. That description refers to leaderboard results and is not a general verdict on audio quality.
How can users contribute listener preferences?
Users can compare audio in the Listen tab and vote after signing in to a Hugging Face account. The project says votes may be used later, but has not specified when or how they will affect rankings.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
