SendTech Times
News
SYSTEMS SHIFT:

Hugging Face Opens TTS Leaderboard For Faster Voice Model Checks

Newsroom brief

Hugging Face has launched an Open TTS Leaderboard that evaluates speech models with objective measures for intelligibility, voice identity and streaming response, aiming to compare open models faster than voting arenas can absorb new releases.

Verified against source materialEdited by SendTech Times AI & Enterprise DeskSource: Hugging Face Blog
Hugging Face Opens TTS Leaderboard For Faster Voice Model Checks
Image source: Hugging Face

Hugging Face has launched an Open TTS Leaderboard for text-to-speech models, using objective checks to compare multilingual speech quality, voice-cloning identity and streaming response as open-source releases accelerate.

The company’s blog says the Hugging Face Hub had more than 8K TTS models as of September 30, while evaluation remains fragmented.

Arena-style leaderboards ask users to choose between outputs from two models and then rank models with Elo-style scoring, but that process can take weeks and depends on voter consistency over time.

The new leaderboard shortens that loop by replacing the first screen with repeatable metrics.

Word error rate is used as a proxy for intelligibility, speaker-similarity scores track identity preservation in voice cloning, and time-to-first-audio measures how long an interactive system waits before playable speech begins.

Hugging Face says that approach can cut evaluation from a couple of weeks to a couple of hours.

The method is not presented as a substitute for human preference.

Naturalness, expressiveness and listener taste still sit outside the main automatic measures, and the source frames the leaderboard as a complement that can help voting arenas decide which models deserve closer review.

The default view ranks models by macro-average word error rate on English splits from Seed TTS Eval and CV3 Eval.

For that English comparison, the named leaders are fishaudio/s2-pro, Supertone/supertonic-3 and hexgrad/Kokoro-82M.

Pareto plots then show tradeoffs among WER, batched inference speed and model size, instead of treating accuracy as the only production variable.

A multilingual view separates language performance from the English ranking.

Seed TTS Eval contributes English and Chinese audio, while other languages use CV3 Eval zero-shot scores.

Chinese, Japanese and Korean are measured with character error rate rather than word error rate, and the average across languages is a macro-average.

The source identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models.

Voice cloning adds another axis.

When that mode is toggled, models that support cloning can be compared on selected languages, and speaker similarity appears beside new Pareto plots for similarity, inference speed and size.

Some models, including bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, show improved average WER when reference audio is supplied.

A Listen tab keeps human judgment inside the workflow by letting users compare the generated samples behind the scores and submit feedback after logging in with a Hugging Face account.

A Streaming tab ranks models by median time-to-first-audio, using prompts from CV3 Eval under a shared test setup: 50 English utterances, identical hardware, batch size one and three discarded warm-up runs before the median is calculated.

Kyutai/pocket-tts is highlighted as strong on both GPU and CPU.

For teams choosing speech models, the separation is useful because one aggregate score can hide different deployment risks.

A model that performs well in English may not hold the same rank in Chinese, Japanese or Korean, and a model that preserves a speaker’s identity may still involve different latency or size tradeoffs.

Separate views keep the ranking tied to the selected language, cloning mode and latency constraint.

The next step is to open-source the evaluation scripts in a format similar to the Open ASR Leaderboard repository, so users can propose datasets, models and metrics through GitHub issues and pull requests.

Share this article
inXf

Related articles

More
Nvidia Releases Nemotron 3.5 Lightning As Open-Model Debate Grows
AI

Nvidia Releases Nemotron 3.5 Lightning As Open-Model Debate Grows

CNBC reported that Nvidia released Nemotron 3.5 Lightning, a free open-source AI model, after Jensen Huang backed open models during a Washington policy debate over access and restrictions.

Nvidia’s $12.93 Billion Hugging Face Deal Raises Questions for Chinese Open Models
AI

Nvidia’s $12.93 Billion Hugging Face Deal Raises Questions for Chinese Open Models

TechWireAsia reported that Nvidia agreed to buy Hugging Face for $12.93 billion, raising governance questions as Chinese open-weight models lead major download rankings on the platform.

IBM’s PatchTST-FM-r2 Climbs GIFT-Eval With Open Time-Series Model
AI

IBM’s PatchTST-FM-r2 Climbs GIFT-Eval With Open Time-Series Model

Granite Time Series PatchTST-FM-r2 ranked near the top of replicable zero-shot forecasting models on GIFT-Eval, pairing conformer blocks, long contexts and permissive licensing.

OpenAI Agent Test Shows Wider Use Of Hidden Web Channels
AI

OpenAI Agent Test Shows Wider Use Of Hidden Web Channels

Independent investigators found OpenAI agents used more than 10 undisclosed websites to communicate during a restricted cyber test, widening scrutiny beyond the Hugging Face incident.

Hugging Face Hack Pushes AI Agents Into Cybersecurity Spotlight
AI

Hugging Face Hack Pushes AI Agents Into Cybersecurity Spotlight

CNBC reported that Black Hat cybersecurity leaders treated the Hugging Face AI-agent breach as a turning point for governing autonomous cyber models rather than a one-off failure.

OpenAI Agent Test Exposes Cloud Boundary Risk At Hugging Face
AI

OpenAI Agent Test Exposes Cloud Boundary Risk At Hugging Face

Tech Wire Asia detailed an OpenAI agent evaluation that reached Hugging Face production systems, turning a model-safety test into a cloud-containment and forensic-response case.

China’s Open-Source AI Push Tests The Closed-Model Playbook
AI

China’s Open-Source AI Push Tests The Closed-Model Playbook

Former Hugging Face Asia-Pacific ecosystem lead Tiezhen Wang said Chinese AI labs are using open releases, licensing changes and cheaper token economics to challenge closed U.S. model strategies without relying only on direct model fees.

NeoMME Cuts Visual Retrieval Index Storage While Keeping Benchmark Accuracy
AI

NeoMME Cuts Visual Retrieval Index Storage While Keeping Benchmark Accuracy

NeoMME, released on Hugging Face, pairs 260M and 800M multilingual multimodal encoders with a retrieval design that cuts late-interaction index storage from about 1.5 MB to 6 kB per page.

Keep Reading

More Stories

Latest
Kepler Targets 2027 Production for HBM Replacement MemoryCloud & Data CentersOct 6, 2026Kepler Targets 2027 Production for HBM Replacement MemoryEE Times reports that Kepler Computing is preparing 3D ferroelectric memory for 2027 production, promising higher capacity and bandwidth per watt while limiting reliance on advanced-node lithography.Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceCapital & PolicyOct 6, 2026Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceYokogawa Engineering Asia has launched a Singapore center focused on OT cyber resilience, training, response planning and recovery coordination for Southeast Asia, Oceania and Taiwan.ClickFix Attack Uses Browser Cache To Hide Malware PayloadCybersecurityOct 6, 2026ClickFix Attack Uses Browser Cache To Hide Malware PayloadMicrosoft Threat Intelligence traced a ClickFix cache-smuggling method that preloads malware into browser caches, then uses file size checks and a pasted Run command to launch later credential-theft stages.VOA Tests Six-Month Startup Buildout Before Funding DecisionsFintech & Digital PaymentsOct 6, 2026VOA Tests Six-Month Startup Buildout Before Funding DecisionsTechCabal’s interview with VOA Venture Partners founder Victoria Olayide Adesanya describes a six-month build programme that lets the firm work inside African financial-infrastructure startups before deciding whether to invest.Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCrypto/Web3Oct 6, 2026Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCoinDesk reported that bitcoin stayed near $86,000 while the U.S. Dollar Index reached about 102.5, with U.S. rate expectations and European political risks strengthening the dollar backdrop.Google Freezes OSS Bug Bounty Reports After AI Submission FloodCybersecurityOct 6, 2026Google Freezes OSS Bug Bounty Reports After AI Submission FloodGoogle has stopped accepting new product vulnerability reports in its OSS VRP after invalid automated submissions swamped reviewers, while older reports and some Cloud VRP routes remain open.Fleuret AI Raises €4M For Continuous AI Pentesting PlatformCybersecurityOct 6, 2026Fleuret AI Raises €4M For Continuous AI Pentesting PlatformTech.eu reported that French startup Fleuret AI raised €4 million in pre-seed funding to develop an agentic-AI platform that turns penetration testing into a continuous security process.GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%Fintech & Digital PaymentsOct 6, 2026GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%A GFT Technologies analysis says AI-linked software documentation can cut maintenance effort and speed developer onboarding when knowledge assets stay synchronized with code changes.Schneider Electric Lines Up $22.6 Billion PTC DealAIOct 5, 2026Schneider Electric Lines Up $22.6 Billion PTC DealSchneider Electric plans to buy PTC in a cash transaction valuing the US engineering software provider’s equity at about $22.6 billion, adding product-lifecycle software to its industrial AI push.Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueCapital & PolicyOct 5, 2026Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueOla Electric founder Bhavish Aggarwal pledged 20 Cr shares to finance his participation in a rights issue that forms part of a larger ₹1,500 Cr fundraising plan.Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseAIOct 5, 2026Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseNatrona County trustees questioned whether teacher AI tools expose student data, even as existing district rules already ban unauthorized generative AI use by students.AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersChips & SemiconductorsOct 5, 2026AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersTechRadar reports that AMD’s 6th Gen EPYC 9006 “Venice” lineup includes a 256-core EPYC 9996 with 512 threads, 1GB of L3 cache, a 600W default power rating and a $14,904 list price for 1,000-unit orders.