Hugging Face Can Now Rank Every Open-Source Voice AI in Hours
More than 8,000 text-to-speech models exist, but testing them has always taken too long. A new automated scorecard changes that, and puts open-source voices on equal footing for the first time.

Key points
- As of 30 September 2026, the Hugging Face Hub hosts more than 8,000 text-to-speech models, yet only 16 of the 92 models on the arena-style leaderboard Artificial Analysis are open-weights.
- The new Open TTS Leaderboard cuts evaluation time from several weeks to a few hours by using automated metrics instead of human votes.
- Kokoro-82M and Supertonic-3 lead on English accuracy; OmniVoice and Fun-CosyVoice3-0.5B top multilingual performance.
- A dedicated streaming tab ranks models by time-to-first-audio, the delay a user actually feels when talking to a voice assistant, with Kyutai's Pocket-TTS performing best across both GPU and CPU hardware.
- The leaderboard is built specifically to surface open-source models that existing voting-based systems routinely overlook.
Text-to-speech AI, software that turns written words into spoken audio, is growing faster than anyone can keep score. Hugging Face, the platform where researchers share AI models, now hosts more than 8,000 such models. The standard way to rank them has been to ask real people to listen to pairs of outputs and pick a winner, a process that takes weeks and quietly favours commercial products over free, open-source alternatives.
The new Open TTS Leaderboard, published by Hugging Face, swaps slow crowd votes for automated measurements that run in hours.
What does the leaderboard actually measure?
Three things: accuracy, speed, and fidelity. Accuracy is measured by feeding the AI's audio output back through a speech-recognition model, then counting how many words came out wrong. The lower the word error rate (WER), the better. Speed splits into two figures: how fast a model processes a batch of audio on a powerful H200 GPU (a specialised chip used for AI computing), and how quickly the first slice of audio arrives when a user sends a single request, which is the number that matters for voice assistants and live calls. Fidelity measures how closely a cloned voice matches the original recording.
It's the first leaderboard in this space to break out a separate tab for streaming performance, directly relevant to any app where a person speaks and expects a quick reply. Kyutai's Pocket-TTS ranked best there on both GPU and ordinary CPU hardware.
Why does it matter that open-source models were being left behind?
Practical economics, mostly. Adding a commercial model to an existing arena-style leaderboard takes one API key. Hosting an open-source model requires the arena operator to pay for servers and do the technical work themselves. Only 16 of the 92 models ranked on Artificial Analysis were open-weights as of 30 September 2026. This leaderboard flips that priority, evaluating open models first.
Voice AI isn't a research curiosity anymore. When Pocket FM reported on 10 September that 99% of its new audio content is AI-generated and revenue has reached $500 million, it underlined how fast this technology is moving into products people use daily. A fast, objective way to rank the underlying models is overdue. For developers, the practical upshot is that they can now compare dozens of open-source options in hours rather than weeks, which should push better, cheaper voice AI into consumer products faster.
The leaderboard doesn't claim to replace human judgement. Automated word-error-rate testing can't score whether a voice sounds warm or expressive, and the team says as much. A "Listen" tab lets visitors play back sample audio and cast their own votes, logged to Hugging Face accounts to filter out bots. Those votes may feed back into rankings over time.
Evaluation scripts will be open-sourced shortly, following the model of the existing Open ASR Leaderboard, so the research community can propose new datasets and metrics directly through GitHub.
Common questions
Does this replace the existing TTS Arena or Artificial Analysis rankings?
No. The leaderboard fills a gap rather than competing directly. Its automated metrics are fast and scalable; human preference arenas remain the gold standard for naturalness and expressiveness. The two approaches are meant to work alongside each other.
Which models performed best overall?
On English accuracy, Kokoro-82M, Supertonic-3, and Fish Audio S2-Pro led. For multilingual use, OmniVoice and Fun-CosyVoice3-0.5B topped the rankings, with Fish Audio S2-Pro also strong across languages. For real-time streaming response, Kyutai's Pocket-TTS came out ahead on both GPU and CPU hardware.
Should you worry that automated scores miss the point?
Somewhat. WER tells you a model pronounced words correctly; it won't tell you the voice sounded human. The "Listen" tab exists precisely because the team knows numbers only go so far. Watch whether the community vote data eventually shifts the rankings, because that's when this leaderboard will get genuinely interesting.



