Add results from Hugging Face's Open TTS Leaderboard
#28
by bezzam HF Staff - opened
README.md
CHANGED
|
@@ -143,6 +143,18 @@ On a single NVIDIA H200 GPU:
|
|
| 143 |
- **Time-to-first-audio:** ~100 ms
|
| 144 |
- **Throughput:** 3,000+ acoustic tokens/s while maintaining RTF below 0.5
|
| 145 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 146 |
## Links
|
| 147 |
|
| 148 |
- [Fish Speech GitHub](https://github.com/fishaudio/fish-speech)
|
|
|
|
| 143 |
- **Time-to-first-audio:** ~100 ms
|
| 144 |
- **Throughput:** 3,000+ acoustic tokens/s while maintaining RTF below 0.5
|
| 145 |
|
| 146 |
+
## Evaluation
|
| 147 |
+
|
| 148 |
+
As of **September 9, 2026**, on Hugging Face's [Open TTS Leaderboard](https://huggingface.co/spaces/hf-audio/open_tts_leaderboard) — which scores open-source TTS models on the same texts across two multilingual eval sets ([Seed-TTS eval](https://github.com/BytedanceSpeech/seed-tts-eval) and [CV3-Eval](https://github.com/QwenAudio/CV3-Eval)) — S2 Pro ranks:
|
| 149 |
+
|
| 150 |
+
- 🥈 **2nd on intelligibility (WER)**, averaged across all 9 evaluated languages (English, Chinese, French, Spanish, German, Italian, Japanese, Korean, Russian).
|
| 151 |
+
- 🥇 **1st on WER** specifically for **German**.
|
| 152 |
+
|
| 153 |
+
<p align="center">
|
| 154 |
+
<img width="90%" alt="Top 9 models by WER" src="https://huggingface.co/datasets/bezzam/tts_leaderboard_screenshots/resolve/main/fishaudio-s2-pro/top_models.png" />
|
| 155 |
+
</p>
|
| 156 |
+
|
| 157 |
+
|
| 158 |
## Links
|
| 159 |
|
| 160 |
- [Fish Speech GitHub](https://github.com/fishaudio/fish-speech)
|