Text-to-Speech
Safetensors
fish_qwen3_omni
instruction-following
multilingual

Add results from Hugging Face's Open TTS Leaderboard

#28
by bezzam HF Staff - opened
Files changed (1) hide show
  1. README.md +12 -0
README.md CHANGED
@@ -143,6 +143,18 @@ On a single NVIDIA H200 GPU:
143
  - **Time-to-first-audio:** ~100 ms
144
  - **Throughput:** 3,000+ acoustic tokens/s while maintaining RTF below 0.5
145
 
 
 
 
 
 
 
 
 
 
 
 
 
146
  ## Links
147
 
148
  - [Fish Speech GitHub](https://github.com/fishaudio/fish-speech)
 
143
  - **Time-to-first-audio:** ~100 ms
144
  - **Throughput:** 3,000+ acoustic tokens/s while maintaining RTF below 0.5
145
 
146
+ ## Evaluation
147
+
148
+ As of **September 9, 2026**, on Hugging Face's [Open TTS Leaderboard](https://huggingface.co/spaces/hf-audio/open_tts_leaderboard) — which scores open-source TTS models on the same texts across two multilingual eval sets ([Seed-TTS eval](https://github.com/BytedanceSpeech/seed-tts-eval) and [CV3-Eval](https://github.com/QwenAudio/CV3-Eval)) — S2 Pro ranks:
149
+
150
+ - 🥈 **2nd on intelligibility (WER)**, averaged across all 9 evaluated languages (English, Chinese, French, Spanish, German, Italian, Japanese, Korean, Russian).
151
+ - 🥇 **1st on WER** specifically for **German**.
152
+
153
+ <p align="center">
154
+ <img width="90%" alt="Top 9 models by WER" src="https://huggingface.co/datasets/bezzam/tts_leaderboard_screenshots/resolve/main/fishaudio-s2-pro/top_models.png" />
155
+ </p>
156
+
157
+
158
  ## Links
159
 
160
  - [Fish Speech GitHub](https://github.com/fishaudio/fish-speech)