Speech
Speech recognition and synthesis on EcoHash. WER, real-time factor and time to first audio, all measured by us end-to-end.
Text-to-Speech • Updated • 11.4M • • 7.11kNote 120 ms to first audio, streaming, $1 per 1M tokens. At 82M parameters it starts speaking inside the window a listener still reads as thinking rather than broken.
Qwen/Qwen3-TTS-12Hz-1.7B-Base
2B • Updated • 3.62M • 540Note 2621 ms to first audio, streaming, $2 per 1M tokens. Richer prosody and broader language coverage; in an interactive agent you can hear what the latency costs.
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.37M • • 3.41kNote WER 4.37%, RTFx 59, $0.006 per audio minute. About 30% more throughput than v3 for roughly 0.7 points of WER - the right trade for interactive use, the wrong one for an archive you transcribe once.
FunAudioLLM/Fun-ASR-Nano-2512
Automatic Speech Recognition • Updated • 2.35k • 233Note WER 3.83%, RTFx 21, $0.05 per 1M tokens. More accurate than Whisper Turbo and nearly three times slower. Neither model dominates, which is why we built a Space that runs both on your own clip.
Speech Recognition Benchmark
🚀Whisper v3 Turbo vs Fun-ASR-Nano, measured side by side
Note Live demo. Runs Whisper Large v3 Turbo and Fun-ASR-Nano on the same clip at once and puts transcripts, real-time factor and cost side by side.
Text-to-Speech Studio
🚀Kokoro-82M and Qwen3-TTS with measured latency and cost
Note Live demo. Both TTS models with switchable voices, reporting real-time factor and cost per second of audio for every generation.
Qwen/Qwen3-ASR-1.7B
Automatic Speech Recognition • 2B • Updated • 1.53M • • 1.14kNote WER 3.28% on LibriSpeech test-clean, RTFx 360, $0.05 per 1M tokens. The most accurate and the fastest of the speech-to-text models we serve - measured by us end-to-end.