Automatic Speech Recognition
ASR models for multilingual transcription and streaming speech recognition, with notes on use cases and licenses.
Automatic Speech Recognition • 2B • Updated • 1.53M • • 1.15kNote Multilingual ASR: 30 languages and 22 Chinese dialects. Offline and streaming inference; streaming uses the vLLM backend. Apache-2.0.
nvidia/parakeet-tdt-0.6b-v3
Automatic Speech Recognition • 0.6B • Updated • 569k • • 1.17kNote Compact 0.6B multilingual ASR covering 25 European languages. Useful transcription baseline in the NeMo/Transformers ecosystem. CC-BY-4.0.
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.37M • • 3.41kNote Widely supported multilingual transcription baseline. Faster variant of Whisper large-v3. MIT.
kyutai/stt-1b-en_fr
Automatic Speech Recognition • 1.0B • Updated • 137Note Native streaming ASR for English and French. Produces punctuated transcripts while audio arrives. CC-BY-4.0.
nvidia/nemotron-3.5-asr-streaming-0.6b
Automatic Speech Recognition • 0.6B • Updated • 1.26M • • 1.16kNote Native cache-aware streaming ASR with configurable chunk sizes and punctuation. Multilingual coverage is tiered; some locales require fine-tuning. OpenMDW-1.1.
mistralai/Voxtral-Mini-4B-Realtime-2602
Automatic Speech Recognition • 4B • Updated • 1.89M • 998Note Native streaming multilingual ASR with a causal audio encoder and configurable transcription delay. Useful for latency/accuracy comparisons. Apache-2.0.
nvidia/canary-qwen-2.5b
Automatic Speech Recognition • 3B • Updated • 45.6k • 466Note English ASR with punctuation and capitalization. Separate LLM mode can process the transcript; raw-audio understanding is not retained in that mode. CC-BY-4.0.
FunAudioLLM/SenseVoiceSmall
Automatic Speech Recognition • Updated • 18k • 491Note Non-autoregressive ASR with emotion recognition and audio-event tags. Useful for richer transcription than words alone. Custom model license.
FireRedTeam/FireRedASR2-AED
Automatic Speech Recognition • Updated • 828 • 33Note Mandarin, Chinese dialects/accents, English and code-switching; speech and singing transcription. ASR component of FireRedASR2S. Apache-2.0.
distil-whisper/distil-large-v3
Automatic Speech Recognition • 0.8B • Updated • 561k • 383Note Distilled Whisper large-v3 for English transcription, including long-form audio. Useful speed-oriented Whisper baseline. MIT.
-
Edge0/Audio8-ASR-Infinite
Automatic Speech Recognition • 4B • Updated • 37k • 2.42k