⚡ Vaani-2 Flash ASR
High-Throughput, Low-Latency Speech Foundation Model for Production APIs & Real-Time Streaming
Developed by Sorika Labs
🌟 Overview & Key Breakthroughs
Vaani-2 Flash is Sorika Labs' flagship high-throughput speech recognition foundation model. Built on top of OpenAI's 809M-parameter whisper-large-v3-turbo, Vaani-2 Flash delivers an unmatched combination of extreme inference velocity (up to 30.3x Real-Time) and flawless English and Indian recognition accuracy.
In our comprehensive 10-benchmark evaluation, Vaani-2 Flash achieved the #1 lowest English Average WER (5.90%) across all evaluated models, beating both Alibaba's Qwen3-ASR (6.19%) and OpenAI Whisper baselines (8.05%).
Key Highlights:
- 🏆 #1 Overall English Accuracy (5.90% Avg WER):
- VoxPopuli: 2.22% WER (Rank 1)
- SPGISpeech: 3.55% WER (Rank 1)
- Earnings22: 8.62% WER (Rank 1)
- LibriSpeech Clean: 4.84% WER
- ⚡ Blazing Inference Velocity (21.5x Real-Time): Thanks to the optimized 4-layer decoder of Whisper-Turbo, Vaani-2 Flash processes 30 seconds of audio in under 1.0 second on GPU, making it ideal for live streaming and enterprise APIs.
- 🇮🇳 Robust Indian Speech Recognition: Achieves 16.38% on FLEURS Hindi and 14.44% on Private Conversational Hindi, slashing Whisper-Small's 47.8% error rate by more than 60%.
- 🛡️ Zero Catastrophic Forgetting: Preserves global multilingual capabilities across European and Asian languages without weight drift.
🧭 Sorika Labs Vaani Speech Series Portfolio
The Vaani family is engineered by Sorika Labs to provide production-ready speech intelligence across diverse computational constraints:
| Series | Model Name | Backbone | Params | Primary Strength | Recommended Use Case |
|---|---|---|---|---|---|
| Lite Series | vaani1-lite-asr |
Whisper-Small | 244M | Ultra-lightweight, 99+ global languages | Edge devices, mobile apps, low-latency CPU |
| Lite Series | vaani1.1-lite-asr |
Qwen3-ASR | 600M | SOTA Hindi accuracy (11.78% WER), LLM-grade reasoning | High-accuracy transcription, noisy audio, Devanagari purity |
| Flash Series | vaani2-flash |
Whisper-Turbo | 809M | #1 English WER (5.90%), 21.5x Real-Time speed | High-throughput production APIs, real-time live streaming |
| Flash Series | vaani2.1-flash (Coming Soon) |
Whisper-Turbo | 809M | Curated Hindi + Hinglish fine-tune targeting <10% WER | Ultra-fast bilingual enterprise applications |
| Pro Series | Vaani-Pro (Research) | Dense Foundation | 1.5B+ | Long-form context, speaker diarization, multi-dialect | Complex enterprise contact centers & broadcasts |
📊 Grand Benchmark Leaderboard (10 Standardized Benchmarks)
All models were evaluated under identical real-world conditions across 7 English benchmarks (meeting, financial, podcast, clean read, noisy speech, parliamentary) and 3 Indic/Hindi benchmarks:
| Benchmark Dataset | Domain / Category | OpenAI Whisper-Small (244M) | Vaani-1 Lite (244M) | Vaani-2 Flash (809M) | Alibaba Qwen3-Base (600M) | Vaani-1.1 Lite (600M) | Vaani-2 Speed |
|---|---|---|---|---|---|---|---|
| AMI-Cleaned | English (Meeting Speech) | 12.74% | 12.99% | 10.41% | 9.05% | 15.13% | 29.4x RTF ⚡ |
| Earnings22 | English (Financial Calls) | 8.62% | 9.90% | 8.62% 🥇 | 10.99% | 10.50% | 30.3x RTF ⚡ |
| GigaSpeech | English (Podcasts / Books) | 8.90% | 8.43% | 7.28% | 6.77% | 7.58% | 19.1x RTF ⚡ |
| LibriSpeech Clean | English (Clean Read) | 5.63% | 9.67% | 4.84% | 4.49% | 4.75% | 15.9x RTF ⚡ |
| LibriSpeech Other | English (Noisy / Accents) | 11.30% | 8.70% | 4.35% | 3.93% | 3.60% 🥇 | 15.3x RTF ⚡ |
| SPGISpeech | English (Financial Disclosures) | 3.89% | 5.41% | 3.55% 🥇 | 4.74% | 5.58% | 18.9x RTF ⚡ |
| VoxPopuli | English (Parliamentary / Live) | 5.29% | 9.20% | 2.22% 🥇 | 3.38% | 8.99% | 28.1x RTF ⚡ |
| Kathbath (Hindi) | Hindi (Indic Crowdsourced) | 50.00% | 16.29% | 22.47% | 14.61% | 12.36% 🥇 | 5.9x RTF |
| FLEURS (Hindi) | Hindi (Speech-to-Text) | 48.88% | 25.56% | 16.38% | 14.39% | 14.64% | 8.5x RTF |
| Vaani Private (Hindi) | Hindi (Conversational Real) | 44.44% | 16.11% | 14.44% | 7.78% | 8.33% | 6.8x RTF |
| English Average WER | 7 Datasets | 8.05% | 9.18% | 5.90% 🏆 | 6.19% | 8.02% | 21.5x Avg 🚀 |
| Hindi Average WER | 3 Datasets | 47.77% | 19.32% | 17.76% | 12.26% | 11.78% 🏆 | 7.1x Avg |
| Overall Average WER | All 10 Datasets | 19.97% | 12.23% | 9.46% | 8.01% | 9.15% | 17.2x Avg |
Lower WER indicates superior recognition accuracy. Speed is measured in Real-Time Factor (RTFx) on NVIDIA GPU.
🚀 Quickstart & Inference
1. Using Hugging Face pipeline
import torch
from transformers import pipeline
device = "cuda:0" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if torch.cuda.is_available() else torch.float32
pipe = pipeline(
"automatic-speech-recognition",
model="sorika-labs/vaani2-flash",
torch_dtype=dtype,
device=device,
chunk_length_s=30,
trust_remote_code=True
)
# 1. Transcribe English Audio (Ultra-Accurate 5.90% Avg WER)
result_en = pipe("english_call.wav", generate_kwargs={"language": "en", "task": "transcribe"})
print("English:", result_en["text"])
# 2. Transcribe Hindi Audio
result_hi = pipe("hindi_audio.wav", generate_kwargs={"language": "hi", "task": "transcribe"})
print("Hindi:", result_hi["text"])
# 3. Automatic Language Detection & Transcription
result_auto = pipe("mixed_speech.mp3", generate_kwargs={"task": "transcribe"})
print("Transcription:", result_auto["text"])
2. Using AutoModelForSpeechSeq2Seq & AutoProcessor
import torch
import torchaudio
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
model_id = "sorika-labs/vaani2-flash"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if torch.cuda.is_available() else torch.float32
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForSpeechSeq2Seq.from_pretrained(
model_id,
torch_dtype=dtype,
low_cpu_mem_usage=True,
trust_remote_code=True
).to(device)
# Load 16 kHz Mono Audio
audio, sr = torchaudio.load("sample.wav")
if sr != 16000:
audio = torchaudio.transforms.Resample(sr, 16000)(audio)
inputs = processor(audio.squeeze().numpy(), sampling_rate=16000, return_tensors="pt")
input_features = inputs.input_features.to(device, dtype=dtype)
with torch.no_grad():
predicted_ids = model.generate(input_features, language="hi", task="transcribe")
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)
print("Transcription:", transcription[0])
🔮 Roadmap: Vaani-2.1 Flash
Sorika Labs is actively fine-tuning the next revision, vaani2.1-flash, which pairs the blazing 25x+ RTF Whisper Turbo architecture with a high-fidelity curated Hindi & Hinglish conversational corpus. Target goals:
- Sub-10% Hindi Average WER
- Native Hinglish code-switching normalization
- Retaining the #1 English benchmark ranking
⚙️ Architecture & Model Specs
- Base Model:
openai/whisper-large-v3-turbo(809 Million Parameters) - Audio Encoder: 32-layer Transformer, 1280 hidden size, 128 mel bins
- Decoder: 4-layer fast Transformer decoder (4x faster than standard large-v3)
- Weights: 100% permanently fused
model.safetensors(1.62 GB FP16) - License: Apache 2.0
🏢 Developed by Sorika Labs
Sorika Labs — Advancing open-source speech AI for Indian and global enterprises.
- Downloads last month
- 58
Model tree for sorika-labs/vaani2-flash
Base model
openai/whisper-large-v3Datasets used to train sorika-labs/vaani2-flash
openslr/librispeech_asr
Collection including sorika-labs/vaani2-flash
Evaluation results
- WER on LibriSpeech Cleanself-reported4.840
- WER on SPGISpeechself-reported3.550
- WER on VoxPopuliself-reported2.220
- WER on FLEURS Hindiself-reported16.380
- WER on Kathbath Hindiself-reported22.470