⚡ Vaani-2 Flash ASR

High-Throughput, Low-Latency Speech Foundation Model for Production APIs & Real-Time Streaming

Developed by Sorika Labs

Series: Flash Base Model: Whisper Turbo English SOTA Throughput Weights: FP16 Standalone License: Apache 2.0

Hugging Face Model • Sorika Labs Organization


🌟 Overview & Key Breakthroughs

Vaani-2 Flash is Sorika Labs' flagship high-throughput speech recognition foundation model. Built on top of OpenAI's 809M-parameter whisper-large-v3-turbo, Vaani-2 Flash delivers an unmatched combination of extreme inference velocity (up to 30.3x Real-Time) and flawless English and Indian recognition accuracy.

In our comprehensive 10-benchmark evaluation, Vaani-2 Flash achieved the #1 lowest English Average WER (5.90%) across all evaluated models, beating both Alibaba's Qwen3-ASR (6.19%) and OpenAI Whisper baselines (8.05%).

Key Highlights:

  • 🏆 #1 Overall English Accuracy (5.90% Avg WER):
    • VoxPopuli: 2.22% WER (Rank 1)
    • SPGISpeech: 3.55% WER (Rank 1)
    • Earnings22: 8.62% WER (Rank 1)
    • LibriSpeech Clean: 4.84% WER
  • ⚡ Blazing Inference Velocity (21.5x Real-Time): Thanks to the optimized 4-layer decoder of Whisper-Turbo, Vaani-2 Flash processes 30 seconds of audio in under 1.0 second on GPU, making it ideal for live streaming and enterprise APIs.
  • 🇮🇳 Robust Indian Speech Recognition: Achieves 16.38% on FLEURS Hindi and 14.44% on Private Conversational Hindi, slashing Whisper-Small's 47.8% error rate by more than 60%.
  • 🛡️ Zero Catastrophic Forgetting: Preserves global multilingual capabilities across European and Asian languages without weight drift.

🧭 Sorika Labs Vaani Speech Series Portfolio

The Vaani family is engineered by Sorika Labs to provide production-ready speech intelligence across diverse computational constraints:

Series Model Name Backbone Params Primary Strength Recommended Use Case
Lite Series vaani1-lite-asr Whisper-Small 244M Ultra-lightweight, 99+ global languages Edge devices, mobile apps, low-latency CPU
Lite Series vaani1.1-lite-asr Qwen3-ASR 600M SOTA Hindi accuracy (11.78% WER), LLM-grade reasoning High-accuracy transcription, noisy audio, Devanagari purity
Flash Series vaani2-flash Whisper-Turbo 809M #1 English WER (5.90%), 21.5x Real-Time speed High-throughput production APIs, real-time live streaming
Flash Series vaani2.1-flash (Coming Soon) Whisper-Turbo 809M Curated Hindi + Hinglish fine-tune targeting <10% WER Ultra-fast bilingual enterprise applications
Pro Series Vaani-Pro (Research) Dense Foundation 1.5B+ Long-form context, speaker diarization, multi-dialect Complex enterprise contact centers & broadcasts

📊 Grand Benchmark Leaderboard (10 Standardized Benchmarks)

All models were evaluated under identical real-world conditions across 7 English benchmarks (meeting, financial, podcast, clean read, noisy speech, parliamentary) and 3 Indic/Hindi benchmarks:

Benchmark Dataset Domain / Category OpenAI Whisper-Small (244M) Vaani-1 Lite (244M) Vaani-2 Flash (809M) Alibaba Qwen3-Base (600M) Vaani-1.1 Lite (600M) Vaani-2 Speed
AMI-Cleaned English (Meeting Speech) 12.74% 12.99% 10.41% 9.05% 15.13% 29.4x RTF ⚡
Earnings22 English (Financial Calls) 8.62% 9.90% 8.62% 🥇 10.99% 10.50% 30.3x RTF ⚡
GigaSpeech English (Podcasts / Books) 8.90% 8.43% 7.28% 6.77% 7.58% 19.1x RTF ⚡
LibriSpeech Clean English (Clean Read) 5.63% 9.67% 4.84% 4.49% 4.75% 15.9x RTF ⚡
LibriSpeech Other English (Noisy / Accents) 11.30% 8.70% 4.35% 3.93% 3.60% 🥇 15.3x RTF ⚡
SPGISpeech English (Financial Disclosures) 3.89% 5.41% 3.55% 🥇 4.74% 5.58% 18.9x RTF ⚡
VoxPopuli English (Parliamentary / Live) 5.29% 9.20% 2.22% 🥇 3.38% 8.99% 28.1x RTF ⚡
Kathbath (Hindi) Hindi (Indic Crowdsourced) 50.00% 16.29% 22.47% 14.61% 12.36% 🥇 5.9x RTF
FLEURS (Hindi) Hindi (Speech-to-Text) 48.88% 25.56% 16.38% 14.39% 14.64% 8.5x RTF
Vaani Private (Hindi) Hindi (Conversational Real) 44.44% 16.11% 14.44% 7.78% 8.33% 6.8x RTF
English Average WER 7 Datasets 8.05% 9.18% 5.90% 🏆 6.19% 8.02% 21.5x Avg 🚀
Hindi Average WER 3 Datasets 47.77% 19.32% 17.76% 12.26% 11.78% 🏆 7.1x Avg
Overall Average WER All 10 Datasets 19.97% 12.23% 9.46% 8.01% 9.15% 17.2x Avg

Lower WER indicates superior recognition accuracy. Speed is measured in Real-Time Factor (RTFx) on NVIDIA GPU.


🚀 Quickstart & Inference

1. Using Hugging Face pipeline

import torch
from transformers import pipeline

device = "cuda:0" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if torch.cuda.is_available() else torch.float32

pipe = pipeline(
    "automatic-speech-recognition",
    model="sorika-labs/vaani2-flash",
    torch_dtype=dtype,
    device=device,
    chunk_length_s=30,
    trust_remote_code=True
)

# 1. Transcribe English Audio (Ultra-Accurate 5.90% Avg WER)
result_en = pipe("english_call.wav", generate_kwargs={"language": "en", "task": "transcribe"})
print("English:", result_en["text"])

# 2. Transcribe Hindi Audio
result_hi = pipe("hindi_audio.wav", generate_kwargs={"language": "hi", "task": "transcribe"})
print("Hindi:", result_hi["text"])

# 3. Automatic Language Detection & Transcription
result_auto = pipe("mixed_speech.mp3", generate_kwargs={"task": "transcribe"})
print("Transcription:", result_auto["text"])

2. Using AutoModelForSpeechSeq2Seq & AutoProcessor

import torch
import torchaudio
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

model_id = "sorika-labs/vaani2-flash"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if torch.cuda.is_available() else torch.float32

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    model_id,
    torch_dtype=dtype,
    low_cpu_mem_usage=True,
    trust_remote_code=True
).to(device)

# Load 16 kHz Mono Audio
audio, sr = torchaudio.load("sample.wav")
if sr != 16000:
    audio = torchaudio.transforms.Resample(sr, 16000)(audio)

inputs = processor(audio.squeeze().numpy(), sampling_rate=16000, return_tensors="pt")
input_features = inputs.input_features.to(device, dtype=dtype)

with torch.no_grad():
    predicted_ids = model.generate(input_features, language="hi", task="transcribe")

transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)
print("Transcription:", transcription[0])

🔮 Roadmap: Vaani-2.1 Flash

Sorika Labs is actively fine-tuning the next revision, vaani2.1-flash, which pairs the blazing 25x+ RTF Whisper Turbo architecture with a high-fidelity curated Hindi & Hinglish conversational corpus. Target goals:

  • Sub-10% Hindi Average WER
  • Native Hinglish code-switching normalization
  • Retaining the #1 English benchmark ranking

⚙️ Architecture & Model Specs

  • Base Model: openai/whisper-large-v3-turbo (809 Million Parameters)
  • Audio Encoder: 32-layer Transformer, 1280 hidden size, 128 mel bins
  • Decoder: 4-layer fast Transformer decoder (4x faster than standard large-v3)
  • Weights: 100% permanently fused model.safetensors (1.62 GB FP16)
  • License: Apache 2.0

🏢 Developed by Sorika Labs

Sorika Labs — Advancing open-source speech AI for Indian and global enterprises.

Downloads last month
58
Safetensors
Model size
0.8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sorika-labs/vaani2-flash

Finetuned
(644)
this model

Datasets used to train sorika-labs/vaani2-flash

Collection including sorika-labs/vaani2-flash

Evaluation results