You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

parakeet-hi-v2

Hindi (and Hinglish) speech recognition, fine-tuned from nvidia/parakeet-tdt-0.6b-v3 (FastConformer encoder + TDT decoder, 0.6B parameters) on ~17K hours of Hindi from kapturecx/bolAIndia, mostly call-centre speech.

Status: final checkpoint (step 75,000) of the first full training run: 22,200 h of audio seen, learning rate fully decayed. It replaces parakeet-hi-v1-pretrain (step 20,000). Internal use only.

v2 vs v1-pretrain

test v1-pretrain v2
Devanagari-only unseen, human references (6,300 clips): WER / CER 18.9 / 7.3 16.7 / 6.4

Files

file size use
parakeet-hi-v2.nemo 2.4 GB NeMo checkpoint: GPU inference, further fine-tuning
onnx/encoder.onnx + onnx/encoder.weights 2.4 GB sherpa-onnx, fp32 (best accuracy)
onnx/encoder.int8.onnx 0.84 GB sherpa-onnx, int8 (fastest on CPU)
onnx/decoder*.onnx, onnx/joiner*.onnx small TDT prediction network and joint, fp32 and int8
onnx/tokens.txt 2,048 SentencePiece tokens + blank
transcribe.py CPU transcription script (sherpa-onnx; no PyTorch or NeMo needed)

Quick start: CPU with sherpa-onnx (recommended for local use)

sherpa-onnx is a C++ runtime with Python bindings that runs NeMo TDT models natively. It needs no GPU, PyTorch or NeMo.

pip install sherpa-onnx soundfile numpy huggingface_hub
hf auth login                    # the repo is private
# int8 only (~0.9 GB download). Repeat --include for every pattern:
hf download amn-raw/parakeet-hi-v2 --local-dir parakeet-hi-v2 \
    --include "onnx/*.int8.onnx" --include "onnx/tokens.txt" --include "transcribe.py"
# or everything, including fp32 ONNX and the .nemo (~5.6 GB):
# hf download amn-raw/parakeet-hi-v2 --local-dir parakeet-hi-v2
cd parakeet-hi-v2
python transcribe.py call.wav                       # int8, 4 threads
python transcribe.py --precision fp32 --threads 8 call.wav   # needs the fp32 files too

transcribe.py takes audio of any length, sample rate and channel count:

  • Long audio is cut at pauses into pieces of at most 20 s (the clip length the model was trained on) and decoded in batches of 16.
  • Stereo files are transcribed per channel (agent / customer); --mix merges them.
  • Resampling to 16 kHz is automatic.

From Python:

import soundfile as sf
from transcribe import load_recognizer, transcribe

rec = load_recognizer("onnx", precision="int8", threads=4)
audio, sr = sf.read("call.wav", dtype="float32")     # mono, 16 kHz
print(transcribe(rec, [audio])[0])

sherpa-onnx also runs these same files from C++, Java/Android, iOS, C#, Go and others. Use model_type="nemo_transducer", feature_dim=128 and greedy_search decoding.

Which precision?

Measured on 200 held-out clips (Lahaja + call-centre, 18.5 min of audio), AMD EPYC Genoa CPU, 16 threads:

backend WER CER speed
sherpa-onnx fp32 19.8 9.7 62× real time
sherpa-onnx int8 20.2 9.7 54× (62–71× on an idle machine with v1)
NeMo / PyTorch fp32 (reference) 19.7 9.7 20×
  • fp32 ONNX gives the same accuracy as NeMo.
  • int8 costs about 0.4 WER points; it is the smaller download (0.9 GB vs 2.4 GB) and, on an idle CPU, the fastest.
  • The speeds in this table were measured while GPU jobs were using the same CPU, so they're lower than on an idle machine.
  • The int8 encoder quantises only MatMul/Gemm layers, per channel. Quantising every layer (sherpa-onnx's default script) gave WER 26.1 and was slower.
  • Peak RAM with --batch 16 is about 8 GB.

GPU: NeMo

pip install "nemo_toolkit[asr]>=2.4" "numba>=0.61,<0.63"   # numba 0.68 breaks the TDT loss kernel
import nemo.collections.asr as nemo_asr
from omegaconf import open_dict

model = nemo_asr.models.ASRModel.restore_from("parakeet-hi-v2.nemo").cuda().eval()
# NeMo's CUDA-graph greedy decoder hit an illegal memory access in our runs; plain greedy is safe.
cfg = model.cfg.decoding
with open_dict(cfg):
    cfg.greedy.use_cuda_graph_decoder = False
model.change_decoding_strategy(cfg)

print(model.transcribe(["clip.wav"])[0].text)      # clips up to ~20 s

For long recordings with NeMo, cut the audio first (for example with segment() from transcribe.py) and transcribe the pieces in a batch: the model has only seen clips of 20 s or less. On an H200 (bf16), batch decoding runs at about 600–1,000× real time.

Output conventions

  • Hindi in Devanagari; English words mostly in lowercase Latin script (payment, loan, hello). The training transcripts are inconsistent here, so the model sometimes writes English words in Devanagari (हेलो, लोन).
  • Numbers as digits (3 दिन, 2500).
  • Light punctuation (। , ?) appears but is not reliable. All scores below ignore punctuation.

Evaluation

All WERs below are on normalised text: NFC, punctuation removed, Latin lowercased, Devanagari digits mapped to 0–9. Devanagari vowel signs are kept. Scores are corpus-level (total errors / total reference words).

Devanagari-only, unseen (kapturecx/hindi-devanagari-asr-eval-v1)

The main Hindi number. Every clip's reference is pure Devanagari (no code-mixed English), from held-out splits, with possibly_seen_in_training clips removed, human-transcribed sources only. Up to 1,000 clips per source: 6,300 clips.

source clips WER CER
vaani-test 1,000 13.2 5.3
fleurs-test 401 13.2 4.8
svq 1,000 16.4 6.9
lahaja 1,000 16.8 5.8
navana-hindi 1,000 17.1 5.7
indicvoices-valid 1,000 18.5 7.3
spring-inx-eval 899 19.4 8.0
all 6,300 16.7 6.4

Reproduce with python -m scripts.eval_asr --model <ckpt> --datasets devanagari --max-per-source 1000 in the training repo.

Held-out test sets (500 clips each, including code-mixed clips; never trained on)

set domain WER CER
vaani-test read / descriptive speech, 165 districts 10.6 4.3
fleurs-test read sentences 13.4 5.0
navana-hindi mixed benchmark 15.0 4.6
lahaja read + extempore, 83 districts 16.0 5.4
svq short spoken questions on phones 16.8 7.4
indicvoices-valid read / extempore / conversational 19.2 7.9
spring-inx-eval phone conversations (54% code-mixed) 21.9 11.1
mean 16.1 6.5

Limitations

  • Translation on English speech: on English-only audio (for example recorded IVR prompts) the model sometimes outputs a Hindi translation instead of an English transcript.
  • Background speech and noise produce extra words. Trimming silence or applying VAD before transcription helps.
  • Script for English words is inconsistent (Latin vs Devanagari).
  • No timestamps in transcribe.py. Word timings are available through NeMo (timestamps=True).

Training

Base nvidia/parakeet-tdt-0.6b-v3 encoder, LSTM prediction network and joint projections
New SentencePiece BPE tokenizer, 2,048 tokens (Devanagari + Latin + digits); new token embedding and output layer
Data kapturecx/bolAIndia Hindi (one pass over the call-centre data, ~3 over human-labelled), yodas3 excluded, filtered to 17,176 h. 11.9K h call-centre, 2.0K h human-labelled (IndicVoices, Vaani, SPRING-INX, Kathbath, SYSPIN, Rasa…), 3.3K h aligned / ASR-labelled (Shrutilipi, Chaashini, WorldSpeech, NPTEL…). Sampled 50 / 25 / 25.
Filters 0.8–20 s, at most 25 characters/s, call-centre vendor confidence ≥ 0.80, test splits held out
Recipe TDT loss (durations 0–4, sigma 0.02, omega 0.1). AdamW (0.9, 0.98), weight decay 1e-3. LR 1e-3 for decoder/joint, 1e-4 for encoder. Encoder frozen for the first 2k steps, 2k warmup, cosine schedule. bf16, gradient clip 1.0, ~1,400 s of audio per batch. SpecAugment, plus telephone-band simulation on 30% of non-call-centre audio. 1× H200.

Base model: NVIDIA parakeet-tdt-0.6b-v3, CC-BY-4.0.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for amn-raw/parakeet-hi-v2

Quantized
(104)
this model

Evaluation results