Instructions to use amn-raw/parakeet-hi-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use amn-raw/parakeet-hi-v2 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("amn-raw/parakeet-hi-v2") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
parakeet-hi-v2
Hindi (and Hinglish) speech recognition, fine-tuned from nvidia/parakeet-tdt-0.6b-v3 (FastConformer encoder + TDT decoder, 0.6B parameters) on ~17K hours of Hindi from kapturecx/bolAIndia, mostly call-centre speech.
Status: final checkpoint (step 75,000) of the first full training run: 22,200 h of audio seen, learning rate fully decayed. It replaces parakeet-hi-v1-pretrain (step 20,000). Internal use only.
v2 vs v1-pretrain
| test | v1-pretrain | v2 |
|---|---|---|
| Devanagari-only unseen, human references (6,300 clips): WER / CER | 18.9 / 7.3 | 16.7 / 6.4 |
Files
| file | size | use |
|---|---|---|
parakeet-hi-v2.nemo |
2.4 GB | NeMo checkpoint: GPU inference, further fine-tuning |
onnx/encoder.onnx + onnx/encoder.weights |
2.4 GB | sherpa-onnx, fp32 (best accuracy) |
onnx/encoder.int8.onnx |
0.84 GB | sherpa-onnx, int8 (fastest on CPU) |
onnx/decoder*.onnx, onnx/joiner*.onnx |
small | TDT prediction network and joint, fp32 and int8 |
onnx/tokens.txt |
2,048 SentencePiece tokens + blank | |
transcribe.py |
CPU transcription script (sherpa-onnx; no PyTorch or NeMo needed) |
Quick start: CPU with sherpa-onnx (recommended for local use)
sherpa-onnx is a C++ runtime with Python bindings that runs NeMo TDT models natively. It needs no GPU, PyTorch or NeMo.
pip install sherpa-onnx soundfile numpy huggingface_hub
hf auth login # the repo is private
# int8 only (~0.9 GB download). Repeat --include for every pattern:
hf download amn-raw/parakeet-hi-v2 --local-dir parakeet-hi-v2 \
--include "onnx/*.int8.onnx" --include "onnx/tokens.txt" --include "transcribe.py"
# or everything, including fp32 ONNX and the .nemo (~5.6 GB):
# hf download amn-raw/parakeet-hi-v2 --local-dir parakeet-hi-v2
cd parakeet-hi-v2
python transcribe.py call.wav # int8, 4 threads
python transcribe.py --precision fp32 --threads 8 call.wav # needs the fp32 files too
transcribe.py takes audio of any length, sample rate and channel count:
- Long audio is cut at pauses into pieces of at most 20 s (the clip length the model was trained on) and decoded in batches of 16.
- Stereo files are transcribed per channel (agent / customer);
--mixmerges them. - Resampling to 16 kHz is automatic.
From Python:
import soundfile as sf
from transcribe import load_recognizer, transcribe
rec = load_recognizer("onnx", precision="int8", threads=4)
audio, sr = sf.read("call.wav", dtype="float32") # mono, 16 kHz
print(transcribe(rec, [audio])[0])
sherpa-onnx also runs these same files from C++, Java/Android, iOS, C#, Go and others. Use
model_type="nemo_transducer", feature_dim=128 and greedy_search decoding.
Which precision?
Measured on 200 held-out clips (Lahaja + call-centre, 18.5 min of audio), AMD EPYC Genoa CPU, 16 threads:
| backend | WER | CER | speed |
|---|---|---|---|
| sherpa-onnx fp32 | 19.8 | 9.7 | 62× real time |
| sherpa-onnx int8 | 20.2 | 9.7 | 54× (62–71× on an idle machine with v1) |
| NeMo / PyTorch fp32 (reference) | 19.7 | 9.7 | 20× |
- fp32 ONNX gives the same accuracy as NeMo.
- int8 costs about 0.4 WER points; it is the smaller download (0.9 GB vs 2.4 GB) and, on an idle CPU, the fastest.
- The speeds in this table were measured while GPU jobs were using the same CPU, so they're lower than on an idle machine.
- The int8 encoder quantises only MatMul/Gemm layers, per channel. Quantising every layer (sherpa-onnx's default script) gave WER 26.1 and was slower.
- Peak RAM with
--batch 16is about 8 GB.
GPU: NeMo
pip install "nemo_toolkit[asr]>=2.4" "numba>=0.61,<0.63" # numba 0.68 breaks the TDT loss kernel
import nemo.collections.asr as nemo_asr
from omegaconf import open_dict
model = nemo_asr.models.ASRModel.restore_from("parakeet-hi-v2.nemo").cuda().eval()
# NeMo's CUDA-graph greedy decoder hit an illegal memory access in our runs; plain greedy is safe.
cfg = model.cfg.decoding
with open_dict(cfg):
cfg.greedy.use_cuda_graph_decoder = False
model.change_decoding_strategy(cfg)
print(model.transcribe(["clip.wav"])[0].text) # clips up to ~20 s
For long recordings with NeMo, cut the audio first (for example with segment() from
transcribe.py) and transcribe the pieces in a batch: the model has only seen clips of 20 s or less.
On an H200 (bf16), batch decoding runs at about 600–1,000× real time.
Output conventions
- Hindi in Devanagari; English words mostly in lowercase Latin script (
payment,loan,hello). The training transcripts are inconsistent here, so the model sometimes writes English words in Devanagari (हेलो,लोन). - Numbers as digits (
3 दिन,2500). - Light punctuation (
। , ?) appears but is not reliable. All scores below ignore punctuation.
Evaluation
All WERs below are on normalised text: NFC, punctuation removed, Latin lowercased, Devanagari digits mapped to 0–9. Devanagari vowel signs are kept. Scores are corpus-level (total errors / total reference words).
Devanagari-only, unseen (kapturecx/hindi-devanagari-asr-eval-v1)
The main Hindi number. Every clip's reference is pure Devanagari (no code-mixed English), from
held-out splits, with possibly_seen_in_training clips removed, human-transcribed sources only. Up to 1,000
clips per source: 6,300 clips.
| source | clips | WER | CER |
|---|---|---|---|
| vaani-test | 1,000 | 13.2 | 5.3 |
| fleurs-test | 401 | 13.2 | 4.8 |
| svq | 1,000 | 16.4 | 6.9 |
| lahaja | 1,000 | 16.8 | 5.8 |
| navana-hindi | 1,000 | 17.1 | 5.7 |
| indicvoices-valid | 1,000 | 18.5 | 7.3 |
| spring-inx-eval | 899 | 19.4 | 8.0 |
| all | 6,300 | 16.7 | 6.4 |
Reproduce with
python -m scripts.eval_asr --model <ckpt> --datasets devanagari --max-per-source 1000
in the training repo.
Held-out test sets (500 clips each, including code-mixed clips; never trained on)
| set | domain | WER | CER |
|---|---|---|---|
| vaani-test | read / descriptive speech, 165 districts | 10.6 | 4.3 |
| fleurs-test | read sentences | 13.4 | 5.0 |
| navana-hindi | mixed benchmark | 15.0 | 4.6 |
| lahaja | read + extempore, 83 districts | 16.0 | 5.4 |
| svq | short spoken questions on phones | 16.8 | 7.4 |
| indicvoices-valid | read / extempore / conversational | 19.2 | 7.9 |
| spring-inx-eval | phone conversations (54% code-mixed) | 21.9 | 11.1 |
| mean | 16.1 | 6.5 |
Limitations
- Translation on English speech: on English-only audio (for example recorded IVR prompts) the model sometimes outputs a Hindi translation instead of an English transcript.
- Background speech and noise produce extra words. Trimming silence or applying VAD before transcription helps.
- Script for English words is inconsistent (Latin vs Devanagari).
- No timestamps in
transcribe.py. Word timings are available through NeMo (timestamps=True).
Training
| Base | nvidia/parakeet-tdt-0.6b-v3 encoder, LSTM prediction network and joint projections |
| New | SentencePiece BPE tokenizer, 2,048 tokens (Devanagari + Latin + digits); new token embedding and output layer |
| Data | kapturecx/bolAIndia Hindi (one pass over the call-centre data, ~3 over human-labelled), yodas3 excluded, filtered to 17,176 h. 11.9K h call-centre, 2.0K h human-labelled (IndicVoices, Vaani, SPRING-INX, Kathbath, SYSPIN, Rasa…), 3.3K h aligned / ASR-labelled (Shrutilipi, Chaashini, WorldSpeech, NPTEL…). Sampled 50 / 25 / 25. |
| Filters | 0.8–20 s, at most 25 characters/s, call-centre vendor confidence ≥ 0.80, test splits held out |
| Recipe | TDT loss (durations 0–4, sigma 0.02, omega 0.1). AdamW (0.9, 0.98), weight decay 1e-3. LR 1e-3 for decoder/joint, 1e-4 for encoder. Encoder frozen for the first 2k steps, 2k warmup, cosine schedule. bf16, gradient clip 1.0, ~1,400 s of audio per batch. SpecAugment, plus telephone-band simulation on 30% of non-call-centre audio. 1× H200. |
Base model: NVIDIA parakeet-tdt-0.6b-v3, CC-BY-4.0.
- Downloads last month
- -
Model tree for amn-raw/parakeet-hi-v2
Base model
nvidia/parakeet-tdt-0.6b-v3Evaluation results
- WER on Hindi Devanagari-only, unseen, human references (6,300 clips)test set self-reported16.740
- CER on Hindi Devanagari-only, unseen, human references (6,300 clips)test set self-reported6.380
- WER on Devanagari-only unseen: vaani-testtest set self-reported13.230
- CER on Devanagari-only unseen: vaani-testtest set self-reported5.310
- WER on Devanagari-only unseen: fleurs-testtest set self-reported13.180
- CER on Devanagari-only unseen: fleurs-testtest set self-reported4.830
- WER on Devanagari-only unseen: svqtest set self-reported16.350
- CER on Devanagari-only unseen: svqtest set self-reported6.870