Instructions to use happyman11/TacRFC-fastconformer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use happyman11/TacRFC-fastconformer with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("happyman11/TacRFC-fastconformer") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
FastConformer Hybrid (RNNT+CTC) β Fine-tuned for Tactical Radio ASR (Hindi-English Code-Switched)
This is an NVIDIA NeMo FastConformer Hybrid Transducer-CTC model fine-tuned to transcribe code-switched Hindi-English tactical radio communications over VHF/UHF push-to-talk channels, decoded to Latin/Roman-script text. It is one of eight architectures (FastConformer, SeamlessM4T v2 Large, IndicConformer, Whisper base/small/medium/large-v3, and wav2vec2-XLSR-53) benchmarked on the TacRFC dataset.
Updated September 2026. This repository now holds a retrained checkpoint. The model was fine-tuned again on chunked audio (each recording split into shorter segments paired with their transcript segments), with early stopping and checkpoint selection on validation loss, and is now scored on a held-out test split with a zero-shot baseline for comparison. The numbers in earlier versions of this card came from a validation split and are not comparable. The previous weights are still available in this repo's commit history.
Base checkpoint note: based on the exported config's architecture fingerprint (17 Conformer encoder layers, d_model=512, hybrid RNNT+CTC decoding, 1024-token BPE tokenizer trained on NeMo's
ASR_SET 3.0English corpus), this checkpoint is consistent with NVIDIA'sstt_en_fastconformer_hybrid_large_pc. The exact base checkpoint name is not recorded in the exported archive.
Model Details
- Architecture:
EncDecHybridRNNTCTCBPEModel(NeMo) β FastConformer encoder (17 layers, d_model=512) with a joint RNN-Transducer decoder and an auxiliary CTC head (CTC loss weight 0.3 during training) - Parameters: ~114M
- Tokenizer: 1024-token SentencePiece BPE inherited from the base checkpoint's
NeMo_ASR_SET/English/asr_set_3.0tokenizer (not rebuilt for this dataset) - Sample rate: 16 kHz mono
- Default decoding: RNNT branch,
greedy_batch(the CTC branch is available viachange_decoding_strategy, see How to Use) - Format: single
.nemoarchive (config + weights + tokenizer bundled together), exported from the best validation checkpoint
Intended Uses
- Research and benchmarking of Conformer/streaming-capable ASR robustness on noisy, code-switched, tactical radio speech.
- A starting checkpoint for further fine-tuning on related radio-communications or code-switched Hindi-English speech tasks.
Out of scope: general-purpose English transcription, and real-time/streaming deployment without re-validating latency and decoding configuration.
Training Data
This model was fine-tuned on TacRFC
β "TacRFC: A Hindi-English Code-Mixed Tactical Radio Frequency Communication Corpus
for Low-Resource Domain ASR" (internally referred to during data collection as
merged-data-augmented-radio-bg). Apart from the base checkpoint's original
pretraining, it was not trained on any public speech dataset.
- Domain: simulated/exercise military tactical voice traffic over VHF and UHF push-to-talk radio nets (e.g. "Delta one, report to DS that the net is through, over").
- Language: code-switched HindiβEnglish, transcribed entirely in Latin/Roman script (romanized Hindi + English). The training labels contain no Devanagari.
- Source corpus: 804 transcribed recordings matched to audio (out of 805 ground-truth rows), augmented with realistic radio background noise/static to reflect real net conditions.
- Chunking (this run): each recording was segmented into shorter chunks, each paired
with the matching part of its transcript (e.g.
record660_chunk_3). The chunks were split 80/10/10 with seed 42: about 1,100 training chunks, about 138 validation chunks, and a 138-chunk test set drawn from 117 different recordings. - Split caveat: this run made its own chunk-level split. It did not use the recording-level 80/10/10 split and 5-fold assignment published with the dataset. Other chunks of a test recording can therefore appear in training, so the test scores may be optimistic for completely unseen recordings or speakers.
- Annotation: every recording is labeled with a channel type (
VHF/UHF; fraction VHF β 0.92, UHF β 0.08); a speaker count is available for a fraction β 0.89 of recordings and an English translation for β 0.92. The source spreadsheet also defines a 30-category tactical operation taxonomy (Ambush, Contact with enemy, Casevac/medevac, etc.), but those columns are not populated in the released data. - Availability: released as a private Hugging Face dataset alongside this model. It was collected/simulated for tactical-communications ASR research and is not derived from real intercepted or classified traffic.
Training Procedure
- Framework: NVIDIA NeMo on PyTorch Lightning 2.4
- Optimizer / schedule: AdamW, peak learning rate 5e-6, weight decay 1e-3, cosine annealing after 50 warmup steps (minimum learning rate 1e-7)
- Batch size: 2
- Loss: joint RNNT + CTC (
ctc_loss_weight=0.3,ctc_reduction="mean_batch") - Checkpoint selection: early stopping (patience 10) and
ModelCheckpointboth monitor validation WER of the CTC branch (val_wer_ctc). The best checkpoint was at epoch 88, and the published.nemoarchive contains exactly those weights (verified tensor-by-tensor againstbest-epoch=087.ckpt).
Evaluation Results
Held-out test split
137 scored test chunks (138 in the split; one chunk has an empty reference and is skipped). WER uses this project's standard normalization (lowercase, punctuation removed, whitespace collapsed); CER is computed on the raw text. All error rates on this card are fractions between 0 and 1 (lower is better); values above 1 mean more errors than reference words or characters.
| Metric | Zero-shot baseline | Fine-tuned (this model) |
|---|---|---|
| Word Error Rate (WER) | 0.9643 | 0.6174 |
| Character Error Rate (CER) | 0.9302 | 0.5636 |
The zero-shot baseline is the pretrained FastConformer checkpoint before fine-tuning, decoded with the same settings on the same chunks.
Fine-tuning improves test WER from 0.9643 to 0.6174 and CER from 0.9302 to 0.5636. The zero-shot baseline returned an empty transcript for 77 of the 138 test chunks; after fine-tuning, 8 are still empty. With only ~114M parameters, this is the smallest model in the table, yet it ranks second on WER, ahead of Whisper Medium and Whisper Small. Its CER, however, is higher than that of any Whisper size.
Comparison with the other retrained models (same 138-chunk test set)
| Model | WER zero-shot β fine-tuned | CER zero-shot β fine-tuned |
|---|---|---|
| Whisper Large-v3 | 0.7553 β 0.5779 | 0.6630 β 0.4847 |
| FastConformer Hybrid (RNNT+CTC) (this model) | 0.9643 β 0.6174 | 0.9302 β 0.5636 |
| Whisper Medium | 0.8417 β 0.6263 | 0.7543 β 0.4945 |
| Whisper Small | 0.8761 β 0.6532 | 0.7522 β 0.5230 |
| wav2vec2-XLSR-53 | 0.9717 β 0.8799 | 0.7068 β 0.5769 |
Whisper Base, SeamlessM4T v2 Large and IndicConformer were not retrained in this round. Their published numbers come from different evaluation splits, so they are not included in this table.
Validation split (best checkpoint)
| Metric | Value |
|---|---|
Word Error Rate, CTC branch (val_wer_ctc) |
0.6487 |
| Selected checkpoint | epoch 88 (Lightning epoch index 87) |
These values were logged during training. They use the training loop's own decoding and scoring, not the normalized test-time evaluation above, so compare models on the test numbers.
How to Use
from huggingface_hub import hf_hub_download
import nemo.collections.asr as nemo_asr
model_path = hf_hub_download(repo_id="happyman11/TacRFC-fastconformer", filename="model.nemo")
model = nemo_asr.models.EncDecHybridRNNTCTCBPEModel.restore_from(model_path)
model.eval()
# Optional: decode with the CTC branch (the one monitored for checkpoint selection)
# instead of the default RNNT branch.
# model.change_decoding_strategy(decoder_type="ctc")
transcripts = model.transcribe(["your_radio_clip.wav"])
print(transcripts)
(Requires nemo_toolkit[asr].)
Limitations and Ethical Considerations
- Research artifact, not a production system. Fine-tuned on a very small corpus (804 recordings, about 1,100 training chunks) to benchmark how different architectures adapt to noisy, code-switched tactical radio speech. Absolute error rates remain high; treat outputs as a strong hint at content, not verbatim transcripts.
- Narrow domain. Performance is expected to degrade sharply outside VHF/UHF radio traffic, tactical/military phraseology, and Hindi-English code-switching in Latin script. It is not tuned for conversational, broadcast, or clean studio speech.
- Dual-use / sensitive domain. This model transcribes simulated military tactical radio communications. It is released for defense-communications ASR research and benchmarking purposes. Do not use it to process real operational, classified, or intercepted communications, and do not deploy it in any surveillance or targeting system without appropriate legal authorization and human oversight.
- No PII/OPSEC review guarantee. While the underlying dataset is a research corpus rather than real intercepted traffic, no formal OPSEC/PII audit of model outputs has been performed. Review outputs before publishing them verbatim in downstream applications.
- Small, chunk-level test set. 137 scored chunks is a small sample, and the chunk-level split lets other parts of a test recording appear in training (see Training Data). Treat the error rates as noisy estimates.
Citation
If you use this model, please cite the corresponding paper.
@misc{tacrfc_fastconformer,
title = {TacRFC: A Hindi-English Code-Mixed Tactical Radio Frequency Communication Corpus for Low-Resource Domain ASR},
author = {[Your name / lab], IIIT-Delhi},
year = {2026},
note = {Fine-tuned FastConformer Hybrid (RNNT+CTC) checkpoint, TacRFC corpus}
}
(Replace the author field above with the correct attribution before publishing.)
- Downloads last month
- -
Evaluation results
- Test WER on TacRFC (chunked test split)test set self-reported0.617
- Test CER on TacRFC (chunked test split)test set self-reported0.564