You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

IndicConformer β€” Fine-tuned for Tactical Radio ASR (Hindi-English Code-Switched)

This is a LoRA adapter + custom CTC head for ai4bharat/IndicConformer (a NeMo hybrid RNNT+CTC checkpoint), fine-tuned to transcribe code-switched Hindi-English tactical radio communications over VHF/UHF push-to-talk channels, decoded to Latin/Roman-script text. It is part of an eight-architecture benchmark (this model, SeamlessM4T v2 Large, NVIDIA FastConformer, Whisper base/small/medium/large-v3, and wav2vec2-XLSR-53) evaluated on the same corpus β€” see the TacRFC dataset.

This repository contains only the fine-tuned delta on top of the base checkpoint: a LoRA adapter for the encoder's attention projections, a from-scratch CTC head (the original head cannot be reused β€” see below), and a replacement tokenizer. It is not a standalone model file and cannot be loaded with a plain from_pretrained() call β€” see How to Use.

Why this isn't a simple adapter load

  • CTC branch only. ai4bharat/IndicConformer is a hybrid RNNT+CTC model; only the CTC branch was fine-tuned here (a single linear head over encoder outputs β€” RNNT's prediction/joint network would add complexity with no benefit on an ~800-recording corpus).
  • New CTC head. The base checkpoint's CTC head emits Devanagari script sized for its original ~5,632-piece Indic-script BPE vocabulary. This corpus's transcripts are entirely Latin/Roman script (English + romanized Hindi), which that head cannot represent at all β€” so the CTC head was replaced and trained from scratch (ctc_head.pt), not merely adapted.
  • New tokenizer. A 128-token SentencePiece BPE tokenizer was trained from scratch on this corpus's own transcripts (tokenizer/) β€” small on purpose, given the small corpus. Its vocabulary is visibly phonetic-alphabet-heavy (▁alpha, ▁bravo, ▁over, ...), reflecting real radio-procedure speech.
  • Gated base checkpoint. ai4bharat/IndicConformer requires accepting its terms on the Hugging Face model page before it can be downloaded.
  • AI4Bharat's NeMo fork required. The base checkpoint needs AI4Bharat's fork of NeMo (git clone https://github.com/AI4Bharat/NeMo && git checkout nemo-v2), not upstream nemo_toolkit.

Model Details

  • Base model: ai4bharat/IndicConformer (multilingual hybrid RNNT+CTC NeMo checkpoint; EncDecHybridRNNTCTCBPEModel)
  • Fine-tuning method: LoRA, rank=16, alpha=32, dropout=0.05, applied to the Conformer encoder's attention projections (linear_q, linear_k, linear_v, linear_out); adapter left unmerged
  • CTC head: fully retrained linear head (ctc_head.pt) over the 128-token replacement vocabulary
  • Tokenizer: 128-token SentencePiece BPE, trained from this corpus's transcripts (tokenizer/tokenizer.model, vocab.txt)
  • Decoding: greedy CTC, no external language model

Intended Uses

  • Research and benchmarking of Indic-pretrained Conformer ASR robustness on noisy, code-switched, tactical radio speech.
  • A starting point for further fine-tuning on related radio-communications or code-switched Hindi-English speech tasks (would need re-fitting the CTC head/ tokenizer to any new target vocabulary).

Out of scope: general-purpose Indic-language transcription, and any input using a character/subword set outside this model's 128-token vocabulary (see vocabulary.json).

Training Data

Fine-tuned on the TacRFC dataset β€” "TacRFC: A Hindi-English Code-Mixed Tactical Radio Frequency Communication Corpus for Low-Resource Domain ASR" β€” a private, in-house corpus of 804 transcribed VHF/UHF tactical radio recordings, code-switched Hindi-English in Latin script, augmented with realistic radio background noise. See the dataset card for full details (including which annotation fields are actually populated), the official 80/10/10 stratified train/val/test split, and 5-fold cross-validation assignment.

Training Procedure

  • Framework: custom PyTorch training loop (train_indic_conformer.py), LoRA via peft, NeMo EncDecHybridRNNTCTCBPEModel
  • Batch size: 8 per device, gradient accumulation 1
  • Learning rate: 1e-4
  • Epochs: up to 200, with early stopping (patience=10 epochs on validation WER); this run's best checkpoint (by validation WER) was at epoch 200 β€” i.e. training was still improving when the run ended, so further training may improve this further
  • Data split: 80/10/10 train/val/test, stratified by VHF/UHF channel type, seed 42 (this run's own internal split; see the dataset card for the dataset's official released split, which may not be bit-identical to this training run's)
  • Max audio length: 30s

Evaluation Results

Computed on the held-out test split (127-138 samples depending on run), word/ character error rate using the same normalization (lowercase, punctuation removed, whitespace collapsed) used throughout this project. Error rates are fractions (0–1 scale, lower is better); values above 1 mean more errors than reference words or characters.

Metric Zero-shot baseline Fine-tuned (this adapter)
Word Error Rate (WER) 1.0089 0.8208
Character Error Rate (CER) 0.9391 0.7143

Fine-tuning improves substantially over the zero-shot baseline (which cannot emit Latin script at all without the replaced head/tokenizer), but this checkpoint is one of the weaker performers in the eight-model benchmark on this dataset β€” noticeably behind SeamlessM4T v2 Large (test WER 0.5699) and the retrained Whisper Large-v3 (test WER 0.5779). The retrained Whisper, wav2vec2 and FastConformer models are scored on a separate chunk-level test split, so that comparison is approximate. Given training had not plateaued by epoch 200 (see Training Procedure), this likely reflects an undertrained/smaller-capacity fit rather than a hard architectural ceiling.

Validation-split numbers at the best checkpoint (epoch 200): WER 0.8024, CER 0.7004.

How to Use

Requires AI4Bharat's NeMo fork (pip install upstream nemo_toolkit[asr] will not work β€” the base checkpoint needs AI4Bharat's nemo-v2 branch) and peft:

git clone https://github.com/AI4Bharat/NeMo.git
cd NeMo && git checkout nemo-v2 && bash reinstall.sh
from huggingface_hub import hf_hub_download, snapshot_download
import torch
import nemo.collections.asr as nemo_asr
from peft import PeftModel, LoraConfig, get_peft_model

# 1. Base checkpoint (gated β€” accept terms on the ai4bharat/IndicConformer page first)
base_nemo_path = hf_hub_download("ai4bharat/IndicConformer", "IndicConformer.nemo")
model = nemo_asr.models.ASRModel.restore_from(base_nemo_path, map_location="cpu")

# 2. Swap in this repo's 128-token tokenizer (replaces the base Indic-script BPE vocab)
adapter_dir = snapshot_download("happyman11/TacRFC-indicconformer")
model.change_vocabulary(new_tokenizer_dir=f"{adapter_dir}/tokenizer", new_tokenizer_type="bpe")

# 3. Apply the LoRA adapter to the encoder attention projections
lora_cfg = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
                      target_modules=["linear_q", "linear_k", "linear_v", "linear_out"], bias="none")
model = get_peft_model(model, lora_cfg)
model = PeftModel.from_pretrained(model, adapter_dir)

# 4. Load the retrained CTC head over the new 128-token vocabulary
ctc_head_state = torch.load(f"{adapter_dir}/ctc_head.pt", map_location="cpu")
model.base_model.model.ctc_decoder.load_state_dict(ctc_head_state)

model.eval()
transcripts = model.transcribe(audio=["your_radio_clip.wav"])
print(transcripts)

(This mirrors the loading logic in this project's own indic_conformer_model.py β€” consult it directly if the exact attribute path above has drifted from your installed NeMo/PEFT versions.)

Limitations and Ethical Considerations

  • Research artifact, not a production system, and one of the weaker models benchmarked on this corpus. See Evaluation Results β€” do not use this checkpoint where SeamlessM4T v2 Large or Whisper Large-v3 (both fine-tuned on the same data) are available instead.
  • Narrow domain and narrow vocabulary. The 128-token tokenizer was sized for an ~800-recording corpus; it will not represent out-of-domain vocabulary well.
  • Dual-use / sensitive domain. This model transcribes simulated military tactical radio communications. It is released for defense-communications ASR research and benchmarking purposes. Do not use it to process real operational, classified, or intercepted communications, and do not deploy it in any surveillance or targeting system without appropriate legal authorization and human oversight.
  • No PII/OPSEC review guarantee. While the underlying dataset is a research corpus rather than real intercepted traffic, no formal OPSEC/PII audit of model outputs has been performed.

Citation

If you use this model, please cite the corresponding paper.

@misc{tacrfc_indicconformer,
  title  = {TacRFC: A Hindi-English Code-Mixed Tactical Radio Frequency Communication Corpus for Low-Resource Domain ASR},
  author = {[Your name / lab], IIIT-Delhi},
  year   = {2026},
  note   = {Fine-tuned IndicConformer LoRA adapter + custom CTC head, TacRFC corpus}
}

(Replace the author field above with the correct attribution before publishing.)

Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for happyman11/TacRFC-indicconformer

Adapter
(1)
this model