Vedic chanting aligner (wav2vec2-CTC, Devanagari)

addy88/wav2vec2-sanskrit-stt fine-tuned on Vedavani (43.5 h of chanted Sanskrit) for forced alignment of a known Devanagari transcript onto a recitation.

It is used as the alignment and pronunciation backbone of a Vedic chanting assessment app: syllable boundaries come from torchaudio.forced_align over its emissions, and pronunciation is Goodness-of-Pronunciation off its posteriors. It is trained to place chanted Sanskrit accurately, not to be a general-purpose Sanskrit ASR system.

Two things you must know before using it

1. The CTC blank is <s> (id 0), not <pad>. This is inherited from the base checkpoint, whose config names <pad> while the token that actually behaves as the blank is <s> -- it argmaxes on ~89 % of frames while <pad> never fires. Pass blank=0 to torchaudio.forced_align, and strip <s> yourself when decoding: processor.batch_decode strips the nominal blank and will leave literal <s> in every hypothesis.

2. Speed perturbation is why the boundaries hold. Vedavani is single verses at one tempo. Without augmentation the model places the elongated final syllable of a chanted line correctly at 1.0x and truncates it past 1.06x -- which reads as a much worse rhythm measurement, not as a worse loss. Trained here with resampling over 0.85-1.25x.

Results

base this
held-out CER (Vedavani test) 0.537 0.084
CER on a held-out chanted teacher recording 0.191 0.064
syllables the recognizer mis-hears on a correct rendition 8.6 % 1.7 %

Training

12 epochs, lr 1e-4 linear with 500 warmup steps, effective batch 24, fp16, feature encoder frozen, ctc_loss_reduction="mean", clips capped at 16 s, speed perturbation 0.85-1.25x. The vocab is reused byte for byte from the base checkpoint -- 67 tokens, never rebuilt.

Usage

from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor

processor = Wav2Vec2Processor.from_pretrained("chiranjeevisagi/wav2vec2-vedic-aligner")
model = Wav2Vec2ForCTC.from_pretrained("chiranjeevisagi/wav2vec2-vedic-aligner")
# ... and remember blank = <s> = id 0

License

Apache-2.0, matching the base checkpoint and the corpus.

Downloads last month
1,200
Safetensors
Model size
94.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for chiranjeevisagi/wav2vec2-vedic-aligner

Finetuned
(1)
this model