Instructions to use chiranjeevisagi/wav2vec2-vedic-aligner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use chiranjeevisagi/wav2vec2-vedic-aligner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="chiranjeevisagi/wav2vec2-vedic-aligner")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForCTC processor = AutoProcessor.from_pretrained("chiranjeevisagi/wav2vec2-vedic-aligner") model = AutoModelForCTC.from_pretrained("chiranjeevisagi/wav2vec2-vedic-aligner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Vedic chanting aligner (wav2vec2-CTC, Devanagari)
addy88/wav2vec2-sanskrit-stt fine-tuned on Vedavani (43.5 h of chanted Sanskrit) for
forced alignment of a known Devanagari transcript onto a recitation.
It is used as the alignment and pronunciation backbone of a Vedic chanting
assessment app: syllable boundaries come from torchaudio.forced_align over
its emissions, and pronunciation is Goodness-of-Pronunciation off its
posteriors. It is trained to place chanted Sanskrit accurately, not to be a
general-purpose Sanskrit ASR system.
Two things you must know before using it
1. The CTC blank is <s> (id 0), not <pad>. This is inherited from the
base checkpoint, whose config names <pad> while the token that actually
behaves as the blank is <s> -- it argmaxes on ~89 % of frames while <pad>
never fires. Pass blank=0 to torchaudio.forced_align, and strip <s>
yourself when decoding: processor.batch_decode strips the nominal blank
and will leave literal <s> in every hypothesis.
2. Speed perturbation is why the boundaries hold. Vedavani is single verses at one tempo. Without augmentation the model places the elongated final syllable of a chanted line correctly at 1.0x and truncates it past 1.06x -- which reads as a much worse rhythm measurement, not as a worse loss. Trained here with resampling over 0.85-1.25x.
Results
| base | this | |
|---|---|---|
| held-out CER (Vedavani test) | 0.537 | 0.084 |
| CER on a held-out chanted teacher recording | 0.191 | 0.064 |
| syllables the recognizer mis-hears on a correct rendition | 8.6 % | 1.7 % |
Training
12 epochs, lr 1e-4 linear with 500 warmup steps, effective batch 24,
fp16, feature encoder frozen, ctc_loss_reduction="mean", clips capped at
16 s, speed perturbation 0.85-1.25x. The vocab is reused byte for byte from
the base checkpoint -- 67 tokens, never rebuilt.
Usage
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
processor = Wav2Vec2Processor.from_pretrained("chiranjeevisagi/wav2vec2-vedic-aligner")
model = Wav2Vec2ForCTC.from_pretrained("chiranjeevisagi/wav2vec2-vedic-aligner")
# ... and remember blank = <s> = id 0
License
Apache-2.0, matching the base checkpoint and the corpus.
- Downloads last month
- 1,200
Model tree for chiranjeevisagi/wav2vec2-vedic-aligner
Base model
addy88/wav2vec2-sanskrit-stt