w2v-BERT Ethiopian ASR (recipe v5, ep14 best)

Fine-tune of facebook/w2v-bert-2.0 (~606M params) on EthioSpeech for multilingual ASR across six Ethiopian languages: amharic, tigrinya, oromo, sidama, somali, afar. Fully fine-tuned (not LoRA) with a fresh char-level CTC head over a 319-symbol vocabulary covering the Latin alphabet, Ge'ez syllables, and six [LID:*] language-conditioning tokens.

Benchmarks

All decoding is greedy CTC argmax (no LM, no waterfall post-processing), text normalised with normalizer v1.0.0, digits kept unfused. WER/CER are percentages; lower is better. Ξ” columns are this model minus whisper-fullft-v3, so negative numbers are improvements.

EthioSpeech legacy test split (main comparison; held out from training)

WER (%) on the EthioSpeech legacy test split. "β€”" = language unsupported by the model. Baseline rows are from the project's paper Table I; the Whisper FT row is ethiospeech01/whisper-fullft-v3-ethiopian.

Model am ti so sid aa om Avg
Whisper ZS 122.1 124.6 92.3 β€” β€” 105.6 β€”
MMS 54.3 67.9 45.7 39.0 β€” 54.1 52.2
W2V2-BERT FT (this model) 25.2 21.7 24.1 17.8 43.8 32.5 27.5
Whisper FT 34.5 30.1 27.9 18.4 55.9 26.3 32.2
Whisper LoRA 44.3 41.9 35.6 21.6 54.3 31.5 38.2

Direct comparison against Whisper FT on this split: better by 9.33pp (amharic), 8.44pp (tigrinya), 3.75pp (somali), 0.64pp (sidama), 12.12pp (afar); worse on oromo (+6.18pp, 32.5 vs 26.3, the only baseline win). Macro: 27.50 vs 32.2 β†’ βˆ’4.72pp. Same-split CER for this model (not measurable for the baseline rows): amharic 6.89, tigrinya 6.99, somali 7.74, sidama 3.86, afar 16.78, oromo 7.33, macro 8.27.

waxal/test (zero-shot, same split as ethiospeech01/whisper-fullft-v3-ethiopian)

Out-of-corpus transfer check on the Waxal data never seen in training.

Language n This model WER whisper-fullft-v3 WER Ξ” WER (pp) This model CER whisper-fullft-v3 CER Ξ” CER (pp)
amharic 3420 35.69 36.98 -1.29 13.24 14.33 -1.09
tigrinya 5034 49.23 52.83 -3.60 20.32 25.67 -5.35
oromo 3782 36.50 38.67 -2.17 8.34 15.87 -7.53
sidama 3561 37.28 37.12 +0.16 8.52 8.58 -0.06
macro β€” 39.67 41.40 -1.73 12.60 16.11 -3.51

Training-time dev

Best checkpoint recorded a dev WER of 30.72 at step 48000 / epoch 12 (selected by "is_best": true).

Quick start: inference

import soundfile as sf
import torch
from transformers import SeamlessM4TFeatureExtractor, Wav2Vec2BertForCTC, Wav2Vec2CTCTokenizer

MODEL_ID = "anshulsc/w2v-bert-ethio-v5"
LID_TOKEN_IDS = [311, 312, 313, 314, 315, 316]

speech, sr = sf.read("audio.wav", dtype="float32")
if speech.ndim > 1:
    speech = speech.mean(axis=1)

fe = SeamlessM4TFeatureExtractor.from_pretrained("facebook/w2v-bert-2.0")  # 16 kHz mel features
if sr != fe.sampling_rate:
    import librosa
    speech = librosa.resample(speech, orig_sr=sr, target_sr=fe.sampling_rate)

inputs = fe(speech, sampling_rate=fe.sampling_rate, return_tensors="pt", return_attention_mask=True)
model = Wav2Vec2BertForCTC.from_pretrained(MODEL_ID, torch_dtype=torch.float16).eval()
tokenizer = Wav2Vec2CTCTokenizer.from_pretrained(MODEL_ID)

with torch.no_grad():
    logits = model(**inputs).logits
pred_ids = logits.argmax(dim=-1)

# The head loves to echo its LID prefix at decode time; wipe those positions to blank first.
lid = torch.tensor(LID_TOKEN_IDS, device=pred_ids.device)
pred_ids = pred_ids.masked_fill(torch.isin(pred_ids, lid), tokenizer.pad_token_id)

print(tokenizer.decode(pred_ids[0]).strip())

inference_example.py in this repo is the same snippet as a runnable script (python inference_example.py path/to/audio.wav).

Token ID reference

Vocabulary size is 319. Decoding is plain CTC, so no BOS/EOS is ever emitted β€” the head produces only chars, the word delimiter, and occasionally its conditioning token.

Token ID Role
[LID:amharic] 311 training-time language conditioning prefix
[LID:tigrinya] 312 training-time language conditioning prefix
[LID:oromo] 313 training-time language conditioning prefix
[LID:somali] 314 training-time language conditioning prefix
[LID:sidama] 315 training-time language conditioning prefix
[LID:afar] 316 training-time language conditioning prefix
[PAD] 0 CTC blank / padding
[UNK] 1 unknown character
` ` 29
<s> 317 BOS (unused at CTC inference)
</s> 318 EOS (unused at CTC inference)

Everything from id 2 through 310 is a literal character token (Latin letters, ' and | plus Ge'ez syllabics). The | word delimiter decodes to a space.

Training

  • Base: facebook/w2v-bert-2.0 with the feature extractor kept frozen at 16 kHz / 80 mel bins.
  • Head: fresh char-level Wav2Vec2BertForCTC.lm_head over the custom 319-token vocab.
  • Conditioning: each label sequence is prefixed with [LID:<lang>]|, which is why the decoder can emit LID tokens and why they must be masked to blank before CTC collapsing.
  • Recipe: full fine-tune (no LoRA), regularisation schedule from the v5 recipe.

Caveats

  • Waxal numbers are zero-shot; EthioSpeech held-out numbers are in-domain. They are not directly comparable.
  • Amharic and Tigrinya were normalised into a mixed-script convention (Ge'ez + latin | punctuation); transcribed hypotheses follow the same convention.
Downloads last month
10
Safetensors
Model size
0.6B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ethiospeech01/w2v2-bert-ethio-v5

Finetuned
(553)
this model