Instructions to use ethiospeech01/w2v2-bert-ethio-v5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ethiospeech01/w2v2-bert-ethio-v5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="ethiospeech01/w2v2-bert-ethio-v5")# Load model directly from transformers import AutoProcessor, AutoModelForCTC processor = AutoProcessor.from_pretrained("ethiospeech01/w2v2-bert-ethio-v5") model = AutoModelForCTC.from_pretrained("ethiospeech01/w2v2-bert-ethio-v5", device_map="auto") - Notebooks
- Google Colab
- Kaggle
w2v-BERT Ethiopian ASR (recipe v5, ep14 best)
Fine-tune of facebook/w2v-bert-2.0 (~606M params) on
EthioSpeech for multilingual ASR across six Ethiopian languages: amharic, tigrinya, oromo, sidama, somali, afar.
Fully fine-tuned (not LoRA) with a fresh char-level CTC head over a 319-symbol vocabulary
covering the Latin alphabet, Ge'ez syllables, and six [LID:*] language-conditioning tokens.
Benchmarks
All decoding is greedy CTC argmax (no LM, no waterfall post-processing), text normalised with normalizer v1.0.0, digits kept unfused. WER/CER are percentages; lower is better. Ξ columns are this model minus whisper-fullft-v3, so negative numbers are improvements.
EthioSpeech legacy test split (main comparison; held out from training)
WER (%) on the EthioSpeech legacy test split. "β" = language unsupported by the model.
Baseline rows are from the project's paper Table I; the Whisper FT row is
ethiospeech01/whisper-fullft-v3-ethiopian.
| Model | am | ti | so | sid | aa | om | Avg |
|---|---|---|---|---|---|---|---|
| Whisper ZS | 122.1 | 124.6 | 92.3 | β | β | 105.6 | β |
| MMS | 54.3 | 67.9 | 45.7 | 39.0 | β | 54.1 | 52.2 |
| W2V2-BERT FT (this model) | 25.2 | 21.7 | 24.1 | 17.8 | 43.8 | 32.5 | 27.5 |
| Whisper FT | 34.5 | 30.1 | 27.9 | 18.4 | 55.9 | 26.3 | 32.2 |
| Whisper LoRA | 44.3 | 41.9 | 35.6 | 21.6 | 54.3 | 31.5 | 38.2 |
Direct comparison against Whisper FT on this split: better by 9.33pp (amharic), 8.44pp (tigrinya), 3.75pp (somali), 0.64pp (sidama), 12.12pp (afar); worse on oromo (+6.18pp, 32.5 vs 26.3, the only baseline win). Macro: 27.50 vs 32.2 β β4.72pp. Same-split CER for this model (not measurable for the baseline rows): amharic 6.89, tigrinya 6.99, somali 7.74, sidama 3.86, afar 16.78, oromo 7.33, macro 8.27.
waxal/test (zero-shot, same split as ethiospeech01/whisper-fullft-v3-ethiopian)
Out-of-corpus transfer check on the Waxal data never seen in training.
| Language | n | This model WER | whisper-fullft-v3 WER | Ξ WER (pp) | This model CER | whisper-fullft-v3 CER | Ξ CER (pp) |
|---|---|---|---|---|---|---|---|
| amharic | 3420 | 35.69 | 36.98 | -1.29 | 13.24 | 14.33 | -1.09 |
| tigrinya | 5034 | 49.23 | 52.83 | -3.60 | 20.32 | 25.67 | -5.35 |
| oromo | 3782 | 36.50 | 38.67 | -2.17 | 8.34 | 15.87 | -7.53 |
| sidama | 3561 | 37.28 | 37.12 | +0.16 | 8.52 | 8.58 | -0.06 |
| macro | β | 39.67 | 41.40 | -1.73 | 12.60 | 16.11 | -3.51 |
Training-time dev
Best checkpoint recorded a dev WER of 30.72 at step 48000 /
epoch 12 (selected by "is_best": true).
Quick start: inference
import soundfile as sf
import torch
from transformers import SeamlessM4TFeatureExtractor, Wav2Vec2BertForCTC, Wav2Vec2CTCTokenizer
MODEL_ID = "anshulsc/w2v-bert-ethio-v5"
LID_TOKEN_IDS = [311, 312, 313, 314, 315, 316]
speech, sr = sf.read("audio.wav", dtype="float32")
if speech.ndim > 1:
speech = speech.mean(axis=1)
fe = SeamlessM4TFeatureExtractor.from_pretrained("facebook/w2v-bert-2.0") # 16 kHz mel features
if sr != fe.sampling_rate:
import librosa
speech = librosa.resample(speech, orig_sr=sr, target_sr=fe.sampling_rate)
inputs = fe(speech, sampling_rate=fe.sampling_rate, return_tensors="pt", return_attention_mask=True)
model = Wav2Vec2BertForCTC.from_pretrained(MODEL_ID, torch_dtype=torch.float16).eval()
tokenizer = Wav2Vec2CTCTokenizer.from_pretrained(MODEL_ID)
with torch.no_grad():
logits = model(**inputs).logits
pred_ids = logits.argmax(dim=-1)
# The head loves to echo its LID prefix at decode time; wipe those positions to blank first.
lid = torch.tensor(LID_TOKEN_IDS, device=pred_ids.device)
pred_ids = pred_ids.masked_fill(torch.isin(pred_ids, lid), tokenizer.pad_token_id)
print(tokenizer.decode(pred_ids[0]).strip())
inference_example.py in this repo is the same snippet as a runnable script
(python inference_example.py path/to/audio.wav).
Token ID reference
Vocabulary size is 319. Decoding is plain CTC, so no BOS/EOS is ever emitted β the head produces only chars, the word delimiter, and occasionally its conditioning token.
| Token | ID | Role |
|---|---|---|
[LID:amharic] |
311 | training-time language conditioning prefix |
[LID:tigrinya] |
312 | training-time language conditioning prefix |
[LID:oromo] |
313 | training-time language conditioning prefix |
[LID:somali] |
314 | training-time language conditioning prefix |
[LID:sidama] |
315 | training-time language conditioning prefix |
[LID:afar] |
316 | training-time language conditioning prefix |
[PAD] |
0 | CTC blank / padding |
[UNK] |
1 | unknown character |
| ` | ` | 29 |
<s> |
317 | BOS (unused at CTC inference) |
</s> |
318 | EOS (unused at CTC inference) |
Everything from id 2 through 310 is a literal character token (Latin letters, ' and | plus
Ge'ez syllabics). The | word delimiter decodes to a space.
Training
- Base:
facebook/w2v-bert-2.0with the feature extractor kept frozen at 16 kHz / 80 mel bins. - Head: fresh char-level
Wav2Vec2BertForCTC.lm_headover the custom 319-token vocab. - Conditioning: each label sequence is prefixed with
[LID:<lang>]|, which is why the decoder can emit LID tokens and why they must be masked to blank before CTC collapsing. - Recipe: full fine-tune (no LoRA), regularisation schedule from the v5 recipe.
Caveats
- Waxal numbers are zero-shot; EthioSpeech held-out numbers are in-domain. They are not directly comparable.
- Amharic and Tigrinya were normalised into a mixed-script convention (Ge'ez + latin | punctuation); transcribed hypotheses follow the same convention.
- Downloads last month
- 10
Model tree for ethiospeech01/w2v2-bert-ethio-v5
Base model
facebook/w2v-bert-2.0