UBUZIMA AI โ€” Kinyarwanda ASR LoRA adapter

A LoRA adapter over badrex/w2v-bert-2.0-kinyarwanda-asr, fine-tuned for a Kinyarwanda voice health assistant. Part of the UBUZIMA AI capstone project at the African Leadership University, Kigali.

Model details

  • Developed by: Ganza Didier (African Leadership University)
  • Model type: LoRA adapter for Wav2Vec2-BERT 2.0 CTC speech recognition
  • Language: Kinyarwanda (rw)
  • License: Apache 2.0
  • Fine-tuned from: badrex/w2v-bert-2.0-kinyarwanda-asr

Adapter configuration

r = 32, alpha = 64, dropout 0.05, target modules linear_q and linear_v. 3,145,728 trainable parameters (0.54% of the 583.6M base). Feature extractor and feature projection frozen.

Training

Fine-tuned in two sequential stages on pooled Afrivoice (health-domain) and Common Voice (general-domain) Kinyarwanda speech:

  • Stage 1: 100 hours at a 1:1 Afrivoice:Common Voice ratio.
  • Stage 2: warm-started from Stage 1, 500 hours at a 1:4 ratio. This adapter is the Stage 2 result.

Training regime: fp16 mixed precision on an NVIDIA T4 (Google Colab).

Evaluation

Test set n clips WER 95% CI CER
Afrivoice (in-domain, health) 1,560 6.76% [6.44, 7.11] 1.39%
Common Voice (cross-domain, general) 16,205 31.29% [30.86, 31.72] 8.99%

WER = word error rate, CER = character error rate. 95% confidence intervals from 500-resample bootstrapping.

Contamination note. The base badrex checkpoint was pretrained on Kinyarwanda audio that included portions of the Digital Umuganda health corpus from which the Afrivoice subset is drawn. The in-domain figure therefore reflects both this adaptation and residual pretraining familiarity; the Common Voice figure is the cleaner cross-domain measure of what the adapter added.

How to use

from transformers import Wav2Vec2BertForCTC, AutoProcessor
from peft import PeftModel

base = "badrex/w2v-bert-2.0-kinyarwanda-asr"
model = Wav2Vec2BertForCTC.from_pretrained(base)
model = PeftModel.from_pretrained(model, "dganza/Ubuzima_AI")
model = model.merge_and_unload()          # fold LoRA in for faster inference
processor = AutoProcessor.from_pretrained("dganza/Ubuzima_AI")

Intended use and limitations

Research prototype for Kinyarwanda health-information delivery. Not a medical device; does not diagnose. Dialect coverage is limited to the standard Kinyarwanda represented in the training corpora. Cross-domain performance (Common Voice, 31% WER) is substantially weaker than in-domain (Afrivoice, 7% WER), so out-of-domain use should expect that gap.

Citation

Ganza, D. (2026). UBUZIMA AI: A Kinyarwanda voice health assistant [Capstone project]. African Leadership University. Supervisor: Emmanuel Adjei.

Framework versions

  • PEFT 0.12.0
  • Transformers 4.44.2
Downloads last month
36
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for dganza/Ubuzima_AI

Adapter
(1)
this model