faster-whisper conversion of SunflowerASR-51-african-languages

This is a CTranslate2 conversion of Sunbird/SunflowerASR-51-african-languages (a Whisper-based ASR model covering 51 African languages), quantized to int8 for use with faster-whisper. All credit for the model goes to Sunbird AI; see the original model card for training details, languages and evaluation results.

How it was converted

ct2-transformers-converter \
  --model Sunbird/SunflowerASR-51-african-languages \
  --output_dir faster-quantum-asr \
  --copy_files tokenizer.json processor_config.json generation_config.json \
  --quantization int8

Note on processor_config.json and preprocessor_config.json

The original repo has no preprocessor_config.json. It ships a small processor_config.json, which holds the feature extractor settings in the newer transformers processor format. faster-whisper does not read that file. It looks for preprocessor_config.json and, when it is missing, defaults to 80 mel bins.

This model expects 128 mel bins (large-v3 style), so transcription fails with:

ValueError: Invalid input features shape: expected an input with shape (1, 128, 3000), but got an input with shape (1, 80, 3000) instead

To fix this, this repo includes a hand-written preprocessor_config.json with the standard Whisper large-v3 feature extractor settings (feature_size: 128, 16 kHz, 30 s chunks). processor_config.json is kept alongside it for reference and for transformers compatibility. If you convert the model yourself, add this file to the output folder afterwards.

Usage

import json
from faster_whisper import WhisperModel
from huggingface_hub import hf_hub_download

repo = "rlabz/faster-quantum-asr"
lang_map = json.load(open(hf_hub_download(repo, "language_map.json")))

model = WhisperModel(repo, device="cpu", compute_type="int8")
segments, info = model.transcribe(
    "audio.mp3",
    language=lang_map["swa"],   # Swahili
    task="transcribe",
    beam_size=5,
    vad_filter=False,
    condition_on_previous_text=False,
)
for segment in segments:
    print(f"[{segment.start:.2f} -> {segment.end:.2f}] {segment.text}")

language_map.json maps language codes (e.g. swa, lug, eng) to the language tokens the model expects. compute_type can be changed at load time (e.g. float16 on GPU); int8 weights are converted on the fly.

Files

File Purpose
model.bin, config.json CTranslate2 int8 weights and model config
tokenizer.json Tokenizer (copied from the original)
generation_config.json Generation settings (copied from the original)
processor_config.json Original processor config (copied, not read by faster-whisper)
preprocessor_config.json Added by hand: 128-mel feature extractor settings for faster-whisper
language_map.json Language code to language token mapping

License and acknowledgements

This conversion is released under the Apache 2.0 licence, the same as Sunbird/SunflowerASR-51-african-languages, which follows the licence of its Whisper large-v3 base model. This repo only changes the weight format and precision (int8 CTranslate2) and adds the config files described above. It does not change the model's behaviour by design.

The original model was trained on publicly available and community-collected speech datasets, credited in the Training Datasets section of the original model card. If you use this model, please also respect the licence and citation requirements of those sources.

Thanks to Sunbird AI for the model, and to the organisations and community contributors who collected and released the speech data that made it possible.

Citation

This repo is a format conversion. If you use this model, please cite the original work by Sunbird AI:

@misc{sunflowerasr2026,
  title        = {Group Relative Policy Optimisation Improves Multilingual Speech
                  Recognition in Low-Resource Languages},
  author       = {Hu, Tim Wenjie and Ouma, Evelyn Nafula and Akera, Benjamin and
                  Quinn, John A.},
  year         = {2026},
  institution  = {Sunbird AI},
  howpublished = {\url{https://huggingface.co/Sunbird/SunflowerASR-51-african-languages}}
}

And, if you use the evaluation set:

@misc{sunbird_speech_benchmark_2026,
  title        = {Sunbird Speech Benchmark: a multilingual ASR test set for 51
                  African languages},
  author       = {Sunbird AI},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/Sunbird/speech-benchmark}}
}
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rlabz/faster-quantum-asr