phcodeswitch-ceb-dvo

Whisper small fine-tuned for automatic speech recognition (ASR) of Davao Cebuano, including English–Cebuano code-switching.

  • Base model: openai/whisper-small
  • Fine-tuned language: Cebuano (ceb), BCP-47 ceb
  • Task: Automatic speech recognition (transcribe)
  • Best WER: 20.86% (test split, step 1200)
  • License: CC BY-NC 4.0 — non-commercial research / educational use only

How to use

Install dependencies:

pip install --upgrade transformers torch librosa

Option 1 – pipeline (quickstart)

from transformers import pipeline
import librosa

pipe = pipeline(
    "automatic-speech-recognition",
    model="eemberda/phcodeswitch-ceb-dvo",
    device=0,  # use -1 for CPU
)

audio, sr = librosa.load("sample.wav", sr=16_000, mono=True)
result = pipe(
    audio,
    generate_kwargs={
        "language": "tl",   # Cebuano is not native to Whisper; Tagalog prompt works best
        "task": "transcribe",
        "num_beams": 5,
    },
)
print(result["text"])

Option 2 – manual inference with processor + model

from transformers import WhisperForConditionalGeneration, WhisperProcessor
import librosa
import torch

model_id = "eemberda/phcodeswitch-ceb-dvo"

processor = WhisperProcessor.from_pretrained(model_id)
model = WhisperForConditionalGeneration.from_pretrained(model_id)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device).eval()

audio, sr = librosa.load("sample.wav", sr=16_000, mono=True)
inputs = processor.feature_extractor(
    audio, sampling_rate=16_000, return_tensors="pt"
).input_features.to(device)

forced_decoder_ids = processor.get_decoder_prompt_ids(
    language="tl", task="transcribe"
)

with torch.no_grad():
    predicted_ids = model.generate(
        inputs,
        forced_decoder_ids=forced_decoder_ids,
        num_beams=5,
    )

transcription = processor.batch_decode(
    predicted_ids, skip_special_tokens=True
)[0].strip()
print(transcription)

Language notes

  • Whisper has no native Cebuano language token. The model is fine-tuned on Cebuano audio but uses the Tagalog (tl) decoder prompt, which Whisper treats as the closest supported related language.
  • The model also handles English and English–Cebuano code-switched speech.
  • Audio is expected at 16 kHz mono (resampled automatically by librosa in the examples above).

Training details

  • Base model: openai/whisper-small
  • Optimizer: AdamW, learning rate 1e-5, warmup 100 steps
  • Batch size 4 with gradient accumulation 4 (effective batch 16)
  • Max steps 1800 (best checkpoint at step 1200), early stopping patience 3
  • Mixed precision (fp16), gradient checkpointing, beam search (5) decoding

Limitations

  • Trained on a small, community-contributed dataset; coverage of accents and vocabulary is limited.
  • For non-commercial research and educational use only (CC BY-NC 4.0).
  • Contributed speaker data must not be used for voice cloning, impersonation, or voice synthesis. See the project repository's compliance documents.
Downloads last month
1,842
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for eemberda/phcodeswitch-ceb-dvo

Finetuned
(3677)
this model