eemberda's picture
Upload README.md with huggingface_hub
fa7f84e verified
|
Raw
History Blame Contribute Delete
3.38 kB
---
license: cc-by-nc-4.0
language:
- tl
- ceb
- en
base_model:
- openai/whisper-small
pipeline_tag: automatic-speech-recognition
tags:
- cebuano
- filipino
- tagalog
- cebuano-asr
- bisaya
- bisaya-asr
- code-switching
---
# phcodeswitch-ceb-dvo
Whisper `small` fine-tuned for automatic speech recognition (ASR) of **Davao Cebuano**, including English–Cebuano code-switching.
- **Base model:** [openai/whisper-small](https://huggingface.co/openai/whisper-small)
- **Fine-tuned language:** Cebuano (`ceb`), BCP-47 `ceb`
- **Task:** Automatic speech recognition (`transcribe`)
- **Best WER:** 20.86% (test split, step 1200)
- **License:** CC BY-NC 4.0 — non-commercial research / educational use only
## How to use
Install dependencies:
```bash
pip install --upgrade transformers torch librosa
```
### Option 1 – `pipeline` (quickstart)
```python
from transformers import pipeline
import librosa
pipe = pipeline(
"automatic-speech-recognition",
model="eemberda/phcodeswitch-ceb-dvo",
device=0, # use -1 for CPU
)
audio, sr = librosa.load("sample.wav", sr=16_000, mono=True)
result = pipe(
audio,
generate_kwargs={
"language": "tl", # Cebuano is not native to Whisper; Tagalog prompt works best
"task": "transcribe",
"num_beams": 5,
},
)
print(result["text"])
```
### Option 2 – manual inference with processor + model
```python
from transformers import WhisperForConditionalGeneration, WhisperProcessor
import librosa
import torch
model_id = "eemberda/phcodeswitch-ceb-dvo"
processor = WhisperProcessor.from_pretrained(model_id)
model = WhisperForConditionalGeneration.from_pretrained(model_id)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device).eval()
audio, sr = librosa.load("sample.wav", sr=16_000, mono=True)
inputs = processor.feature_extractor(
audio, sampling_rate=16_000, return_tensors="pt"
).input_features.to(device)
forced_decoder_ids = processor.get_decoder_prompt_ids(
language="tl", task="transcribe"
)
with torch.no_grad():
predicted_ids = model.generate(
inputs,
forced_decoder_ids=forced_decoder_ids,
num_beams=5,
)
transcription = processor.batch_decode(
predicted_ids, skip_special_tokens=True
)[0].strip()
print(transcription)
```
## Language notes
- Whisper has **no native Cebuano language token**. The model is fine-tuned on
Cebuano audio but uses the **Tagalog (`tl`)** decoder prompt, which Whisper
treats as the closest supported related language.
- The model also handles English and English–Cebuano code-switched speech.
- Audio is expected at **16 kHz mono** (resampled automatically by `librosa`
in the examples above).
## Training details
- Base model: `openai/whisper-small`
- Optimizer: AdamW, learning rate `1e-5`, warmup 100 steps
- Batch size 4 with gradient accumulation 4 (effective batch 16)
- Max steps 1800 (best checkpoint at step 1200), early stopping patience 3
- Mixed precision (fp16), gradient checkpointing, beam search (5) decoding
## Limitations
- Trained on a small, community-contributed dataset; coverage of accents and
vocabulary is limited.
- For non-commercial research and educational use only (CC BY-NC 4.0).
- Contributed speaker data must not be used for voice cloning, impersonation,
or voice synthesis. See the project repository's compliance documents.