whisper-tiny-id

openai/whisper-tiny (39M params) fine-tuned for Indonesian, for realtime ASR on weak ARM CPUs such as the Cortex-A53 (Orange Pi Zero3, Raspberry Pi 3 class). Code: github.com/porcupine-md/whisper-id. Bigger siblings: whisper-small-id.

Training: 284.8 h (Common Voice 17 id, FLEURS id_id, YODAS2 id000 filtered with large-v3-turbo, YODAS labels = turbo transcripts), 8000 steps; telephone / Opus / AAC / noise augmentation; 80% of batches with short encoder context (mel cut to the clip, as whisper.cpp -ac).

Results

WER / CER (%), full test sets, lowercase + punctuation stripped:

Test set whisper-tiny whisper-tiny-id whisper-tiny-id, short context
FLEURS id (684 utts) 61.6 / 27.6 19.1 / 6.3 19.2 / 6.4
Common Voice 17 id (3629 utts) 59.1 / 30.1 18.0 / 6.7 18.3 / 6.8

Latency, whisper.cpp q8_0, 4 s utterance:

CPU Settings ms / utterance RTF
Cortex-A53 x4 1.4 GHz (Orange Pi Zero3) -t 4 -ac 256 898 0.22
Cortex-A53 x4 -t 3 -ac 256 1095 0.27

Output is lowercase without punctuation.

Files

  • model.safetensors (fp16) + tokenizer/config: transformers
  • ggml-tiny-id-q8_0.bin (43 MB): whisper.cpp. On ARM, q8_0 beats q5_0 on speed.

Usage

# -ac = encoder positions: 50 per second of audio, round up to a multiple of 64 (4 s -> 256, 8 s -> 448)
whisper-cli -m ggml-tiny-id-q8_0.bin -l id -t 4 -ac 256 -bs 1 -bo 1 -nt -f audio16k.wav
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="anak10thn/whisper-tiny-id")
print(asr("audio.wav", generate_kwargs={"language": "indonesian", "task": "transcribe"})["text"])

Limitations

  • About 6 WER points worse on telephone-band audio than on clean audio (training-time eval).
  • Test sets are read speech; conversational audio will be harder.
  • Orange Pi Zero3 has no thermal throttling (only a 110 C shutdown); sustained 4-thread decoding runs near 80 C. Use a heatsink, or -t 3.
Downloads last month
3
Safetensors
Model size
37.8M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for anak10thn/whisper-tiny-id

Finetuned
(1918)
this model

Datasets used to train anak10thn/whisper-tiny-id