google/fleurs
Viewer • Updated • 768k • 110k • 473
How to use anak10thn/whisper-tiny-id with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition", model="anak10thn/whisper-tiny-id") # Load model directly
from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq
processor = AutoProcessor.from_pretrained("anak10thn/whisper-tiny-id")
model = AutoModelForSpeechSeq2Seq.from_pretrained("anak10thn/whisper-tiny-id", device_map="auto")openai/whisper-tiny (39M params) fine-tuned for Indonesian, for realtime ASR on weak ARM CPUs such as the Cortex-A53 (Orange Pi Zero3, Raspberry Pi 3 class). Code: github.com/porcupine-md/whisper-id. Bigger siblings: whisper-small-id.
Training: 284.8 h (Common Voice 17 id, FLEURS id_id, YODAS2 id000 filtered with large-v3-turbo, YODAS labels = turbo transcripts), 8000 steps; telephone / Opus / AAC / noise augmentation; 80% of batches with short encoder context (mel cut to the clip, as whisper.cpp -ac).
WER / CER (%), full test sets, lowercase + punctuation stripped:
| Test set | whisper-tiny | whisper-tiny-id | whisper-tiny-id, short context |
|---|---|---|---|
| FLEURS id (684 utts) | 61.6 / 27.6 | 19.1 / 6.3 | 19.2 / 6.4 |
| Common Voice 17 id (3629 utts) | 59.1 / 30.1 | 18.0 / 6.7 | 18.3 / 6.8 |
Latency, whisper.cpp q8_0, 4 s utterance:
| CPU | Settings | ms / utterance | RTF |
|---|---|---|---|
| Cortex-A53 x4 1.4 GHz (Orange Pi Zero3) | -t 4 -ac 256 |
898 | 0.22 |
| Cortex-A53 x4 | -t 3 -ac 256 |
1095 | 0.27 |
Output is lowercase without punctuation.
model.safetensors (fp16) + tokenizer/config: transformersggml-tiny-id-q8_0.bin (43 MB): whisper.cpp. On ARM, q8_0 beats q5_0 on speed.# -ac = encoder positions: 50 per second of audio, round up to a multiple of 64 (4 s -> 256, 8 s -> 448)
whisper-cli -m ggml-tiny-id-q8_0.bin -l id -t 4 -ac 256 -bs 1 -bo 1 -nt -f audio16k.wav
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="anak10thn/whisper-tiny-id")
print(asr("audio.wav", generate_kwargs={"language": "indonesian", "task": "transcribe"})["text"])
-t 3.Base model
openai/whisper-tiny