google/fleurs
Viewer • Updated • 768k • 116k • 474
Fine-tuned openai/whisper-small for Hebrew automatic speech recognition (ASR).
| Parameter | Value |
|---|---|
| Batch size | 8 |
| Gradient accumulation | 4 |
| Effective batch size | 32 |
| Learning rate | 1e-5 |
| Scheduler | Linear |
| Warmup steps | 500 |
| Epochs | 3 |
| Precision | bf16 |
| Trainable params | 154M (decoder only) |
from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch
processor = WhisperProcessor.from_pretrained("dm15/whisper-small-hebrew")
model = WhisperForConditionalGeneration.from_pretrained("dm15/whisper-small-hebrew")
# Transcribe
input_features = processor(audio_array, sampling_rate=16000, return_tensors="pt").input_features
predicted_ids = model.generate(input_features, language="he", task="transcribe")
text = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]
Convert to CTranslate2 first:
ct2-whisper-converter --model dm15/whisper-small-hebrew --output_dir whisper-hebrew-ct2 --quantization float16
Then use faster-whisper for 4-5x faster inference:
from faster_whisper import WhisperModel
model = WhisperModel("whisper-hebrew-ct2", device="cuda", compute_type="float16")
segments, info = model.transcribe("audio.wav", language="he", beam_size=1)
text = " ".join(s.text for s in segments)
Base model
openai/whisper-small