Instructions to use itayinbar/Ozen-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use itayinbar/Ozen-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="itayinbar/Ozen-v1")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("itayinbar/Ozen-v1") model = AutoModelForSpeechSeq2Seq.from_pretrained("itayinbar/Ozen-v1", device_map="auto") - Transformers.js
How to use itayinbar/Ozen-v1 with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('automatic-speech-recognition', 'itayinbar/Ozen-v1'); - Notebooks
- Google Colab
- Kaggle
Ozen-v1
Ozen (ืืืื, "ear") is a small Hebrew speech recognition model built to run on a phone or in a browser. At 60.7M parameters, 13 times smaller than the ivrit.ai model it learned from, it reaches 8.59% word error rate on ivrit.ai's eval-d1 and 15.94% on WhatsApp voice messages, and transcribes at about 40 times realtime on four CPU threads.
Results
Scored with a harness that reproduces the ivrit.ai Hebrew leaderboard on 40 of 40 published model and dataset pairs.
| Benchmark | Ozen-v1, 61M | whisper-small, 242M | whisper-base, 73M | ivrit.ai turbo, 809M |
|---|---|---|---|---|
ivrit-ai/eval-d1 |
8.59% | 29.93% | 48.32% | 5.5% |
ivrit-ai/eval-whatsapp |
15.94% | 39.64% | 56.60% | 6.1% |
imvladikon/hebrew_speech_kan |
13.40% | 37.42% | 68.12% | 8.10% |
The ivrit.ai column is the model Ozen was distilled from, and the ceiling it is measured against: it closes much of the gap at a fraction of the size, not all of it. Stock Whisper at comparable sizes does not work for Hebrew.
How it was built
Distillation from ivrit.ai. ivrit.ai's Hebrew fine-tune of Whisper
large-v3-turbo transcribed 3,519 hours of ivrit-ai/audio-v2, mostly podcasts
and interviews, decoded as whole episodes, cut into windows of up to 28 seconds
and filtered on the decoder's own confidence. No source contributes more than
10% of the audio. Ozen learns from those transcripts plus 356 hours of human
transcription. Teacher labels on conversational speech were chosen over the
much larger Knesset corpus, whose plenum protocols are aligned records rather
than verbatim speech.
Architecture. A Whisper encoder of 12 layers and a decoder of 4, width 512,
with an 8,192-token Hebrew byte-level BPE in place of Whisper's multilingual
vocabulary (1.76 tokens per Hebrew word against Whisper's 3.17). It descends
from openai/whisper-base: the base model was given the Hebrew tokenizer,
doubled in depth, trained on Hebrew, and then cut to four decoder layers. The
decoder depth was chosen by measurement rather than convention:
| Decoder layers | Held-out WER after 12k steps | CPU time per 30 s window |
|---|---|---|
| 2 | 24.95% | 0.59 s |
| 3 | 21.31% | 0.66 s |
| 4 | 20.13% | 0.73 s |
The encoder dominates inference cost, so each extra decoder layer is cheap while buying about a point of accuracy. Learning rate was swept the same way (5e-5, 1e-4 and 2e-4; 1e-4 kept).
Training. 75% teacher labels, 15% ivrit-ai/crowd-transcribe-v5, 4%
ivrit-ai/crowd-recital, 3% google/fleurs he and 3%
imvladikon/hebrew_speech_kan, as shares of audio heard. Schedule-free AdamW,
learning rate 1e-4, effective batch 16, bf16, on a single RTX 5070 Laptop GPU.
The weights are from step 220,000 of 303,562, chosen on 15 hours of held-out
podcast audio never trained on, where it scores 13.24% against the teacher's
own transcripts.
Usage
from transformers import AutoModelForSpeechSeq2Seq, AutoTokenizer, AutoFeatureExtractor
from transformers.generation.utils import GenerationMixin
model = AutoModelForSpeechSeq2Seq.from_pretrained("itayinbar/Ozen-v1")
tokenizer = AutoTokenizer.from_pretrained("itayinbar/Ozen-v1")
features = AutoFeatureExtractor.from_pretrained("itayinbar/Ozen-v1")
inputs = features(audio_16khz, sampling_rate=16000,
return_tensors="pt", padding="max_length")
# Monolingual: no language token, no task token, no timestamps.
ids = GenerationMixin.generate(model, **inputs, max_new_tokens=220)
print(tokenizer.decode(ids[0], skip_special_tokens=True))
Three things differ from stock Whisper:
- Call the generic
generate. There are no language or task tokens. - Pad mel features to the full 30-second window.
- Cut audio longer than 30 seconds into overlapping windows and stitch them; one second of overlap with a repeated phrase dropped at the seam works well.
For the browser (transformers.js, WebGPU or WASM), load the fp16 encoder with the fp32 decoder, a 164.3 MB download. Only those weights are shipped: an fp16 decoder fails to load in ONNX Runtime and an int8 encoder loads but changes the transcript. Every combination offered here was checked by transcribing with it and comparing against the fp32 output.
Limitations
- It inherits the teacher's conventions: punctuation, and English words sometimes written in Latin letters.
- It transcribes Hebrew only. The teacher it learned from writes English speech as Hebrew-letter transliteration, so Ozen may do the same.
- On fast, dense speech a small decoder can occasionally drop a phrase inside a window. This was seen in a two-layer predecessor; the four-layer decoder may reduce it, but that has not been measured separately.
Licence and provenance
Weights are Apache-2.0, following openai/whisper-base and
ivrit-ai/whisper-large-v3-turbo.
Training audio and transcripts from ivrit.ai under the
ivrit.ai licence, which permits training
models including commercially and requires attribution. FLEURS is CC-BY-4.0.
imvladikon/hebrew_speech_kan declares no licence on the Hub; it contributed 3%
of the audio heard and is named so anyone relying on provenance can judge it.
- Downloads last month
- 24
Model tree for itayinbar/Ozen-v1
Base model
openai/whisper-base