Ozen-v1

Ozen (ืื•ื–ืŸ, "ear") is a small Hebrew speech recognition model built to run on a phone or in a browser. At 60.7M parameters, 13 times smaller than the ivrit.ai model it learned from, it reaches 8.59% word error rate on ivrit.ai's eval-d1 and 15.94% on WhatsApp voice messages, and transcribes at about 40 times realtime on four CPU threads.

Results

Scored with a harness that reproduces the ivrit.ai Hebrew leaderboard on 40 of 40 published model and dataset pairs.

Benchmark Ozen-v1, 61M whisper-small, 242M whisper-base, 73M ivrit.ai turbo, 809M
ivrit-ai/eval-d1 8.59% 29.93% 48.32% 5.5%
ivrit-ai/eval-whatsapp 15.94% 39.64% 56.60% 6.1%
imvladikon/hebrew_speech_kan 13.40% 37.42% 68.12% 8.10%

The ivrit.ai column is the model Ozen was distilled from, and the ceiling it is measured against: it closes much of the gap at a fraction of the size, not all of it. Stock Whisper at comparable sizes does not work for Hebrew.

How it was built

Distillation from ivrit.ai. ivrit.ai's Hebrew fine-tune of Whisper large-v3-turbo transcribed 3,519 hours of ivrit-ai/audio-v2, mostly podcasts and interviews, decoded as whole episodes, cut into windows of up to 28 seconds and filtered on the decoder's own confidence. No source contributes more than 10% of the audio. Ozen learns from those transcripts plus 356 hours of human transcription. Teacher labels on conversational speech were chosen over the much larger Knesset corpus, whose plenum protocols are aligned records rather than verbatim speech.

Architecture. A Whisper encoder of 12 layers and a decoder of 4, width 512, with an 8,192-token Hebrew byte-level BPE in place of Whisper's multilingual vocabulary (1.76 tokens per Hebrew word against Whisper's 3.17). It descends from openai/whisper-base: the base model was given the Hebrew tokenizer, doubled in depth, trained on Hebrew, and then cut to four decoder layers. The decoder depth was chosen by measurement rather than convention:

Decoder layers Held-out WER after 12k steps CPU time per 30 s window
2 24.95% 0.59 s
3 21.31% 0.66 s
4 20.13% 0.73 s

The encoder dominates inference cost, so each extra decoder layer is cheap while buying about a point of accuracy. Learning rate was swept the same way (5e-5, 1e-4 and 2e-4; 1e-4 kept).

Training. 75% teacher labels, 15% ivrit-ai/crowd-transcribe-v5, 4% ivrit-ai/crowd-recital, 3% google/fleurs he and 3% imvladikon/hebrew_speech_kan, as shares of audio heard. Schedule-free AdamW, learning rate 1e-4, effective batch 16, bf16, on a single RTX 5070 Laptop GPU. The weights are from step 220,000 of 303,562, chosen on 15 hours of held-out podcast audio never trained on, where it scores 13.24% against the teacher's own transcripts.

Usage

from transformers import AutoModelForSpeechSeq2Seq, AutoTokenizer, AutoFeatureExtractor
from transformers.generation.utils import GenerationMixin

model = AutoModelForSpeechSeq2Seq.from_pretrained("itayinbar/Ozen-v1")
tokenizer = AutoTokenizer.from_pretrained("itayinbar/Ozen-v1")
features = AutoFeatureExtractor.from_pretrained("itayinbar/Ozen-v1")

inputs = features(audio_16khz, sampling_rate=16000,
                  return_tensors="pt", padding="max_length")
# Monolingual: no language token, no task token, no timestamps.
ids = GenerationMixin.generate(model, **inputs, max_new_tokens=220)
print(tokenizer.decode(ids[0], skip_special_tokens=True))

Three things differ from stock Whisper:

  1. Call the generic generate. There are no language or task tokens.
  2. Pad mel features to the full 30-second window.
  3. Cut audio longer than 30 seconds into overlapping windows and stitch them; one second of overlap with a repeated phrase dropped at the seam works well.

For the browser (transformers.js, WebGPU or WASM), load the fp16 encoder with the fp32 decoder, a 164.3 MB download. Only those weights are shipped: an fp16 decoder fails to load in ONNX Runtime and an int8 encoder loads but changes the transcript. Every combination offered here was checked by transcribing with it and comparing against the fp32 output.

Limitations

  • It inherits the teacher's conventions: punctuation, and English words sometimes written in Latin letters.
  • It transcribes Hebrew only. The teacher it learned from writes English speech as Hebrew-letter transliteration, so Ozen may do the same.
  • On fast, dense speech a small decoder can occasionally drop a phrase inside a window. This was seen in a two-layer predecessor; the four-layer decoder may reduce it, but that has not been measured separately.

Licence and provenance

Weights are Apache-2.0, following openai/whisper-base and ivrit-ai/whisper-large-v3-turbo.

Training audio and transcripts from ivrit.ai under the ivrit.ai licence, which permits training models including commercially and requires attribution. FLEURS is CC-BY-4.0. imvladikon/hebrew_speech_kan declares no licence on the Hub; it contributed 3% of the audio heard and is named so anyone relying on provenance can judge it.

Downloads last month
24
Safetensors
Model size
60.7M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for itayinbar/Ozen-v1

Quantized
(249)
this model