Addis Scribe Streaming

Realtime Amharic speech recognition by Addis AI. Text appears while people speak, in 0.32-second steps, and is final about 30 ms after they stop. Built on NVIDIA Nemotron 3.5 ASR Streaming (0.6B), trained on 5,000+ hours of quality-checked Amharic speech.

Addis Scribe Streaming Google Chirp 3 (streaming)
Word error rate, natural speech 28.4% 32.5%
Character error rate, natural speech 13.1% 16.5%
Final text after the speaker stops 30 ms 1,411 ms
First words on screen 1.56 s 6.67 s

Streaming support

"Streaming" here means the model reads each new piece of audio once, keeps what it learned in a cache, and updates the text without re-reading earlier audio. Cost and delay stay flat however long someone talks.

System Streaming How it handles live audio
Addis Scribe Streaming Native Cache-aware: each 0.32 s (or 1.12 s) chunk is processed once.
Google Chirp 3 Native (cloud API) Server-side streaming over the network.
Shook (Whisper medium) Not supported Whisper models transcribe complete audio, not a live stream. Our benchmark runs it live using the method of WhisperLiveKit, a widely used open-source tool for live Whisper: once per second it transcribes all audio received so far and displays a word once two consecutive transcriptions contain it. Because each pass covers the whole recording so far, the delay grows the longer someone speaks: measured median 10.8 s after speech ends.
Hohe (wav2vec2-BERT) No Offline only: needs the whole recording before it can answer.

Benchmarks

Every system is scored with the same text normaliser (Unicode NFC, punctuation and symbols removed). WER = word error rate, CER = character error rate; lower is better.

Natural speech: 100 WAXAL Amharic test clips

22 speakers, 29 minutes, human transcripts. None of it was used in training. Clips were picked at random (fixed seed) from WAXAL amh test, 2 to 30 s long.

System Mode WER CER Final text after speech ends First words
Addis Scribe Streaming streaming 0.32 s 28.4% 13.1% 30 ms 1.56 s
Addis Scribe Streaming streaming 1.12 s 27.3% 12.4% 34 ms 2.21 s
Google Chirp 3 (am-ET) streaming, realtime 32.5% 16.5% 1,411 ms 6.67 s
Google Chirp 3 (am-ET) batch, whole file 29.2% 13.5% n/a n/a
Shook (Whisper medium) live, WhisperLiveKit method (1 s updates) 35.0% 16.6% 10,839 ms 3.66 s
Shook (Whisper medium) offline 35.0% 16.6% n/a n/a
Hohe (wav2vec2-BERT) offline only 25.4% 11.2% n/a n/a

Latency for Addis Scribe and Shook is measured on one NVIDIA A100 with no network hop, with audio fed at real-time speed; Google's figures include the network round trip to its EU endpoint from the benchmark machine.

FLEURS

System Mode WER CER
Addis Scribe Streaming streaming 0.32 s (first 150 clips) 19.8% 8.2%
Hohe offline (same 150 clips) 20.6% 7.9%
Addis Scribe Streaming offline (all 516 clips) 20.6% 7.7%
Hohe offline (all 516 clips) 20.2% 7.3%
Previous Addis AI Nemotron streaming 0.32 s (first 150 clips) 35.0% 17.0%

Every per-clip transcript, latency and the scoring script are in benchmark/, so these numbers can be reproduced exactly.

Quickstart

from stream_infer import AmharicStreamer

scribe = AmharicStreamer("addis-scribe-streaming.nemo", chunk="320ms")   # or "1120ms"
for frame in microphone_frames():        # float32 mono, 16 kHz, any frame size
    print(scribe.feed(frame))            # transcript so far
final_text = scribe.finish()             # at end of speech; scribe.reset() for the next utterance

Command line: python stream_infer.py addis-scribe-streaming.nemo speech_16k.wav [320ms|1120ms]. Requires NVIDIA NeMo 3.0 with numba-cuda[cu12], numpy<2.4 and nvidia-nvjitlink-cu12>=12.9. Fed 100 ms frames like a live microphone, AmharicStreamer reproduces the reference streaming transcripts exactly (100 of 100 benchmark clips).

Model

Architecture Cache-aware FastConformer encoder, hybrid RNNT + CTC, 0.6B parameters
Decoding Greedy RNNT; language prompt am-ET (index 49)
Chunk sizes 0.32 s (default) or 1.12 s
Input / output 16 kHz mono audio / Amharic text without punctuation
Files addis-scribe-streaming.nemo, serving.json, stream_infer.py, SHA256SUMS

Training

Trained on 5,000+ hours of quality-checked Amharic speech, then refined on human-transcribed speech. FLEURS, WAXAL test and the evaluation speakers were never used in training.

Limits

  • On natural conversational speech it is about 3 WER points behind Hohe, which is offline only.
  • No punctuation (training text had punctuation removed).
  • Tested on Amharic only; other languages are out of scope.

References

Base model NVIDIA Nemotron 3.5 ASR Streaming 0.6B
Toolkit NVIDIA NeMo
Google Chirp 3 Chirp 3 model documentation
Hohe snapwre/hohe-asr-amharic
Shook b1n1yam/shook-medium-amharic-2k
WhisperLiveKit QuentinFuxa/WhisperLiveKit
WAXAL (natural-speech test set) google/WaxalNLP, amh test split, CC-BY-SA-4.0
FLEURS (read-speech test set) google/fleurs, am_et test split

License

Built on NVIDIA Nemotron 3.5 ASR Streaming 0.6B, used under the OpenMDW-1.1 license. © Addis AI.

Downloads last month
134
GGUF
Model size
0.6B params
Architecture
asr
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for addisai/addis-scribe-streaming

Quantized
(64)
this model