Jackrabbit, the ears of Surogate Speech 🎧

Community Article
Published September 25, 2026

Jackrabbit 110M

Jackrabbit does the listening in Surogate Speech, one model per language, and Romanian goes first. I ran five speech recognisers over the Romanian FLEURS test set with the Open ASR Leaderboard's own runner. Ours came first, at less than an eighth the size of the next one.

  • Surogate Jackrabbit 110M, CTC + 4-gram: 5.69% WER, 2,531× real time on one RTX 5090
  • NVIDIA Canary 1B v2: 5.95%
  • Jackrabbit 110M, TDT greedy, no LM: 7.56%, 2,774× real time
  • Whisper large-v3: 8.42%, 102× real time
  • SpeD 110M: 8.60%
  • Parakeet TDT 0.6B v3: 11.58%

883 clips, bf16, the leaderboard's MultilingualNormalizer, commit b2e6f04, batch 64 for the NeMo models and 32 for Whisper.

What's in the Romanian release

Two models from one 116M FastConformer, both writing cased, punctuated Romanian:

Both are plain .nemo files and run fine on a CPU. The offline model also ships as GGUF for NVIDIA's NeMo-Speech.cpp runtime, no PyTorch needed, with the same accuracy as the .nemo in TDT mode.

How it was measured

Romanian has no speech leaderboard, so we ran the public protocols ourselves and put every model through the same code.

The FLEURS numbers above use the leaderboard's nemo_asr/run_eval_ml.py untouched for every row except one. For Jackrabbit's CTC + LM row we added seven lines that switch the decoder, and they're published with the results. The first time I added them the runner reset the decoder a few lines later and the "LM" run came back identical to greedy, same WER and same speed. Moving the switch after the runner's own setup fixed it.

Then SpeD's Common Voice 21 clip list (3,929 clips), scored with their normalizer and with the leaderboard's:

Model leaderboard norm SpeD norm
Jackrabbit 110M, CTC + 4-gram 2.07% 2.19%
Jackrabbit 110M, TDT greedy 2.19% 2.30%
SpeD 110M 3.47% 3.58%
Parakeet TDT 0.6B v3 10.06% 10.19%
NVIDIA Canary 1B v2 8.65% 8.87%
OpenAI Whisper large-v3 8.94% 9.16%

Their paper says 3.57% for SpeD on that set and our run says 3.58%, so the scoring checks out.

Streaming, scored on what you actually get

The streaming model shows partial text while you talk. When you pause, it re-reads the whole utterance with full attention and the LM and writes the final sentence. We score those finals.

Live 32 ms packets over FLEURS-ro:

  • final text: 7.03% WER
  • final sentence: about 0.72 s after you stop talking on an RTX 5090. That's 640 ms of silence to decide you've paused, then 84 ms to finalize (p95 182 ms).

The leaderboard runner has no idea a model streams. It only sees the partial text, which scores 10.67% on FLEURS-ro. It's in the results with that label.

Run it

surogate serve --stt surogate/jackrabbit-110m-ro             # files, /v1/audio/transcriptions
surogate serve --stt surogate/jackrabbit-110m-ro-streaming   # live, /v1/audio/streams

pip install "surogate-speech[asr] @ git+https://github.com/invergent-ai/surogate-speech"
surogate-speech transcribe interviu.wav --lm
surogate-speech listen

Or straight from NeMo:

from nemo.collections.asr.models import ASRModel
model = ASRModel.from_pretrained("surogate/jackrabbit-110m-ro")
print(model.transcribe(["audio.wav"])[0].text)

Transcripts for every clip and every model are in surogate-speech-evals, and the harness is in surogate-speech. Weights are CC-BY-NC-4.0, commercial licensing through Invergent. Fine-tuned from NVIDIA's parakeet-tdt_ctc-110m.

Community

Sign up or log in to comment