Betterflow — English streaming FastConformer, ONNX for sherpa-onnx

An ONNX export of nvidia/stt_en_fastconformer_hybrid_large_streaming_multi, prepared so it loads in sherpa_onnx.OnlineRecognizer and produces live partials for English.

We are not the authors of the weights. Upstream is NVIDIA; this repo is a format conversion plus quantization.

Provenance and licence

Upstream nvidia/stt_en_fastconformer_hybrid_large_streaming_multi
Upstream licence CC-BY-4.0 (read off the model card, not inferred)
This repo CC-BY-4.0, inherited — attribution required
What changed .nemo → ONNX via k2-fsa's own export path · set_default_att_context_size([70,13]) · int8 quantization
What did NOT change the weights — no fine-tuning

Please cite NVIDIA for the underlying model.

Contents — a transducer bundle (three graphs)

encoder.int8.onnx   131,507,640 B      encoder.onnx   456,772,215 B
decoder.int8.onnx     3,955,863 B      decoder.onnx    15,753,087 B
joiner.int8.onnx      1,408,183 B      joiner.onnx      5,584,035 B
tokens.txt               11,896 B

int8 total ≈ 137 MB.

Measured

librispeech-en, n=50, through sherpa with a padded tail:

WER 7.7% pooled · 5.3% median
RTF 0.021
peak RSS 662 MB
empty 0/50
script 100% Latin

Sample decode (int8, 2 s tail pad):

'concord returned to its place amidst the tents'
'congratulations were poured in upon the princess everywhere during her journey'

⚠️ Three things worth knowing

1. Pad the tail — and this bundle tells you exactly how much. sherpa's online recogniser only decodes when num_frames_ready - num_processed >= window_size, and input_finished() does not pad to a whole window, so up to window_size - 1 frames of every utterance are never decoded.

This encoder declares window_size = 121, chunk_shift = 112, subsampling_factor = 8 in its ONNX metadata. At a 10 ms hop that is 1.21 s, so pad ≥ ~1.3 s; we use 2,000 ms. Shorter pads lose words as deletions, which read as poor model quality rather than as a configuration error.

‼️ Read window_size off the graph rather than copying a number. An earlier version of this card quoted a 0 → 20.7% · 500 → 10.4% · 2,000 → 5.4% sweep as if it were measured on this bundle. It was not — those are third-party figures from a streaming zipformer on Android, a different architecture whose chunk length we never read. The advice was right; the numbers were not ours.

2. downloadMb is not peakRssMb. 137 MB on disk, 662 MB resident — a 4.9× gap. Budget on the resident figure.

3. Peak RSS is FLAT in utterance length671.2 MB at 5 s, 671.5 MB at 240 s. No utterance-length cap is needed for this bundle.

The reason is the cache-aware architecture — bounded left context plus a fixed cache — not the fact that it streams. ‼️ "Streaming ⇒ bounded memory" is false as a general rule: we measured a streaming decoder-only model whose peak RSS scales T^1.49 and walls at ~10.6 s of audio. Flat memory is a property of this family (cache-aware Conformer), and an offline Conformer's attention is O(T²). Check the scaling; do not infer it from the word "streaming."

Not evaluated

Device/Android verification · lookaheads other than [70,13] (0/80/480 ms are exportable via the same script) · languages other than English · dictation-register audio — the numbers above are read speech.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support