FUTO asr4all medium (streaming)

This efficiency-focused model (55.6M params: 50.3M acoustic + 5.3M PCEC) is optimized for edge, mobile and CPU inference that opens automatic speech recognition to all platforms. It is an excellent choice for mobile environments, where low-power consumption is desirable. It comes bundled with a secondary model that applies punctuation, capitalization and light error correction.

These models use a custom, unique architecture that significantly increases the realtime factor (RTFx) compared to other ASR models of comparable size. It is heavily optimized for CPU inference, and delegates operations cleanly to vectorized SIMD and NEON operations.

The model has low and high latency streaming inference modes from the same weights, with runtime-configurable latency (details in the streaming usage section below). A small language model adds punctuation, capitalization and error correction (PCEC) capability. We also include an ultralight VAD for use in applications that need low power voice activity detection.

We export weights for a number of different formats, quantizations and runtimes. This model supports English language only. Commercial use is permitted with limitations for ethical restrictions prohibiting certain activities (eg. mass surveillance). See license for details.

Contents

Benchmark results

Deployment RTFx across the model family and platforms, single-thread

Open ASR Leaderboard dataset results (WER %, c256r64, lower is better):

Word error rate across eight benchmarks for all three model sizes

RTFx (single thread, higher is better, 120 s clip, median of 3), one row per recommended platform configuration:

platform backend stream_c16r4 RTFx stream_c512r8 RTFx
Pixel 4 executorch-xnnpack int8 7.6 16.5
Pixel 10 executorch-xnnpack int8 17.5 50.3
Raspberry Pi 4 executorch-xnnpack int8 1.2 2.4
i9-12900K executorch-xnnpack int8 24.6 64.7
i9-12900K executorch-xnnpack fp32 14.5 39.8
i9-12900K openvino fp32 16.7 44.2
Ryzen 7 PRO 8700GE openvino bf16 24.6 75.8
i9-12900K openvino-igpu fp16 9.8 43.0
RTX 6000 Max-Q (batch 16) cuda 28412.8 27988.3

RTFx is the acoustic decoding. The PCEC adds 2–7 % inference time on x86 at one thread.

Model details

  • Architecture: DFSMN conv superblocks + LDSA & sliding-window attention, causal squeeze-excitation (SE). CTC over BPE-256+blank, 25hz mel features.
  • Streaming: per-layer left-context caches. Each audio frame is encoded once.
  • Distilled from the Granite-Speech teacher (futo-org/gravel-ctc-440m, logit + hidden-state KD).

Voice activity detection (wake)

The model also includes an ultralight VAD, which operates independently of the ASR. This can be used for things like wake detection or triggering the ASR. It is based on kiloVAD (arXiv:2607.25870) and pruned to 2,537 parameters. Each call takes 1s of audio and returns 21 predictions at 40 ms intervals, each describing 200 ms of context.

Our implementation follows the kiloVAD architecture as described in the paper. This model is trained on permissively-licensed data and ships under this repository's license.

benchmark score
AVA-Speech AUC (full 151 clips) 0.8545
AMI frame error (16 meetings) 14.30 %

There are two choices for inference, optimized per architecture. For ARM targets, use vad_dft, while vad is faster for x86.

x86-64 (AVX-512) arm64 (NEON)
vad (conv-STFT) 0.26 ms 1.13 ms
vad_dft (in-graph DFT) 0.65 ms 0.65 ms

PCEC: punctuation, capitalization and error correction

A 5.3M-parameter second-pass head that runs at finalize() and turns the raw CTC hypothesis into formatted text. It reads the decoded tokens and acoustic hidden state pooled from the encoder.

Features:

  • Punctuation: periods, commas, question marks
  • Capitalization: proper nouns and acronyms from the model. Sentence-initial capitalization is applied by the decode examples (first word, and after . / ?).
  • Error correction: repairs the trunk's characteristic mistakes: garbled non-words from greedy CTC (dropped letters, fused or split words), wrong-word substitutions that sound like the intended word, and homophone disambiguation from sentence context (their/there, no/know, weather/whether).

PCEC is a token-space model that runs incrementally as tokens are committed. Its output trails the raw transcript by a small, fixed lag (commit + lookahead + settle, in tokens). Words behind that lag are settled and no longer change. The tail is provisional and is updated as more audio arrives. At the end of a clip a flush settles the remaining tail.

Live output is punctuated and cased, not only the final text. Every runtime here ships it: transformers transcribe(), the ONNX file and the ExecuTorch bundles have it on by default with a flag to disable. Its measured cost is the note under the RTFx table above.

Effect per dataset for this size (WER %, negative delta = better):

dataset plain CTC with PCEC delta
AMI (cleaned) 11.58 11.18 -0.40
GigaSpeech (cleaned) 11.27 10.63 -0.64
VoxPopuli (cleaned) 4.44 4.57 +0.13
Earnings-22 (cleaned, chunked) 11.12 10.48 -0.64
LibriSpeech clean 2.60 2.72 +0.12
LibriSpeech other 7.17 7.22 +0.05
SPGISpeech 5.32 5.15 -0.17
Monsoon en-IN 9.75 9.01 -0.74
macro average 7.91 7.62 -0.29

Exported artifacts

format file(s) target
transformers model.safetensors + modeling_asr4all.py Python / research
ONNX onnx/ (one graph per component, shared weights) general purpose (ORT, OpenVINO)
ExecuTorch XNNPACK int8 executorch/xnnpack_int8/ Mobile (ARM) / lightweight (GPTQ-style rounding)
ExecuTorch XNNPACK fp32 executorch/xnnpack_fp32/ Desktop, Laptop

Recommended configuration per platform

platform run this recommended sizes
Android / modern ARM incl. Raspberry Pi 5 (v8.2+) ExecuTorch xnnpack_int8 all sizes feasible on most devices
Raspberry Pi 4 / ARMv8.0 ExecuTorch xnnpack_int8 use small for low latency streaming
x86 CPU (Intel and AMD) OpenVINO on the onnx/ graphs, fp32 or bf16 (if AVX512-BF16 available) all sizes
Intel iGPU OpenVINO all sizes
NVIDIA GPU transformers, bf16 all sizes

Usage: transformers (one-shot)

input_values is raw 16 kHz mono PCM: the mel frontend lives inside the model, so the paired feature extractor only pads and masks.

from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="futo-org/asr4all-m", trust_remote_code=True)
print(asr("clip.wav")["text"])

transcribe takes a 16 kHz mono float tensor:

import torch
from transformers import AutoModel

# option A: soundfile (lighter install: `pip install soundfile`):
import soundfile as sf
audio, sr = sf.read("clip.wav", dtype="float32")   # numpy [samples] (mono)
wav = torch.from_numpy(audio)

# option B: torchaudio (needs a backend: torchcodec / ffmpeg):
# import torchaudio; wav = torchaudio.load("clip.wav")[0][0]

model = AutoModel.from_pretrained("futo-org/asr4all-m", trust_remote_code=True)
print(model.transcribe(wav))                    # punctuated + cased (PCEC, default)
print(model.transcribe(wav, punctuate=False))   # plain lowercase CTC decode
print(model.detect_speech(wav[None])[0])        # wake VAD: speech prob every 40 ms

Usage: transformers (streaming)

s = model.session()                    # text finalizes every 640 ms (commit_chunk=16 @ 25 Hz)
for chunk in mic_chunks:               # any chunk sizes, 16 kHz float PCM
    s.accept(chunk)                    # cheap: runs the encoder once per completed commit
    print(s.preview())                 # committed text + provisional tail; poll at any rate
print(s.finish())                      # flush the tail

Options and their ranges (in frames at 25 Hz, so 1 frame = 40 ms):

model.session(commit_chunk=32)         # finalize every 1280 ms: fewer encoder calls
model.session(commit_chunk=8)          # finalize every 320 ms: snappier, more compute
model.session(right_ctx=32)            # the accurate tier c256r32's geometry (more latency)

commit_chunk: any value from 1 up. The trade is update granularity against compute, which grows with the window overlap (commit_chunk + lookahead) / commit_chunk. Values over 125 (5 s) trigger a warning suggesting model.transcribe instead.

right_ctx: the tier's r and the latency/accuracy axis. Default 4 (tier stream_c16r4, 920 ms floor latency). The engine derives the frame lookahead from it (right_ctx x superblocks + conv span), so there is nothing to convert by hand. The trunk is trained across a range of right_ctx. The shipped .pte/ONNX graphs are fixed to the tiers in the table, while session() runs any right_ctx in that range. Wider right_ctx buys accuracy at more latency. right_ctx=32 is the accurate c256r32 point.

A tier is two numbers. For low latency use the default c16r4 (commit_chunk=16, right_ctx=4). For accuracy raise right_ctx to 32 (tier c256r32). For the highest throughput on long audio use commit_chunk=512, right_ctx=8 (tier c512r8), or model.transcribe for whole files.

preview() never changes committed text and can be polled at any rate. Each call costs one encoder chunk.

GPU acceleration

This is an edge/CPU model. A GPU is for batched throughput (offline transcription of many files), not lower single-clip latency. The path is bandwidth-bound, so throughput comes from compiling the model (torch.compile(mode="max-autotune")) and batching many files together. The full serving recipe is in the repo. Two optional knobs:

  • FlexAttention: AutoModel.from_pretrained(..., use_flex_attn=True) accelerates the sliding-window streaming path (model.session()). It does not change one-shot model.transcribe(). Default off.
  • Fused Triton kernels: model.use_hub_kernels() routes the LDSA window to futo-org/ldsa (pin version=2) and the causal-SE squeeze to futo-org/causal-trailing-mean. Install kernels in the range your transformers pins (5.x → kernels>=0.12,<0.13). Both fall back to the standard path when not on a CUDA GPU or during training.

Streaming tiers

The model uses sliding-window attention with a cache carried across chunks. It can run at several operating points. A file decode is the same as live dictation at a larger commit and right-context. Tiers are listed in order of their full window (lookahead + commit), which is the single axis along which throughput rises, WER falls and latency rises. In general, the first row is the lowest latency, lowest throughput and highest WER, while the last row is the highest latency, highest throughput, with comparable but slightly higher WER to the 4th tier (c256r32).

tier commit lookahead first result Notes
stream_c16r4 16 fr 920.0 ms 1.56 s Lowest latency (livestream, dictation)
stream_c32r8 32 fr 1720.0 ms 3.00 s
stream_c64r16 64 fr 3320.0 ms 5.88 s
stream_c128r24 128 fr 4920.0 ms 10.04 s
stream_c256r32 256 fr 6520.0 ms 16.76 s Highest accuracy
stream_c512r8 512 fr 1720.0 ms 22.20 s Highest throughput (long audio)

Tier names read c{commit}r{right-context} in 25 Hz frames. right_ctx is the latency axis. commit is the finalization cadence: a larger commit gives the decode fewer block boundaries and a cleaner decoding. The benchmark table above (7.62 average with PCEC) is scored at a single operating point, commit 256 with 64 frames of lookahead, across all eight sets. The .pte/ONNX run a fixed tier instead, using true streaming.

Usage: ONNX

onnx/ holds one graph per component: one per tier (stream_c16r4.onnx, ...), one per PCEC rung and its flush, and the VAD. Every graph reads its weights from one shared file, asr4all.onnx.data, so the backbone is stored once. Load only the components you need, or all of them for instant tier switching: the caches have the same shape for every tier, so a stream can move between tiers mid-utterance. metadata.json maps component names to files.

python onnx/asr4all_onnx.py onnx/ clip.wav                       # default (lowest-latency) tier
python onnx/asr4all_onnx.py onnx/ clip.wav --tier stream_c256r32 # alternate tier (name or index)
python onnx/asr4all_onnx.py onnx/ clip.wav --no-pcec             # raw lowercase decode

Usage: ExecuTorch

executorch/xnnpack_{int8,fp32}/ are XNNPACK backend models for CPU and mobile inference. Each .pte is one multi-method export holding the same components as the ONNX file, and metadata.json carries the vocab, the PCEC control-token ids, the decode contract and each tier's commit / lookahead / window in frames.

A reference decoder for the python runtime is provided:

pip install executorch
python executorch/asr4all_executorch.py executorch/xnnpack_int8 clip.wav             # punctuated (PCEC)
python executorch/asr4all_executorch.py executorch/xnnpack_int8 clip.wav --no-pcec   # plain lowercase decode
python executorch/asr4all_executorch.py executorch/xnnpack_int8 clip.wav --tier stream_c256r32

Training data

About 38,000 hours of English speech across 15 public corpora:

corpus hours
YODAS 17,619
YODAS, colloquial-vocabulary mining 828
People's Speech 5,645
LibriHeavy 4,532
VoxPopuli 4,286
Common Voice v25 2,610
Europarl-ASR 1,065
LibriSpeech 961
AppTek call-center 108
Earnings-22 104
AMI 72
Vystadial 45
SLURP 40
Earnings-21 33
English dialects 31
EdAcc 12

Most corpora receive waveform augmentation (speed perturbation, MUSAN noise, room impulse responses, channel simulation).

Decontamination. Every corpus that overlaps an evaluation set is decontaminated against it before training (8-gram + speaker / video / recording-level, standard Open ASR Leaderboard practice):

corpus decontaminated against
LibriHeavy LibriSpeech test (reader + book-id + 8-gram)
VoxPopuli (unlabeled) VoxPopuli test
YODAS en005 GigaSpeech test (video-level 8-gram) + de-duplicated vs the YODAS-hard subset
Europarl-ASR VoxPopuli test, shared EU-Parliament source. Text-detected overlap dropped at the parent-speech level (audio-clean, not speaker-clean)
Earnings-22 its held-out evaluation calls

Held-out accent evals (EdAcc test, accented-VoxPopuli test) are never trained on. The AppTek call-center data is used in the training data, so it cannot serve as an evaluation benchmark for this model.

Downloads last month
78
Safetensors
Model size
55.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train futo-org/asr4all-m

Collection including futo-org/asr4all-m

Paper for futo-org/asr4all-m

Evaluation results