Instructions to use futo-org/asr4all-s with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use futo-org/asr4all-s with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="futo-org/asr4all-s", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("futo-org/asr4all-s", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- FUTO asr4all small (streaming)
- Contents
- Benchmark results
- Model details
- Voice activity detection (wake)
- PCEC: punctuation, capitalization and error correction
- Exported artifacts
- Recommended configuration per platform
- Usage: transformers (one-shot)
- Usage: transformers (streaming)
- GPU acceleration
- Streaming tiers
- Usage: ONNX
- Usage: ExecuTorch
- Training data
- Contents
FUTO asr4all small (streaming)
This efficiency-focused model (28.1M params: 22.8M acoustic + 5.3M PCEC) is optimized for edge, mobile and CPU inference that opens automatic speech recognition to all platforms. It is an excellent choice for mobile environments, where low-power consumption is desirable. It comes bundled with a secondary model that applies punctuation, capitalization and light error correction.
These models use a custom, unique architecture that significantly increases the realtime factor (RTFx) compared to other ASR models of comparable size. It is heavily optimized for CPU inference, and delegates operations cleanly to vectorized SIMD and NEON operations.
The model has low and high latency streaming inference modes from the same weights, with runtime-configurable latency (details in the streaming usage section below). A small language model adds punctuation, capitalization and error correction (PCEC) capability. We also include an ultralight VAD for use in applications that need low power voice activity detection.
We export weights for a number of different formats, quantizations and runtimes. This model supports English language only. Commercial use is permitted with limitations for ethical restrictions prohibiting certain activities (eg. mass surveillance). See license for details.
Contents
- Benchmark results
- Model details
- Voice activity detection (wake)
- PCEC: punctuation, capitalization and error correction
- Exported artifacts
- Recommended configuration per platform
- Usage: transformers (one-shot)
- Usage: transformers (streaming)
- GPU acceleration
- Streaming tiers
- Usage: ONNX
- Usage: ExecuTorch
- Training data
Benchmark results
Open ASR Leaderboard dataset results (WER %, c256r64, lower is better):
RTFx (single thread, higher is better, 120 s clip, median of 3), one row per recommended platform configuration:
| platform | backend | stream_c16r4 RTFx |
stream_c512r8 RTFx |
|---|---|---|---|
| Pixel 4 | executorch-xnnpack int8 | 14.6 | 30.0 |
| Pixel 10 | executorch-xnnpack int8 | 32.4 | 85.2 |
| Raspberry Pi 4 | executorch-xnnpack int8 | 2.4 | 4.9 |
| i9-12900K | executorch-xnnpack int8 | 52.4 | 128.2 |
| i9-12900K | executorch-xnnpack fp32 | 33.1 | 83.2 |
| i9-12900K | openvino fp32 | 38.1 | 93.0 |
| Ryzen 7 PRO 8700GE | openvino bf16 | 52.2 | 168.2 |
| i9-12900K | openvino-igpu fp16 | 16.7 | 84.0 |
| RTX 6000 Max-Q (batch 32) | cuda | 51968.2 | 51876.6 |
RTFx is the acoustic decoding. The PCEC adds 6–15 % inference time on x86 at one thread.
Model details
- Architecture: DFSMN conv superblocks + LDSA & sliding-window attention, causal squeeze-excitation (SE). CTC over BPE-256+blank, 25hz mel features.
- Streaming: per-layer left-context caches. Each audio frame is encoded once.
- Distilled from the Granite-Speech teacher
(
futo-org/gravel-ctc-440m, logit + hidden-state KD).
Voice activity detection (wake)
The model also includes an ultralight VAD, which operates independently of the ASR. This can be used for things like wake detection or triggering the ASR. It is based on kiloVAD (arXiv:2607.25870) and pruned to 2,537 parameters. Each call takes 1s of audio and returns 21 predictions at 40 ms intervals, each describing 200 ms of context.
Our implementation follows the kiloVAD architecture as described in the paper. This model is trained on permissively-licensed data and ships under this repository's license.
| benchmark | score |
|---|---|
| AVA-Speech AUC (full 151 clips) | 0.8545 |
| AMI frame error (16 meetings) | 14.30 % |
There are two choices for inference, optimized per architecture. For ARM targets, use vad_dft,
while vad is faster for x86.
| x86-64 (AVX-512) | arm64 (NEON) | |
|---|---|---|
vad (conv-STFT) |
0.26 ms | 1.13 ms |
vad_dft (in-graph DFT) |
0.65 ms | 0.65 ms |
PCEC: punctuation, capitalization and error correction
A 5.3M-parameter second-pass head that runs at finalize() and turns the raw CTC hypothesis
into formatted text. It reads the decoded tokens and acoustic hidden state pooled from
the encoder.
Features:
- Punctuation: periods, commas, question marks
- Capitalization: proper nouns and acronyms from the model. Sentence-initial capitalization
is applied by the decode examples (first word, and after
./?). - Error correction: repairs the trunk's characteristic mistakes: garbled non-words from greedy CTC (dropped letters, fused or split words), wrong-word substitutions that sound like the intended word, and homophone disambiguation from sentence context (their/there, no/know, weather/whether).
PCEC is a token-space model that runs incrementally as tokens are committed. Its output trails the raw transcript by a small, fixed lag (commit + lookahead + settle, in tokens). Words behind that lag are settled and no longer change. The tail is provisional and is updated as more audio arrives. At the end of a clip a flush settles the remaining tail.
Live output is punctuated and cased, not only the final text. Every runtime here ships
it: transformers transcribe(), the ONNX file and the ExecuTorch bundles have it on by
default with a flag to disable. Its measured cost is the note under the RTFx table above.
Effect per dataset for this size (WER %, negative delta = better):
| dataset | plain CTC | with PCEC | delta |
|---|---|---|---|
| AMI (cleaned) | 13.98 | 13.05 | -0.93 |
| GigaSpeech (cleaned) | 13.61 | 12.65 | -0.96 |
| VoxPopuli (cleaned) | 5.69 | 5.44 | -0.25 |
| Earnings-22 (cleaned, chunked) | 14.91 | 13.54 | -1.37 |
| LibriSpeech clean | 3.66 | 3.71 | +0.05 |
| LibriSpeech other | 10.45 | 10.07 | -0.38 |
| SPGISpeech | 6.84 | 6.37 | -0.47 |
| Monsoon en-IN | 14.50 | 12.98 | -1.52 |
| macro average | 10.46 | 9.73 | -0.73 |
Exported artifacts
| format | file(s) | target |
|---|---|---|
| transformers | model.safetensors + modeling_asr4all.py |
Python / research |
| ONNX | onnx/ (one graph per component, shared weights) |
general purpose (ORT, OpenVINO) |
| ExecuTorch XNNPACK int8 | executorch/xnnpack_int8/ |
Mobile (ARM) / lightweight (GPTQ-style rounding) |
| ExecuTorch XNNPACK fp32 | executorch/xnnpack_fp32/ |
Desktop, Laptop |
Recommended configuration per platform
| platform | run this | recommended sizes |
|---|---|---|
| Android / modern ARM incl. Raspberry Pi 5 (v8.2+) | ExecuTorch xnnpack_int8 |
all sizes feasible on most devices |
| Raspberry Pi 4 / ARMv8.0 | ExecuTorch xnnpack_int8 |
use small for low latency streaming |
| x86 CPU (Intel and AMD) | OpenVINO on the onnx/ graphs, fp32 or bf16 (if AVX512-BF16 available) |
all sizes |
| Intel iGPU | OpenVINO | all sizes |
| NVIDIA GPU | transformers, bf16 | all sizes |
Usage: transformers (one-shot)
input_values is raw 16 kHz mono PCM: the mel frontend lives inside the
model, so the paired feature extractor only pads and masks.
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="futo-org/asr4all-s", trust_remote_code=True)
print(asr("clip.wav")["text"])
transcribe takes a 16 kHz mono float tensor:
import torch
from transformers import AutoModel
# option A: soundfile (lighter install: `pip install soundfile`):
import soundfile as sf
audio, sr = sf.read("clip.wav", dtype="float32") # numpy [samples] (mono)
wav = torch.from_numpy(audio)
# option B: torchaudio (needs a backend: torchcodec / ffmpeg):
# import torchaudio; wav = torchaudio.load("clip.wav")[0][0]
model = AutoModel.from_pretrained("futo-org/asr4all-s", trust_remote_code=True)
print(model.transcribe(wav)) # punctuated + cased (PCEC, default)
print(model.transcribe(wav, punctuate=False)) # plain lowercase CTC decode
print(model.detect_speech(wav[None])[0]) # wake VAD: speech prob every 40 ms
Usage: transformers (streaming)
s = model.session() # text finalizes every 640 ms (commit_chunk=16 @ 25 Hz)
for chunk in mic_chunks: # any chunk sizes, 16 kHz float PCM
s.accept(chunk) # cheap: runs the encoder once per completed commit
print(s.preview()) # committed text + provisional tail; poll at any rate
print(s.finish()) # flush the tail
Options and their ranges (in frames at 25 Hz, so 1 frame = 40 ms):
model.session(commit_chunk=32) # finalize every 1280 ms: fewer encoder calls
model.session(commit_chunk=8) # finalize every 320 ms: snappier, more compute
model.session(right_ctx=32) # the accurate tier c256r32's geometry (more latency)
commit_chunk: any value from 1 up. The trade is update granularity against compute,
which grows with the window overlap (commit_chunk + lookahead) / commit_chunk.
Values over 125 (5 s) trigger a warning suggesting model.transcribe instead.
right_ctx: the tier's r and the latency/accuracy axis. Default 4 (tier stream_c16r4,
760 ms floor latency). The engine derives the frame lookahead from it
(right_ctx x superblocks + conv span), so there is nothing to convert by hand. The trunk is
trained across a range of right_ctx. The shipped .pte/ONNX graphs are fixed to the tiers in
the table, while session() runs any right_ctx in that range. Wider right_ctx buys accuracy
at more latency. right_ctx=32 is the accurate c256r32 point.
A tier is two numbers. For low latency use the default
c16r4(commit_chunk=16, right_ctx=4). For accuracy raiseright_ctxto 32 (tierc256r32). For the highest throughput on long audio usecommit_chunk=512, right_ctx=8(tierc512r8), ormodel.transcribefor whole files.
preview() never changes committed text and can be polled at any rate. Each call
costs one encoder chunk.
GPU acceleration
This is an edge/CPU model. A GPU is for batched throughput (offline transcription of many
files), not lower single-clip latency. The path is bandwidth-bound, so throughput comes from
compiling the model (torch.compile(mode="max-autotune")) and batching many files together. The
full serving recipe is in the repo. Two optional knobs:
- FlexAttention:
AutoModel.from_pretrained(..., use_flex_attn=True)accelerates the sliding-window streaming path (model.session()). It does not change one-shotmodel.transcribe(). Default off. - Fused Triton kernels:
model.use_hub_kernels()routes the LDSA window tofuto-org/ldsa(pinversion=2) and the causal-SE squeeze tofuto-org/causal-trailing-mean. Installkernelsin the range your transformers pins (5.x →kernels>=0.12,<0.13). Both fall back to the standard path when not on a CUDA GPU or during training.
Streaming tiers
The model uses sliding-window attention with a cache carried across chunks. It can run at several operating points. A file decode is the same as live dictation at a larger commit and right-context. Tiers are listed in order of their full window (lookahead + commit), which is the single axis along which throughput rises, WER falls and latency rises. In general, the first row is the lowest latency, lowest throughput and highest WER, while the last row is the highest latency, highest throughput, with comparable but slightly higher WER to the 4th tier (c256r32).
| tier | commit | lookahead | first result | Notes |
|---|---|---|---|---|
stream_c16r4 |
16 fr | 760.0 ms | 1.40 s | Lowest latency (livestream, dictation) |
stream_c32r8 |
32 fr | 1400.0 ms | 2.68 s | |
stream_c64r16 |
64 fr | 2680.0 ms | 5.24 s | |
stream_c128r24 |
128 fr | 3960.0 ms | 9.08 s | |
stream_c256r32 |
256 fr | 5240.0 ms | 15.48 s | Highest accuracy |
stream_c512r8 |
512 fr | 1400.0 ms | 21.88 s | Highest throughput (long audio) |
Tier names read c{commit}r{right-context} in 25 Hz frames.
right_ctx is the latency axis. commit is the finalization cadence: a larger commit gives the
decode fewer block boundaries and a cleaner decoding. The benchmark table above (9.73 average with PCEC) is scored at a single operating point, commit 256 with 64 frames of lookahead, across all eight sets. The .pte/ONNX run a fixed tier instead, using true streaming. Scored on the same battery without PCEC: stream_c16r4 11.15, stream_c256r32 10.09, stream_c512r8 10.12 -- a spread of 1.06 WER, with stream_c256r32 the most accurate.
Usage: ONNX
onnx/ holds one graph per component: one per tier (stream_c16r4.onnx, ...), one per PCEC
rung and its flush, and the VAD. Every graph reads its weights from one shared file,
asr4all.onnx.data, so the backbone is stored once. Load only the components you need, or all
of them for instant tier switching: the caches have the same shape for every tier, so a stream
can move between tiers mid-utterance. metadata.json maps component names to files.
python onnx/asr4all_onnx.py onnx/ clip.wav # default (lowest-latency) tier
python onnx/asr4all_onnx.py onnx/ clip.wav --tier stream_c256r32 # alternate tier (name or index)
python onnx/asr4all_onnx.py onnx/ clip.wav --no-pcec # raw lowercase decode
Usage: ExecuTorch
executorch/xnnpack_{int8,fp32}/ are XNNPACK backend models for CPU and mobile inference.
Each .pte is one multi-method export holding the same components as the ONNX file,
and metadata.json carries the vocab, the PCEC control-token ids, the decode contract
and each tier's commit / lookahead / window in frames.
A reference decoder for the python runtime is provided:
pip install executorch
python executorch/asr4all_executorch.py executorch/xnnpack_int8 clip.wav # punctuated (PCEC)
python executorch/asr4all_executorch.py executorch/xnnpack_int8 clip.wav --no-pcec # plain lowercase decode
python executorch/asr4all_executorch.py executorch/xnnpack_int8 clip.wav --tier stream_c256r32
Training data
About 38,000 hours of English speech across 15 public corpora:
| corpus | hours |
|---|---|
| YODAS | 17,619 |
| YODAS, colloquial-vocabulary mining | 828 |
| People's Speech | 5,645 |
| LibriHeavy | 4,532 |
| VoxPopuli | 4,286 |
| Common Voice v25 | 2,610 |
| Europarl-ASR | 1,065 |
| LibriSpeech | 961 |
| AppTek call-center | 108 |
| Earnings-22 | 104 |
| AMI | 72 |
| Vystadial | 45 |
| SLURP | 40 |
| Earnings-21 | 33 |
| English dialects | 31 |
| EdAcc | 12 |
Most corpora receive waveform augmentation (speed perturbation, MUSAN noise, room impulse responses, channel simulation).
Decontamination. Every corpus that overlaps an evaluation set is decontaminated against it before training (8-gram + speaker / video / recording-level, standard Open ASR Leaderboard practice):
| corpus | decontaminated against |
|---|---|
| LibriHeavy | LibriSpeech test (reader + book-id + 8-gram) |
| VoxPopuli (unlabeled) | VoxPopuli test |
| YODAS en005 | GigaSpeech test (video-level 8-gram) + de-duplicated vs the YODAS-hard subset |
| Europarl-ASR | VoxPopuli test, shared EU-Parliament source. Text-detected overlap dropped at the parent-speech level (audio-clean, not speaker-clean) |
| Earnings-22 | its held-out evaluation calls |
Held-out accent evals (EdAcc test, accented-VoxPopuli test) are never trained on. The AppTek call-center data is used in the training data, so it cannot serve as an evaluation benchmark for this model.
- Downloads last month
- 77
Datasets used to train futo-org/asr4all-s
openslr/librispeech_asr
MLCommons/peoples_speech
Collection including futo-org/asr4all-s
Paper for futo-org/asr4all-s
Evaluation results
- Test WER (Whisper-normalized) on AMI (cleaned)test set self-reported13.050
- Test WER (Whisper-normalized) on GigaSpeech (cleaned)test set self-reported12.650
- Test WER (Whisper-normalized) on VoxPopuli (cleaned)test set self-reported5.440
- Test WER (Whisper-normalized) on Earnings-22 (cleaned, chunked)test set self-reported13.540
- Test WER (Whisper-normalized) on LibriSpeech cleantest set self-reported3.710
- Test WER (Whisper-normalized) on LibriSpeech othertest set self-reported10.070
- Test WER (Whisper-normalized) on SPGISpeechtest set self-reported6.370