Instructions to use futo-org/asr4all-m-plus with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use futo-org/asr4all-m-plus with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="futo-org/asr4all-m-plus", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("futo-org/asr4all-m-plus", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
FUTO asr4all+ medium
Asr4all "plus" models include additional modules for diarization, speaker verification and
speaker segmentation. The ASR decoders and PCEC are otherwise the same as
futo-org/asr4all-m. See the non-plus model cards for ASR training and
benchmarking details.
These are the additional features provided by the plus models:
- speaker-verification head (
sv) providing voice embeddings - powerset speaker-segmentation head (
seg) indicating who is talking in each frame - FSMN voice-activity detector, larger and more accurate than the base model's wake VAD
diarize.py, a self-contained cascaded diarization example pipeline
55.6M parameters in the transcription path, plus the heads. Commercial use is permitted within ethical-use limitations.
Contents
- Voice activity detection
- Speaker embedding
- Speaker segmentation
- Diarization
- Exported artifacts
- Training data for the speaker heads
Voice activity detection
The plus models package the FSMN VAD from Alibaba's FunASR
(funasr/fsmn-vad, Apache-2.0), 429k
parameters. It supplies one speech probability per 10 ms frame with about
800 ms of causal context. This is less suitable for wake detection than the
2.5k-parameter VAD in the non-plus models, and more accurate at bounding speech.
from transformers import AutoModel
import soundfile as sf, torch
m = AutoModel.from_pretrained("futo-org/asr4all-m-plus", trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32") # 16 kHz mono
p = m.detect_speech_fsmn(torch.from_numpy(audio)[None]) # [1, T] speech probability at 100 Hz
speech = p[0] > 0.5
detect_speech_fsmn is a separate method from the base model's detect_speech because the two
return different rates and tensor shapes. ExecuTorch has method vad, taking 1 s of PCM plus the
model's state and returning 100 probabilities plus the updated state. The state is six tensors that
the caller keeps between calls and zeroes to start a new stream; each is returned with the shape it
came in, so a runtime can discover the pairing from the program. Score i of a call describes the
10 ms frame i - 4 of that call, on the same frame grid as the whole-file call. In ONNX it is vad.onnx, input vad_pcm and output vad_speech_prob, whole file in.
Our provided diarization pipeline does not use the VAD. It is included for other applications and for alternative diarization processes that bound speech first.
Speaker embedding
A 512-dimensional speaker embedding trained with Matryoshka representation learning.
Truncations at 48, 96, 192 dimensions are usable embeddings in their own right
(emb[:, :d], re-normalized).
| VoxCeleb1-O | EER |
|---|---|
| raw cosine, whole utterance | 1.66 % |
| with adaptive score normalization against a cohort | 1.35 % |
The whole-utterance figure is what embed_speaker on a full clip produces. Embeddings get
better with more audio (about 9 % EER at 1.5 s, 4 % at 3 s, 2 % at 6 s on this head), so embed the
longest stretch of one voice you have for best quality.
from transformers import AutoModel
import soundfile as sf, torch
m = AutoModel.from_pretrained("futo-org/asr4all-m-plus", trust_remote_code=True).eval()
a, sr = sf.read("alice.wav", dtype="float32")
b, sr = sf.read("bob.wav", dtype="float32")
ea = m.embed_speaker(torch.from_numpy(a)[None]) # [1, 512], unit-norm
eb = m.embed_speaker(torch.from_numpy(b)[None])
same_person = float(ea @ eb.T) > 0.5 # cosine, tune the threshold on your data
e96 = torch.nn.functional.normalize(ea[:, :96], dim=-1) # a valid 96-d embedding
ExecuTorch has methods sv_1500ms, sv_3000ms, sv_6000ms, the weights traced at 1.5 / 3 / 6 s of PCM. Input lengths
are fixed, so use the longest the audio fills and do not zero-pad up. In ONNX it is sv.onnx, input sv_pcm and
output sv_embedding.
The embedding head's adapters sit on a side branch and the trunk runs unmodified, so embedding can share a forward pass with transcription in a runtime built for it.
Speaker segmentation
Powerset speaker segmentation over an 8 s window, returning 11 classes per 40 ms frame (25 Hz). Concurrent speech is a class, not a threshold. Speaker assignment labels apply only within the 8 s window and are not consistent across windows. There is no inherent limit on the number of speakers in a full recording, but within one window the head distinguishes up to 4.
| class id | label | active slots | meaning |
|---|---|---|---|
| 0 | silence |
[] |
silence, no local speaker active |
| 1 | speaker_0 |
[0] |
speaker slot 0 talking alone |
| 2 | speaker_1 |
[1] |
speaker slot 1 talking alone |
| 3 | speaker_2 |
[2] |
speaker slot 2 talking alone |
| 4 | speaker_3 |
[3] |
speaker slot 3 talking alone |
| 5 | speaker_0+speaker_1 |
[0, 1] |
slots 0 and 1 talking concurrently |
| 6 | speaker_0+speaker_2 |
[0, 2] |
slots 0 and 2 talking concurrently |
| 7 | speaker_0+speaker_3 |
[0, 3] |
slots 0 and 3 talking concurrently |
| 8 | speaker_1+speaker_2 |
[1, 2] |
slots 1 and 2 talking concurrently |
| 9 | speaker_1+speaker_3 |
[1, 3] |
slots 1 and 3 talking concurrently |
| 10 | speaker_2+speaker_3 |
[2, 3] |
slots 2 and 3 talking concurrently |
Slots are local to the window. Slot 0 in one window and slot 0 in the next may be different people. Connecting slots to speakers across a recording is done by the clustering step in the diarization pipeline.
Segmentation-only benchmarks (DER of the segmentation head only, collar 0, overlap scored, 8 s clips).
| dataset | segmentation DER | miss | false alarm | confusion |
|---|---|---|---|---|
| AMI (headset mix) | 15.0 % | 8.0 % | 4.3 % | 2.7 % |
| AMI (single distant mic) | 18.1 % | 9.8 % | 5.1 % | 3.3 % |
| VoxConverse | 8.0 % | 4.1 % | 2.7 % | 1.1 % |
| AISHELL-4 | 11.2 % | 4.0 % | 5.3 % | 1.8 % |
| AliMeeting (far-field) | 15.5 % | 7.5 % | 4.0 % | 4.1 % |
| AISHELL-5 (in-car) | 25.6 % | 12.3 % | 9.5 % | 3.8 % |
m = AutoModel.from_pretrained("futo-org/asr4all-m-plus", trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32")
window = torch.from_numpy(audio[: 16000 * 8])[None] # one 8 s window
logits = m.segment_speakers(window) # [1, 11, 200]
cls = logits.argmax(1)[0] # class id per 40 ms frame
labels = m.config.seg["id2label"] # the table above
who = [labels[str(int(c))] for c in cls] # e.g. ["speaker_0", "speaker_0+speaker_1", ...]
ExecuTorch has method seg_8000ms, 128000 samples in and [1, 11, 200] logits out.
In ONNX it is seg.onnx, input seg_pcm and output seg_logits.
The segmentation head's adapters are summed into the trunk, so segmentation is a second encoder pass and cannot share one with transcription.
Diarization
Diarization composes the components above. NVIDIA NeMo calls this a cascaded pipeline, as opposed to an end-to-end model that emits speaker labels in one step.
The pipeline is diarize.py in this repository, one file depending on numpy, scipy, scikit-learn and
torch. It is the code the benchmark numbers below were measured with.
import sys, soundfile as sf, torch
from huggingface_hub import snapshot_download
from transformers import AutoModel
repo = snapshot_download("futo-org/asr4all-m-plus")
sys.path.append(repo) # append: the repo's executorch/ folder must not shadow the package
from diarize import Diarizer
m = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32") # 16 kHz mono
out = Diarizer(m)(audio)
for start, end, speaker in out.segments:
print(f"{start:8.2f} {end:8.2f} speaker {speaker}")
Benchmarks (DER of this implementation, collar 0, overlap scored).
| dataset | DER |
|---|---|
| AMI (headset mix) | 16.0 % |
| AMI (single distant mic) | 21.2 % |
| VoxConverse | 10.5 % |
| AISHELL-4 | 12.9 % |
| AliMeeting (far-field) | 17.9 % |
| AISHELL-5 (in-car) | 28.6 % |
Exported artifacts
| format | file(s) | target |
|---|---|---|
| transformers | model.safetensors + modeling_asr4all.py |
Python / research |
| ONNX | onnx/ (one graph per component, shared weights) |
general purpose (ORT, OpenVINO) |
| ExecuTorch XNNPACK int8 | executorch/xnnpack_int8/ |
Mobile (ARM) / lightweight (GPTQ-style rounding) |
| ExecuTorch XNNPACK fp32 | executorch/xnnpack_fp32/ |
Desktop, Laptop |
| component | transformers | ONNX (onnx/) |
ExecuTorch (asr_encoder.pte) |
|---|---|---|---|
| VAD | detect_speech_fsmn |
vad_pcm → vad_speech_prob |
method vad |
| speaker embedding | embed_speaker |
sv_pcm → sv_embedding |
sv_1500ms / sv_3000ms / sv_6000ms |
| speaker segmentation | segment_speakers |
seg_pcm → seg_logits |
method seg_8000ms |
| diarization | diarize.py |
n/a | n/a |
Training data for the speaker heads
The ASR training data is documented on the base model. The speaker-verification head was trained
on VoxCeleb2, and the segmentation head on meeting and conversational corpora with speaker turn labels
(AMI, AliMeeting, AISHELL-4, VoxConverse, AVA-AVD, about 520 h), both with waveform augmentation.
The FSMN VAD is third-party (Apache-2.0, funasr/fsmn-vad), trained on Mandarin. The diarization
and segmentation benchmarks above are on the test splits of corpora whose train splits were used.
- Downloads last month
- 27
Datasets used to train futo-org/asr4all-m-plus
openslr/librispeech_asr
MLCommons/peoples_speech
Collection including futo-org/asr4all-m-plus
Evaluation results
- Test WER (Whisper-normalized) on AMI (cleaned)test set self-reported11.060
- Test WER (Whisper-normalized) on GigaSpeech (cleaned)test set self-reported10.630
- Test WER (Whisper-normalized) on VoxPopuli (cleaned)test set self-reported4.540
- Test WER (Whisper-normalized) on Earnings-22 (cleaned, chunked)test set self-reported11.380
- Test WER (Whisper-normalized) on LibriSpeech cleantest set self-reported2.830
- Test WER (Whisper-normalized) on LibriSpeech othertest set self-reported7.250
- Test WER (Whisper-normalized) on SPGISpeechtest set self-reported5.180
- Test WER (Whisper-normalized) on Monsoon en-INtest set self-reported8.900
