Instructions to use futo-org/asr4all-l-plus with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use futo-org/asr4all-l-plus with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="futo-org/asr4all-l-plus", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("futo-org/asr4all-l-plus", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
FUTO asr4all+ large
Asr4all "plus" models include additional modules for diarization, speaker verification and
speaker segmentation. The ASR decoders and PCEC are otherwise the same as
futo-org/asr4all-l. See the non-plus model cards for ASR training and
benchmarking details.
These are the additional features provided by the plus models:
- speaker-verification head (
sv) providing voice embeddings - powerset speaker-segmentation head (
seg) indicating who is talking in each frame - FSMN voice-activity detector, larger and more accurate than the base model's wake VAD
diarize.py, a self-contained cascaded diarization example pipeline
96.8M parameters in the transcription path, plus the heads. Commercial use is permitted within ethical-use limitations.
Contents
- Voice activity detection
- Speaker embedding
- Speaker segmentation
- Diarization
- Exported artifacts
- Training data for the speaker heads
Voice activity detection
The plus models package the FSMN VAD from Alibaba's FunASR
(funasr/fsmn-vad, Apache-2.0), 429k
parameters. It supplies one speech probability per 10 ms frame with about
800 ms of causal context. This is less suitable for wake detection than the
2.5k-parameter VAD in the non-plus models, and more accurate at bounding speech.
from transformers import AutoModel
import soundfile as sf, torch
m = AutoModel.from_pretrained("futo-org/asr4all-l-plus", trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32") # 16 kHz mono
p = m.detect_speech_fsmn(torch.from_numpy(audio)[None]) # [1, T] speech probability at 100 Hz
speech = p[0] > 0.5
detect_speech_fsmn is a separate method from the base model's detect_speech because the two
return different rates and tensor shapes. ExecuTorch has method vad, taking 1 s of PCM plus the
model's state and returning 100 probabilities plus the updated state. The state is six tensors that
the caller keeps between calls and zeroes to start a new stream; each is returned with the shape it
came in, so a runtime can discover the pairing from the program. Score i of a call describes the
10 ms frame i - 4 of that call, on the same frame grid as the whole-file call. In ONNX it is vad.onnx, input vad_pcm and output vad_speech_prob, whole file in.
Our provided diarization pipeline does not use the VAD. It is included for other applications and for alternative diarization processes that bound speech first.
Speaker embedding
A 512-dimensional speaker embedding trained with Matryoshka representation learning.
Truncations at 48, 96, 192 dimensions are usable embeddings in their own right
(emb[:, :d], re-normalized).
| VoxCeleb1-O | EER |
|---|---|
| raw cosine, whole utterance | 1.62 % |
| with adaptive score normalization against a cohort | 1.25 % |
The whole-utterance figure is what embed_speaker on a full clip produces. Embeddings get better with more audio (on this head, 3.15 % EER from a single 3 s crop against 1.62 % on the whole utterance), so embed the longest stretch of one voice you have for best quality.
from transformers import AutoModel
import soundfile as sf, torch
m = AutoModel.from_pretrained("futo-org/asr4all-l-plus", trust_remote_code=True).eval()
a, sr = sf.read("alice.wav", dtype="float32")
b, sr = sf.read("bob.wav", dtype="float32")
ea = m.embed_speaker(torch.from_numpy(a)[None]) # [1, 512], unit-norm
eb = m.embed_speaker(torch.from_numpy(b)[None])
same_person = float(ea @ eb.T) > 0.5 # cosine, tune the threshold on your data
e96 = torch.nn.functional.normalize(ea[:, :96], dim=-1) # a valid 96-d embedding
ExecuTorch has methods sv_1500ms, sv_3000ms, sv_6000ms, the weights traced at 1.5 / 3 / 6 s of PCM. Input lengths
are fixed, so use the longest the audio fills and do not zero-pad up. In ONNX it is sv.onnx, input sv_pcm and
output sv_embedding.
The embedding head's adapters sit on a side branch and the trunk runs unmodified, so embedding can share a forward pass with transcription in a runtime built for it.
Speaker segmentation
Powerset speaker segmentation over an 8 s window, returning 11 classes per 40 ms frame (25 Hz). Concurrent speech is a class, not a threshold. Speaker assignment labels apply only within the 8 s window and are not consistent across windows. There is no inherent limit on the number of speakers in a full recording, but within one window the head distinguishes up to 4.
| class id | label | active slots | meaning |
|---|---|---|---|
| 0 | silence |
[] |
silence, no local speaker active |
| 1 | speaker_0 |
[0] |
speaker slot 0 talking alone |
| 2 | speaker_1 |
[1] |
speaker slot 1 talking alone |
| 3 | speaker_2 |
[2] |
speaker slot 2 talking alone |
| 4 | speaker_3 |
[3] |
speaker slot 3 talking alone |
| 5 | speaker_0+speaker_1 |
[0, 1] |
slots 0 and 1 talking concurrently |
| 6 | speaker_0+speaker_2 |
[0, 2] |
slots 0 and 2 talking concurrently |
| 7 | speaker_0+speaker_3 |
[0, 3] |
slots 0 and 3 talking concurrently |
| 8 | speaker_1+speaker_2 |
[1, 2] |
slots 1 and 2 talking concurrently |
| 9 | speaker_1+speaker_3 |
[1, 3] |
slots 1 and 3 talking concurrently |
| 10 | speaker_2+speaker_3 |
[2, 3] |
slots 2 and 3 talking concurrently |
Slots are local to the window. Slot 0 in one window and slot 0 in the next may be different people. Connecting slots to speakers across a recording is done by the clustering step in the diarization pipeline.
Segmentation-only benchmarks (DER of the segmentation head only, collar 0, overlap scored, 8 s clips).
| dataset | segmentation DER | miss | false alarm | confusion |
|---|---|---|---|---|
| AMI (headset mix) | 15.0 % | 8.3 % | 4.1 % | 2.6 % |
| AMI (single distant mic) | 18.1 % | 10.3 % | 4.7 % | 3.1 % |
| VoxConverse | 7.8 % | 4.0 % | 2.7 % | 1.1 % |
| AISHELL-4 | 10.7 % | 4.1 % | 4.8 % | 1.7 % |
| AliMeeting (far-field) | 15.2 % | 8.3 % | 3.2 % | 3.8 % |
| AISHELL-5 (in-car) | 24.2 % | 11.8 % | 8.6 % | 3.9 % |
m = AutoModel.from_pretrained("futo-org/asr4all-l-plus", trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32")
window = torch.from_numpy(audio[: 16000 * 8])[None] # one 8 s window
logits = m.segment_speakers(window) # [1, 11, 200]
cls = logits.argmax(1)[0] # class id per 40 ms frame
labels = m.config.seg["id2label"] # the table above
who = [labels[str(int(c))] for c in cls] # e.g. ["speaker_0", "speaker_0+speaker_1", ...]
ExecuTorch has method seg_8000ms, 128000 samples in and [1, 11, 200] logits out.
In ONNX it is seg.onnx, input seg_pcm and output seg_logits.
The segmentation head's adapters are summed into the trunk, so segmentation is a second encoder pass and cannot share one with transcription.
Diarization
Diarization composes the components above. NVIDIA NeMo calls this a cascaded pipeline, as opposed to an end-to-end model that emits speaker labels in one step.
The pipeline is diarize.py in this repository, one file depending on numpy, scipy, scikit-learn and
torch. It is the code the benchmark numbers below were measured with.
import sys, soundfile as sf, torch
from huggingface_hub import snapshot_download
from transformers import AutoModel
repo = snapshot_download("futo-org/asr4all-l-plus")
sys.path.append(repo) # append: the repo's executorch/ folder must not shadow the package
from diarize import Diarizer
m = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32") # 16 kHz mono
out = Diarizer(m)(audio)
for start, end, speaker in out.segments:
print(f"{start:8.2f} {end:8.2f} speaker {speaker}")
Benchmarks (DER of this implementation, collar 0, overlap scored).
| dataset | DER |
|---|---|
| AMI (headset mix) | 15.8 % |
| AMI (single distant mic) | 19.7 % |
| VoxConverse | 9.9 % |
| AISHELL-4 | 11.4 % |
| AliMeeting (far-field) | 17.3 % |
| AISHELL-5 (in-car) | 29.4 % |
Exported artifacts
| format | file(s) | target |
|---|---|---|
| transformers | model.safetensors + modeling_asr4all.py |
Python / research |
| ONNX | onnx/ (one graph per component, shared weights) |
general purpose (ORT, OpenVINO) |
| ExecuTorch XNNPACK int8 | executorch/xnnpack_int8/ |
Mobile (ARM) / lightweight (GPTQ-style rounding) |
| ExecuTorch XNNPACK fp32 | executorch/xnnpack_fp32/ |
Desktop, Laptop |
| component | transformers | ONNX (onnx/) |
ExecuTorch (asr_encoder.pte) |
|---|---|---|---|
| VAD | detect_speech_fsmn |
vad_pcm → vad_speech_prob |
method vad |
| speaker embedding | embed_speaker |
sv_pcm → sv_embedding |
sv_1500ms / sv_3000ms / sv_6000ms |
| speaker segmentation | segment_speakers |
seg_pcm → seg_logits |
method seg_8000ms |
| diarization | diarize.py |
n/a | n/a |
Training data for the speaker heads
The ASR training data is documented on the base model. The speaker-verification head was trained
on VoxCeleb2 plus the VoxCeleb1 development split, and the segmentation head on meeting and conversational corpora with speaker turn labels
(AMI, AliMeeting, AISHELL-4, VoxConverse, AVA-AVD, about 520 h), both with waveform augmentation.
The FSMN VAD is third-party (Apache-2.0, funasr/fsmn-vad), trained on Mandarin. The diarization
and segmentation benchmarks above are on the test splits of corpora whose train splits were used.
Credit for the speaker datasets. Training used VoxCeleb2 plus the VoxCeleb1 development split, and the verification numbers above are on VoxCeleb1-O (none of its 40 test speakers is in the VoxCeleb1 development split):
A. Nagrani, J. S. Chung, A. Zisserman, VoxCeleb: a large-scale speaker identification dataset, INTERSPEECH 2017.
J. S. Chung, A. Nagrani, A. Zisserman, VoxCeleb2: Deep Speaker Recognition, INTERSPEECH 2018.
Both datasets are from the Visual Geometry Group, University of Oxford.
- Downloads last month
- -
Datasets used to train futo-org/asr4all-l-plus
openslr/librispeech_asr
MLCommons/peoples_speech
Collection including futo-org/asr4all-l-plus
Evaluation results
- Test WER (Whisper-normalized) on AMI (cleaned)test set self-reported9.940
- Test WER (Whisper-normalized) on GigaSpeech (cleaned)test set self-reported9.840
- Test WER (Whisper-normalized) on VoxPopuli (cleaned)test set self-reported4.050
- Test WER (Whisper-normalized) on Earnings-22 (cleaned, chunked)test set self-reported9.150
- Test WER (Whisper-normalized) on LibriSpeech cleantest set self-reported2.380
- Test WER (Whisper-normalized) on LibriSpeech othertest set self-reported5.940
- Test WER (Whisper-normalized) on SPGISpeechtest set self-reported4.610
- Test WER (Whisper-normalized) on Monsoon en-INtest set self-reported7.400
