FUTO asr4all+ large

Asr4all "plus" models include additional modules for diarization, speaker verification and speaker segmentation. The ASR decoders and PCEC are otherwise the same as futo-org/asr4all-l. See the non-plus model cards for ASR training and benchmarking details.

These are the additional features provided by the plus models:

  • speaker-verification head (sv) providing voice embeddings
  • powerset speaker-segmentation head (seg) indicating who is talking in each frame
  • FSMN voice-activity detector, larger and more accurate than the base model's wake VAD
  • diarize.py, a self-contained cascaded diarization example pipeline

96.8M parameters in the transcription path, plus the heads. Commercial use is permitted within ethical-use limitations.

Contents

Voice activity detection

The plus models package the FSMN VAD from Alibaba's FunASR (funasr/fsmn-vad, Apache-2.0), 429k parameters. It supplies one speech probability per 10 ms frame with about 800 ms of causal context. This is less suitable for wake detection than the 2.5k-parameter VAD in the non-plus models, and more accurate at bounding speech.

from transformers import AutoModel
import soundfile as sf, torch

m = AutoModel.from_pretrained("futo-org/asr4all-l-plus", trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32")          # 16 kHz mono
p = m.detect_speech_fsmn(torch.from_numpy(audio)[None])      # [1, T] speech probability at 100 Hz
speech = p[0] > 0.5

detect_speech_fsmn is a separate method from the base model's detect_speech because the two return different rates and tensor shapes. ExecuTorch has method vad, taking 1 s of PCM plus the model's state and returning 100 probabilities plus the updated state. The state is six tensors that the caller keeps between calls and zeroes to start a new stream; each is returned with the shape it came in, so a runtime can discover the pairing from the program. Score i of a call describes the 10 ms frame i - 4 of that call, on the same frame grid as the whole-file call. In ONNX it is vad.onnx, input vad_pcm and output vad_speech_prob, whole file in.

Our provided diarization pipeline does not use the VAD. It is included for other applications and for alternative diarization processes that bound speech first.

Speaker embedding

A 512-dimensional speaker embedding trained with Matryoshka representation learning. Truncations at 48, 96, 192 dimensions are usable embeddings in their own right (emb[:, :d], re-normalized).

VoxCeleb1-O EER
raw cosine, whole utterance 1.62 %
with adaptive score normalization against a cohort 1.25 %

The whole-utterance figure is what embed_speaker on a full clip produces. Embeddings get better with more audio (on this head, 3.15 % EER from a single 3 s crop against 1.62 % on the whole utterance), so embed the longest stretch of one voice you have for best quality.

from transformers import AutoModel
import soundfile as sf, torch

m = AutoModel.from_pretrained("futo-org/asr4all-l-plus", trust_remote_code=True).eval()
a, sr = sf.read("alice.wav", dtype="float32")
b, sr = sf.read("bob.wav", dtype="float32")
ea = m.embed_speaker(torch.from_numpy(a)[None])          # [1, 512], unit-norm
eb = m.embed_speaker(torch.from_numpy(b)[None])
same_person = float(ea @ eb.T) > 0.5                     # cosine, tune the threshold on your data
e96 = torch.nn.functional.normalize(ea[:, :96], dim=-1)  # a valid 96-d embedding

ExecuTorch has methods sv_1500ms, sv_3000ms, sv_6000ms, the weights traced at 1.5 / 3 / 6 s of PCM. Input lengths are fixed, so use the longest the audio fills and do not zero-pad up. In ONNX it is sv.onnx, input sv_pcm and output sv_embedding.

The embedding head's adapters sit on a side branch and the trunk runs unmodified, so embedding can share a forward pass with transcription in a runtime built for it.

Speaker segmentation

Powerset speaker segmentation over an 8 s window, returning 11 classes per 40 ms frame (25 Hz). Concurrent speech is a class, not a threshold. Speaker assignment labels apply only within the 8 s window and are not consistent across windows. There is no inherent limit on the number of speakers in a full recording, but within one window the head distinguishes up to 4.

class id label active slots meaning
0 silence [] silence, no local speaker active
1 speaker_0 [0] speaker slot 0 talking alone
2 speaker_1 [1] speaker slot 1 talking alone
3 speaker_2 [2] speaker slot 2 talking alone
4 speaker_3 [3] speaker slot 3 talking alone
5 speaker_0+speaker_1 [0, 1] slots 0 and 1 talking concurrently
6 speaker_0+speaker_2 [0, 2] slots 0 and 2 talking concurrently
7 speaker_0+speaker_3 [0, 3] slots 0 and 3 talking concurrently
8 speaker_1+speaker_2 [1, 2] slots 1 and 2 talking concurrently
9 speaker_1+speaker_3 [1, 3] slots 1 and 3 talking concurrently
10 speaker_2+speaker_3 [2, 3] slots 2 and 3 talking concurrently

Slots are local to the window. Slot 0 in one window and slot 0 in the next may be different people. Connecting slots to speakers across a recording is done by the clustering step in the diarization pipeline.

Segmentation-only benchmarks (DER of the segmentation head only, collar 0, overlap scored, 8 s clips).

dataset segmentation DER miss false alarm confusion
AMI (headset mix) 15.0 % 8.3 % 4.1 % 2.6 %
AMI (single distant mic) 18.1 % 10.3 % 4.7 % 3.1 %
VoxConverse 7.8 % 4.0 % 2.7 % 1.1 %
AISHELL-4 10.7 % 4.1 % 4.8 % 1.7 %
AliMeeting (far-field) 15.2 % 8.3 % 3.2 % 3.8 %
AISHELL-5 (in-car) 24.2 % 11.8 % 8.6 % 3.9 %
m = AutoModel.from_pretrained("futo-org/asr4all-l-plus", trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32")
window = torch.from_numpy(audio[: 16000 * 8])[None]      # one 8 s window
logits = m.segment_speakers(window)                          # [1, 11, 200]
cls = logits.argmax(1)[0]                                    # class id per 40 ms frame
labels = m.config.seg["id2label"]                            # the table above
who = [labels[str(int(c))] for c in cls]                     # e.g. ["speaker_0", "speaker_0+speaker_1", ...]

ExecuTorch has method seg_8000ms, 128000 samples in and [1, 11, 200] logits out. In ONNX it is seg.onnx, input seg_pcm and output seg_logits.

The segmentation head's adapters are summed into the trunk, so segmentation is a second encoder pass and cannot share one with transcription.

Diarization

Diarization composes the components above. NVIDIA NeMo calls this a cascaded pipeline, as opposed to an end-to-end model that emits speaker labels in one step.

cascaded diarization pipeline

The pipeline is diarize.py in this repository, one file depending on numpy, scipy, scikit-learn and torch. It is the code the benchmark numbers below were measured with.

import sys, soundfile as sf, torch
from huggingface_hub import snapshot_download
from transformers import AutoModel

repo = snapshot_download("futo-org/asr4all-l-plus")
sys.path.append(repo)  # append: the repo's executorch/ folder must not shadow the package
from diarize import Diarizer

m = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
audio, sr = sf.read("meeting.wav", dtype="float32")          # 16 kHz mono
out = Diarizer(m)(audio)
for start, end, speaker in out.segments:
    print(f"{start:8.2f} {end:8.2f}  speaker {speaker}")

Benchmarks (DER of this implementation, collar 0, overlap scored).

dataset DER
AMI (headset mix) 15.8 %
AMI (single distant mic) 19.7 %
VoxConverse 9.9 %
AISHELL-4 11.4 %
AliMeeting (far-field) 17.3 %
AISHELL-5 (in-car) 29.4 %

Exported artifacts

format file(s) target
transformers model.safetensors + modeling_asr4all.py Python / research
ONNX onnx/ (one graph per component, shared weights) general purpose (ORT, OpenVINO)
ExecuTorch XNNPACK int8 executorch/xnnpack_int8/ Mobile (ARM) / lightweight (GPTQ-style rounding)
ExecuTorch XNNPACK fp32 executorch/xnnpack_fp32/ Desktop, Laptop
component transformers ONNX (onnx/) ExecuTorch (asr_encoder.pte)
VAD detect_speech_fsmn vad_pcm → vad_speech_prob method vad
speaker embedding embed_speaker sv_pcm → sv_embedding sv_1500ms / sv_3000ms / sv_6000ms
speaker segmentation segment_speakers seg_pcm → seg_logits method seg_8000ms
diarization diarize.py n/a n/a

Training data for the speaker heads

The ASR training data is documented on the base model. The speaker-verification head was trained on VoxCeleb2 plus the VoxCeleb1 development split, and the segmentation head on meeting and conversational corpora with speaker turn labels (AMI, AliMeeting, AISHELL-4, VoxConverse, AVA-AVD, about 520 h), both with waveform augmentation. The FSMN VAD is third-party (Apache-2.0, funasr/fsmn-vad), trained on Mandarin. The diarization and segmentation benchmarks above are on the test splits of corpora whose train splits were used.

Credit for the speaker datasets. Training used VoxCeleb2 plus the VoxCeleb1 development split, and the verification numbers above are on VoxCeleb1-O (none of its 40 test speakers is in the VoxCeleb1 development split):

A. Nagrani, J. S. Chung, A. Zisserman, VoxCeleb: a large-scale speaker identification dataset, INTERSPEECH 2017.

J. S. Chung, A. Nagrani, A. Zisserman, VoxCeleb2: Deep Speaker Recognition, INTERSPEECH 2018.

Both datasets are from the Visual Geometry Group, University of Oxford.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train futo-org/asr4all-l-plus

Collection including futo-org/asr4all-l-plus

Evaluation results