DiariZen v2 for CoreML

CoreML conversions of the models behind DiariZen's wavlm-large-s80-md-v2 speaker-diarization pipeline, packaged for Apple silicon. Used by the transcriber desktop app through FluidAudio's offline diarization pipeline. Every conversion is an exact algebraic rewrite of the source checkpoint (reshapes, framing, normalization identities, activation identities); nothing was retrained, fine-tuned or further pruned.

file source checkpoint interface placement
DiariZenSegmentation.mlmodelc BUT-FIT/diarizen-wavlm-large-s80-md-v2 (pruned WavLM-Large + Conformer, 4 local speakers, 16-class powerset) waveform (1, 256000) float32, 16 kHz mono โ†’ logprobs (1, 799, 16) Neural Engine (1253 of 1255 ops)
DiariZenFBank.mlmodelc Kaldi fbank + per-window mean normalization as used by pyannote's WeSpeaker wrapper audio (B, 1, 256000), B in 1โ€ฆ32 โ†’ fbank_features (B, 1, 80, 1598) CPU
DiariZenEmbedding.mlmodelc pyannote/wespeaker-voxceleb-resnet34-LM fbank_features (1, 1, 80, 1598) + weights (4, 799) mask lanes โ†’ embedding (4, 256) Neural Engine
PldaRho.mlmodelc, plda-parameters.json, xvector-transform.json DiariZen's PLDA (plda.npz, xvec_transform.npz), byte-identical to pyannote community-1's 256 โ†’ 128 CPU

The embedding model runs the ResNet trunk once per 16 s window and pools up to four speaker masks from it, which is what DiariZen's per-slot embedding computes, at a quarter of the calls.

Verified against the PyTorch pipeline on ICSI meetings: segmentation log-probabilities within fp16 rounding (argmax agreement โ‰ฅ 99.75 %), embeddings cosine โ‰ฅ 0.99997, and identical per-embedding cluster labels through PLDA + AHC + VBx; meeting-level DER equal to the reference.

Requirements: macOS 14 or later, Apple silicon. First load compiles the segmentation model for the Neural Engine (10โ€“25 s, cached afterwards).

Licenses

  • The segmentation model is a derivative of DiariZen's weights, released under CC BY-NC 4.0 by BUT-FIT. This repository is therefore CC BY-NC 4.0: attribution required, non-commercial use only.
  • The fbank and embedding models derive from pyannote/wespeaker-voxceleb-resnet34-LM (CC BY 4.0), itself a port of WeSpeaker's VoxCeleb ResNet34-LM.
  • PLDA parameters are DiariZen's, distributed with the checkpoint above.

Please cite DiariZen if you use these models:

Jiangyu Han et al., "Leveraging Self-Supervised Learning for Speaker Diarization" and "Fine-tuning Self-Supervised Models for Speaker Diarization with Structured Pruning" (BUT-FIT, 2025).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sl-data-systems/diarizen-v2-coreml

Finetuned
(4)
this model