Community-1 diarization models for NCNN + PLDA/VBx

This repository contains the runtime model assets for the NCNN Community-1 diarization pipeline. Together, they provide speaker segmentation, speaker embeddings, and PLDA parameters for VBx clustering. The linked Python project supplies audio input, filter-bank features, masked pooling, clustering, and RTTM output; these files alone are not an executable pipeline.

Models and file formats

Component Files Format and role
Speaker segmentation segmentation.param, segmentation.bin NCNN text graph and binary FP32 weights. Scores speaker activity in audio windows.
Speaker embedding encoder embedding_encoder.param, embedding_encoder.bin NCNN text graph and binary FP32 weights. Produces frame features from 80-bin filter-bank features.
Embedding projection resnet_seg_1_weight.npy, resnet_seg_1_bias.npy NumPy arrays for the learned projection after masked statistics pooling; produce 256-dimensional speaker embeddings.
PLDA and x-vector transform xvec_transform.npz, plda.npz NumPy archives containing the 256-to-128 transform and PLDA parameters used by VBx clustering. These are used directly, without NCNN conversion.

manifest.json records SHA-256 hashes, tensor contracts, source revisions, and conversion tool versions. Its file paths are relative to this directory. LICENSE contains the CC BY 4.0 license text. The manifest links to the reference ONNX sources by URL; those ONNX files are kept in the code repository and are not part of this runtime bundle.

Sources and changes

Asset Source repository Processing in this bundle
Base diarization pipeline pyannote/speaker-diarization-community-1 Provides the original Community-1 model and pipeline design.
Segmentation NCNN pair FredrikKarlssonSpeech/pyannote-speaker-diarization-onnx Converted its FP32 segmentation ONNX graph to NCNN. The conversion fixes four instance normalization layers and their affine weight order.
Embedding encoder NCNN pair and projection arrays welcomyou/pyannote-community-1-onnx-split, derived from altunenes/speaker-diarization-community-1-onnx Converted the split FP32 encoder to NCNN. Copied the projection arrays without retraining or modifying their values.
PLDA and x-vector transform BUT Speech@FIT DiariZen PLDA, also distributed in Community-1 Copied the two archives unchanged. Their SHA-256 hashes match the BUT Speech@FIT files.

The NCNN models were generated with the project conversion script using pnnx in FP32 mode. The fixed-input NCNN outputs were checked against their ONNX sources; see the evaluation notes for numerical and recording-level comparisons.

The linked pipeline's PLDA/VBx equations adapt BUTSpeechFIT/VBx under Apache-2.0, and its filter-bank implementation uses kaldi-native-fbank under Apache-2.0. These code sources are separate from the weight files listed here.

Use

Place these files together in models/ at the root of the pipeline code repository, then follow that repository's installation and CLI instructions. The runtime loads the two NCNN pairs, both projection arrays, and both PLDA archives from that directory. ONNX, PyTorch, and pyannote are not required for normal NCNN inference. A Vulkan-enabled NCNN build is optional for the neural models.

License and attribution

The upstream model assets in this bundle are licensed under CC BY 4.0, with attribution to their respective creators and the conversion/split authors linked above. The BUT Speech@FIT PLDA directory license explicitly names plda.npz and xvec_transform.npz as CC BY 4.0. The upstream Community-1, segmentation ONNX export, and split embedding export model cards also identify CC BY 4.0. Retain attribution, source links, and the description of conversions when redistributing these assets.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for grikdotnet/Pyannote-Community-1-PLDA-VBx

Quantized
(5)
this model