diarize-embedding-eres2netv2-int8.onnx

Statically int8-quantized ERes2NetV2 speaker embedder (3D-Speaker, zh-cn 16k-common), for speaker diarization under sherpa-onnx.

  • 28 MB (fp32 source: 71 MB), 192-dim embeddings, 16 kHz.
  • Drop-in for SpeakerEmbeddingExtractorConfig: the sherpa metadata_props (framework, sample_rate, output_dim, feature_normalize_type) are preserved, which the extractor requires.

Why quantize

fp32 ERes2NetV2 separates close voices well, but its 2D convolutions cost about 10x CAM++'s embedding time β€” roughly 11 minutes on a 19-minute meeting, untenable on the no-GPU laptops this app targets. Static quantization removes that objection:

build time per 6 s window (ORT CPU, 4 threads)
fp32 223 ms
int8 static (this file) 77 ms β€” 2.9x faster
int8 dynamic 528 ms β€” 2.4x slower

Dynamic quantization is a trap here: it lowers Conv to ConvInteger, which onnxruntime's CPU provider does not optimize. Static quantization lowers to QLinearConv, which it does.

Accuracy

Against the fp32 model on real meeting windows: embedding cosine β‰₯ 0.9956 (mean 0.9973), and the pairwise-similarity matrix β€” what clustering actually consumes β€” drifts by at most 0.019.

End-to-end on diarization bench (DER against hand-annotated references):

fixture CAM++ this model
2-speaker interview, 28 min DER 12.6%, 2 voices DER 12.7%, 2 voices
2-speaker phone call, 8 min DER 14.6%, 2 voices DER 14.5%, 2 voices
multi-speaker meeting, 19 min 2 voices, 80/20 speech split 3 voices, 42/37/21

Two-speaker recordings cannot tell these models apart. The difference appears where it matters β€” a meeting with several voices, where CAM++ collapses 80% of the speech onto one speaker.

Both models still undercount a crowded room (3 of 5 real speakers on that meeting), so lets the user pin the speaker count rather than trust auto-detection.

How it was made

quantize_static with QuantFormat.QOperator, per-channel int8 weights, uint8 activations, Conv only, calibrated on ~40 log-mel fbank windows (600 frames β‰ˆ 6 s, per-window global-mean normalized, matching sherpa's own preprocessing) taken from a real meeting recording. Model metadata is copied back from the fp32 file afterwards, since the quantizer drops it.

The script lives in the app repo (scripts/quantize-eres2netv2.py):

python3 scripts/quantize-eres2netv2.py \
    eres2netv2-fp32.onnx diarize-embedding-eres2netv2-int8.onnx \
    some-real-meeting.mp3

Verifying this file

sha256  be6b162137d8b08854268a97763c007e49882f221e02950242923d40d2be157e

Credits and license

The weights derive from iic/speech_eres2netv2_sv_zh-cn_16k-common by the 3D-Speaker team (Apache-2.0); the fp32 ONNX export came via csukuangfj/speaker-embedding-models. This repository redistributes a quantized derivative under the same Apache-2.0 terms. If you use it, cite the original work:

@inproceedings{eres2netv2,
  title     = {{ERes2NetV2}: Boosting Short-Duration Speaker Verification
               Performance with Computational Efficiency},
  author    = {Chen, Yafeng and Zheng, Siqi and Wang, Hui and Cheng, Luyao and
               Zhu, Tinglong and Huang, Rongjie and Qian, Chong and Chen, Qian
               and Zhang, Wen and Wang, Yanmin},
  booktitle = {Interspeech},
  year      = {2024}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for phoenix124/kept-models