WeSpeaker ResNet34 as TorchScript, with its PLDA
A format conversion made for FineSub's speaker diarization, so
it can run without pyannote.audio. It is not a new model and was not retrained.
| File | What it is | From | Licence |
|---|---|---|---|
wespeaker-cuda.ts |
WeSpeaker ResNet34 speaker-embedding network (the part after fbank), traced on CUDA | pyannote/wespeaker-voxceleb-resnet34-LM @ 837717dd |
CC BY 4.0 |
wespeaker-cpu.ts |
The same network, traced on CPU | same | CC BY 4.0 |
plda/plda.npz, plda/xvec_transform.npz |
PLDA and x-vector transform, byte-identical copies | pyannote/speaker-diarization-community-1 @ 3533c8cf |
CC BY 4.0 |
Two builds
TorchScript records the device it was traced on, so each file runs only on its own device. Both are here because together they are about 54 MB.
What changed from the originals
- Traced with
torch.jit.trace(torch 2.11.0); not frozen, not quantised, weights unchanged. - The traced modules were checked against the eager original bit for bit, repeated calls at batch sizes 1–3, before writing.
- The PLDA files are copied unchanged.
Export script: tools/diarizen_export/export_diarizen.py in the FineSub repository.
Credit
WeSpeaker: Wang et al., "WeSpeaker: A research and production oriented speaker embedding learning
toolkit", ICASSP 2023; the checkpoint as packaged by pyannote. The PLDA is pyannote's, from
speaker-diarization-community-1 (Bredin, pyannote.audio). Please credit them, not this repository.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for spurtcarl/finesub-speaker-embedding
Base model
pyannote/speaker-diarization-community-1