voicehospital-denoise
ONNX exports of speech-enhancement models used by voicehospital, an offline desktop app that transcribes doctor–patient consultations. The app downloads these files on first run and verifies each against a pinned SHA-256.
Each file keeps its upstream licence: MossFormer2 is Apache-2.0; the pyannote WeSpeaker files are CC-BY-4.0 (attribution below).
mossformer2_se48k.onnx
The network of MossFormer2_SE_48K from Alibaba's
ClearVoice
(Apache-2.0), exported unchanged with torch.onnx.export (opset 17).
- Input
fbank[1, 496, 180]: 60-bin Kaldi log-mel filterbank of one 4 s window at 48 kHz (40 ms frames, 8 ms hop, Hamming, dither 1.0 on int16-scaled audio), with deltas and delta-deltas. - Output
mask[1, 496, 961]: a real mask for the 1920-point Hamming STFT (hop 384) of the same window.
Windowing (4 s windows every 3 s, 0.5 s discarded at inner edges), features,
STFT and resynthesis follow ClearVoice's decode_one_audio_mossformer2_se_48k
and are done by the application. Exported output matches PyTorch to a maximum
relative error of 1.5e-5.
SHA-256: aa0a90d954b033d83886c48a32166e395d49152f2b453751caa39af60d00f337
Size: 228,228,370 bytes
All credit for the model belongs to its authors; this repository only repackages it for ONNX Runtime.
pyannote_wespeaker_trunk.onnx and pyannote_wespeaker_pool.onnx
The speaker-embedding model of pyannote/speaker-diarization-3.1,
pyannote/wespeaker-voxceleb-resnet34-LM
(WeSpeaker ResNet34-LM trained on VoxCeleb, CC-BY-4.0, by the WeSpeaker
and pyannote authors), split in two graphs where pyannote.audio itself splits
it, and exported with torch.onnx.export (opset 17). Weights unchanged.
pyannote_wespeaker_trunk.onnx:forward_frames. Inputfbank[B, 998, 80](80-bin Kaldi fbank of a 10 s chunk at 16 kHz, 25/10 ms, Hamming, dither 0, on int16-scaled audio, minus its mean over time); outputframes[B, 256, 10, 125].pyannote_wespeaker_pool.onnx:forward_embedding. Inputsframesandweights[B, 589](a local speaker's activity over the segmentation frames of the chunk); outputembedding[B, 256].
Running the trunk once per chunk and pooling per speaker reproduces pyannote 3.1's embeddings (cosine 1.000000 on every speaker of the test clip).
SHA-256 / size:
- trunk
02ac280e129961f05df467affb8d33b47a9281832628a9e590f12d0731c79125, 21,292,833 bytes - pool
3dbc64b3183560047bfd3f391b2a75b24163a2d10bc6f878d6e4c47728a6ff95, 5,252,349 bytes
Export script: tools/diarize/export_pyannote_wespeaker.py in the
voicehospital repository.
Model tree for tdson/voicehospital-denoise
Base model
alibabasglab/MossFormer2_SE_48K