voicehospital-denoise

ONNX exports of speech-enhancement models used by voicehospital, an offline desktop app that transcribes doctor–patient consultations. The app downloads these files on first run and verifies each against a pinned SHA-256.

Each file keeps its upstream licence: MossFormer2 is Apache-2.0; the pyannote WeSpeaker files are CC-BY-4.0 (attribution below).

mossformer2_se48k.onnx

The network of MossFormer2_SE_48K from Alibaba's ClearVoice (Apache-2.0), exported unchanged with torch.onnx.export (opset 17).

  • Input fbank [1, 496, 180]: 60-bin Kaldi log-mel filterbank of one 4 s window at 48 kHz (40 ms frames, 8 ms hop, Hamming, dither 1.0 on int16-scaled audio), with deltas and delta-deltas.
  • Output mask [1, 496, 961]: a real mask for the 1920-point Hamming STFT (hop 384) of the same window.

Windowing (4 s windows every 3 s, 0.5 s discarded at inner edges), features, STFT and resynthesis follow ClearVoice's decode_one_audio_mossformer2_se_48k and are done by the application. Exported output matches PyTorch to a maximum relative error of 1.5e-5.

SHA-256: aa0a90d954b033d83886c48a32166e395d49152f2b453751caa39af60d00f337 Size: 228,228,370 bytes

All credit for the model belongs to its authors; this repository only repackages it for ONNX Runtime.

pyannote_wespeaker_trunk.onnx and pyannote_wespeaker_pool.onnx

The speaker-embedding model of pyannote/speaker-diarization-3.1, pyannote/wespeaker-voxceleb-resnet34-LM (WeSpeaker ResNet34-LM trained on VoxCeleb, CC-BY-4.0, by the WeSpeaker and pyannote authors), split in two graphs where pyannote.audio itself splits it, and exported with torch.onnx.export (opset 17). Weights unchanged.

  • pyannote_wespeaker_trunk.onnx: forward_frames. Input fbank [B, 998, 80] (80-bin Kaldi fbank of a 10 s chunk at 16 kHz, 25/10 ms, Hamming, dither 0, on int16-scaled audio, minus its mean over time); output frames [B, 256, 10, 125].
  • pyannote_wespeaker_pool.onnx: forward_embedding. Inputs frames and weights [B, 589] (a local speaker's activity over the segmentation frames of the chunk); output embedding [B, 256].

Running the trunk once per chunk and pooling per speaker reproduces pyannote 3.1's embeddings (cosine 1.000000 on every speaker of the test clip).

SHA-256 / size:

  • trunk 02ac280e129961f05df467affb8d33b47a9281832628a9e590f12d0731c79125, 21,292,833 bytes
  • pool 3dbc64b3183560047bfd3f391b2a75b24163a2d10bc6f878d6e4c47728a6ff95, 5,252,349 bytes

Export script: tools/diarize/export_pyannote_wespeaker.py in the voicehospital repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tdson/voicehospital-denoise

Quantized
(1)
this model