Parley speech models

ONNX models used by Parley, a desktop app for live speaker-aware transcription and translation. Parley downloads one variant on first run and runs it natively with ONNX Runtime (DirectML on any DirectX 12 GPU, CPU fallback).

Variant Size Precision Best for
gpu16/ 1.55 GB fp16 weights and compute, LayerNorms in fp32 GPUs (DirectML), about 1.6 GB VRAM
standard/ 2.99 GB fp32 GPUs (DirectML), about 3 GB VRAM
compact/ 0.79 GB int8 (dynamic, per-channel) CPU-only machines

gpu16 is built from standard by export_fp16.py: inputs and outputs stay fp32, and every LayerNorm runs in fp32 because DirectML's fp16 ReduceMean overflows on the conformer's variance. On the 97.6 s eight-speaker demo it produces the same transcript as standard (796 vs 797 tokens, identical utterances) and 99.9% diarization agreement.

Each variant folder has a manifest.json with the size and SHA-256 of every file.

Contents

diarization/ β€” NVIDIA Nemotron-3-Diarization (streaming Sortformer, up to 8 speakers)

Exported from the πŸ€— Transformers implementation of nvidia/Nemotron-3-Diarization:

  • diar_embed.onnx β€” feature stacking: log-mel (1, T, 128), T % 8 == 0 β†’ embeddings (1, T/8, 512)
  • diar_step.onnx β€” encoder + head: embeddings (1, L, 512) and padding mask (1, L) β†’ speaker probabilities (1, L, 8) pooled to 80 ms frames. Right padding is exact (masked attention, zeroed before the upsampling conv), so callers can use fixed shapes on DirectML.
  • silence_embeds.f32 β€” learned silence embedding for the speaker cache
  • diarization.json β€” Arrival-Order Speaker Cache / FIFO parameters and streaming modes
  • mel_filters_128x257.f32 β€” slaney mel filter bank (shared with the ASR front end)

The streaming state (speaker cache + FIFO, with its score-based compression) is not part of the graphs; Parley implements it in Rust, ported from Nemotron3DiarizationSpeakerCache. Streaming output matches the Transformers reference exactly on the fp32 variant (100% frame agreement on the model card's example); the int8 variant agrees on 99.5% of frames.

asr/ β€” NVIDIA Nemotron 3.5 ASR (cache-aware FastConformer RNN-T, 40 locales, 560 ms chunks)

NeMo cache-aware streaming export of nvidia/nemotron-3.5-asr-streaming-0.6b by the sherpa-onnx project (csukuangfj2/sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms[-int8]-2026-06-11): encoder.onnx (+ encoder.data), decoder.onnx, joiner.onnx, tokens.txt. Chunk geometry, cache shapes and the language-prompt dictionary are in the encoder's metadata.

Front end

Both models consume the same features: 16 kHz mono, pre-emphasis 0.97, 512-point FFT, 400-sample symmetric Hann window, hop 160, power spectrum, 128 slaney mel bins, ln(x + 2^-24), no normalization.

License and attribution

The model weights are NVIDIA's, released under the OpenMDW License Agreement 1.1; these are format conversions and quantizations of those weights, redistributed under the same terms.

  • Nemotron-3-Diarization and Nemotron 3.5 ASR Β© NVIDIA Corporation.
  • ASR ONNX export by the sherpa-onnx project (Apache-2.0 tooling).
  • Diarization export and quantization by Parley.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for killameep/parley-models

Quantized
(27)
this model