Notere speech models

Model files downloaded by Notere's on-device speech feature. Every published object is immutable: a new build of a file gets a new name, and Notere's manifest pins each file's sha256, so an overwrite would fail every installed client's verification.

parakeet-tdt-0.6b-v3-int8.encoder.q8b32.onnx

The encoder of NVIDIA's parakeet-tdt-0.6b-v3 (CC-BY-4.0), re-quantized block-wise to 8 bits from the fp32 sherpa-onnx export csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3. Notere pairs it with that export's int8 decoder and joiner.

size 945,740,634 bytes
sha256 86752927dd3b5ecfea95b80ae5984589925e7933170ecc4ca4b5d8d53929453b
recipe onnxruntime MatMulNBitsQuantizer, bits=8, block_size=32, symmetric, accuracy_level=4, over every MatMul weight; the fp32 export's metadata_props copied onto the result
reproduce scripts/speech/quantize-parakeet-encoder.py in the Notere repository (pinned package versions and input digests; reproduces this file byte for byte)

Why not the upstream int8 export: the per-tensor quantize_dynamic encoder returns no tokens or stops emitting mid-window for some inputs, deterministically. This block-wise export decodes every measured window whole and matches fp32 word error rate within a point on English and Polish meeting audio, at 1.45 GB resident versus 2.2 GB for fp32.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Wyvek/notere-speech-models

Quantized
(101)
this model