Notere speech models
Model files downloaded by Notere's on-device speech feature. Every published object is immutable: a new build of a file gets a new name, and Notere's manifest pins each file's sha256, so an overwrite would fail every installed client's verification.
parakeet-tdt-0.6b-v3-int8.encoder.q8b32.onnx
The encoder of NVIDIA's parakeet-tdt-0.6b-v3 (CC-BY-4.0), re-quantized block-wise to 8 bits from the fp32 sherpa-onnx export csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3. Notere pairs it with that export's int8 decoder and joiner.
| size | 945,740,634 bytes |
| sha256 | 86752927dd3b5ecfea95b80ae5984589925e7933170ecc4ca4b5d8d53929453b |
| recipe | onnxruntime MatMulNBitsQuantizer, bits=8, block_size=32, symmetric, accuracy_level=4, over every MatMul weight; the fp32 export's metadata_props copied onto the result |
| reproduce | scripts/speech/quantize-parakeet-encoder.py in the Notere repository (pinned package versions and input digests; reproduces this file byte for byte) |
Why not the upstream int8 export: the per-tensor quantize_dynamic encoder returns no tokens or stops emitting mid-window for some inputs, deterministically. This block-wise export decodes every measured window whole and matches fp32 word error rate within a point on English and Polish meeting audio, at 1.45 GB resident versus 2.2 GB for fp32.
Model tree for Wyvek/notere-speech-models
Base model
nvidia/parakeet-tdt-0.6b-v3