Instructions to use CobbledSteel/moonshine-tiny-qat with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CobbledSteel/moonshine-tiny-qat with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="CobbledSteel/moonshine-tiny-qat")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("CobbledSteel/moonshine-tiny-qat") model = AutoModelForSpeechSeq2Seq.from_pretrained("CobbledSteel/moonshine-tiny-qat", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Moonshine Tiny β QAT for int8 deployment
Quantisation-aware-trained checkpoint of
UsefulSensors/moonshine-tiny,
prepared for int8 inference on a RISC-V SoC (Rocket + a custom RoCC matrix engine on a
Xilinx Zynq-7020).
The weights in this repository are float32. They are trained so that they quantise cleanly to int8 β they are not stored as int8. Use them as the starting point for int8 quantisation, not as a pre-quantised model.
Why this exists, and why the stock checkpoint will not do
Post-training quantisation of the stock checkpoint does not survive this deployment. Measured on LibriSpeech dev-clean (765 utterances), float simulation, at three activation-calibration policies:
| activation calibration | this checkpoint | stock checkpoint |
|---|---|---|
max |
38.14 % | 120.22 % |
p99.99 (what QAT trained at) |
17.43 % | 96.70 % |
p99.9 (deployed) |
12.33 % | 26.19 % |
At the deployed setting the stock checkpoint has roughly double the word error rate. A pipeline that silently fetches the stock model still runs, still produces plausible transcripts, and is twice as wrong β with nothing in the output to say why. That failure mode is the reason this checkpoint is published rather than kept locally.
Deployment settings that matter
Reproducing the numbers above requires the calibration policy as well as the weights.
- Activation ranges:
p99.9, pooled over the whole calibration set β not the per-sample percentile then max, which gives ranges 1.09Γ (median) to 3.11Γ (worst layer) wider and costs about 5 points. - Calibration set: 64 pinned LibriSpeech dev-clean windows.
- Not
p99.99, even though that is what QAT trained against.p99.9over the deployment calibration windows lands about 5.1 points better. The mismatch is deliberate and measured.
Measured results
| configuration | WER |
|---|---|
| float simulation, dev-clean 765 | 12.33 % |
| generated C (int8), dev-clean 765, on the board | 13.39 % |
| file-fed, 80 utterances / 40 speakers, on the board | 12.78 % |
| through a MEMS microphone, PS-side equalised, same 80 | 19.08 % |
The last row is end-to-end from a microphone on the device β acoustic path included.
Provenance
| base | UsefulSensors/moonshine-tiny @ 390624ed33d594443aa4aa221f5b9f283b545b5a |
| training | 162.66 h quantisation-aware training |
| architecture | MoonshineForConditionalGeneration, 6 encoder + 6 decoder layers, hidden 288 |
| tensors | 160, all F32 |
| tokenizer | byte-identical to the base model's |
Files
model.safetensors 108,389,160 B
config.json 864 B
generation_config.json 189 B
preprocessor_config.json 215 B
tokenizer.json 1,985,530 B
Licence
MIT, following the base model. Check the base model's terms if you redistribute.
- Downloads last month
- 32
Model tree for CobbledSteel/moonshine-tiny-qat
Base model
moonshine-ai/moonshine-tiny