Moonshine Tiny β€” QAT for int8 deployment

Quantisation-aware-trained checkpoint of UsefulSensors/moonshine-tiny, prepared for int8 inference on a RISC-V SoC (Rocket + a custom RoCC matrix engine on a Xilinx Zynq-7020).

The weights in this repository are float32. They are trained so that they quantise cleanly to int8 β€” they are not stored as int8. Use them as the starting point for int8 quantisation, not as a pre-quantised model.

Why this exists, and why the stock checkpoint will not do

Post-training quantisation of the stock checkpoint does not survive this deployment. Measured on LibriSpeech dev-clean (765 utterances), float simulation, at three activation-calibration policies:

activation calibration this checkpoint stock checkpoint
max 38.14 % 120.22 %
p99.99 (what QAT trained at) 17.43 % 96.70 %
p99.9 (deployed) 12.33 % 26.19 %

At the deployed setting the stock checkpoint has roughly double the word error rate. A pipeline that silently fetches the stock model still runs, still produces plausible transcripts, and is twice as wrong β€” with nothing in the output to say why. That failure mode is the reason this checkpoint is published rather than kept locally.

Deployment settings that matter

Reproducing the numbers above requires the calibration policy as well as the weights.

  • Activation ranges: p99.9, pooled over the whole calibration set β€” not the per-sample percentile then max, which gives ranges 1.09Γ— (median) to 3.11Γ— (worst layer) wider and costs about 5 points.
  • Calibration set: 64 pinned LibriSpeech dev-clean windows.
  • Not p99.99, even though that is what QAT trained against. p99.9 over the deployment calibration windows lands about 5.1 points better. The mismatch is deliberate and measured.

Measured results

configuration WER
float simulation, dev-clean 765 12.33 %
generated C (int8), dev-clean 765, on the board 13.39 %
file-fed, 80 utterances / 40 speakers, on the board 12.78 %
through a MEMS microphone, PS-side equalised, same 80 19.08 %

The last row is end-to-end from a microphone on the device β€” acoustic path included.

Provenance

base UsefulSensors/moonshine-tiny @ 390624ed33d594443aa4aa221f5b9f283b545b5a
training 162.66 h quantisation-aware training
architecture MoonshineForConditionalGeneration, 6 encoder + 6 decoder layers, hidden 288
tensors 160, all F32
tokenizer byte-identical to the base model's

Files

model.safetensors          108,389,160 B
config.json                        864 B
generation_config.json             189 B
preprocessor_config.json           215 B
tokenizer.json                 1,985,530 B

Licence

MIT, following the base model. Check the base model's terms if you redistribute.

Downloads last month
32
Safetensors
Model size
27.1M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for CobbledSteel/moonshine-tiny-qat

Finetuned
(28)
this model