Instructions to use rajb3/whisper-tiny.en-ternary with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rajb3/whisper-tiny.en-ternary with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="rajb3/whisper-tiny.en-ternary")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("rajb3/whisper-tiny.en-ternary", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Whisper tiny.en with ternary weights {-1, 0, +1}
Every attention and feed-forward projection and the tied token embedding / output projection of openai/whisper-tiny.en are constrained to three values with one FP32 scale per output row, recovered by quantization-aware fine-tuning. The file is 12.2 MB, against 151 MB for the FP32 original.
| Model | LibriSpeech test-clean WER | test-other WER | File |
|---|---|---|---|
| Whisper tiny.en, FP32, zero-shot | 5.66% | 14.54% | 151.0 MB |
| Whisper tiny.en, FP32, fine-tuned with the same recipe (control) | 4.50% | 12.37% | 151.0 MB |
| Same ternary quantizer, no training | 100% | 100% | — |
| This model | 12.12% | 28.42% | 12.2 MB |
WER is corpus-level with the Whisper English text normalizer on both sides, greedy decoding, scored once per arm on the model rebuilt from this exact file.
Code, design document, full results and negative results: github.com/ajbarryiii/wilderness-labs-stt.
Use
Needs torch, transformers, safetensors and, for the command line, soundfile.
hf download rajb3/whisper-tiny.en-ternary --local-dir whisper-ternary
python whisper-ternary/load_ternary.py clip.wav # 16 kHz audio
import sys; sys.path.insert(0, "whisper-ternary")
from load_ternary import load, transcribe
model, processor = load("whisper-ternary")
print(transcribe(model, processor, audio_16khz_float32))
The model is rebuilt in FP32 from the ternary codes, so it reproduces the accuracy above on any hardware. It does not run on packed ternary kernels; that needs activation quantization too, which this model does not have.
What is in the file
export.safetensors: 65 ternary weight matrices (36.4M parameters; the output projection shares the token embedding's matrix), each as 2-bit codes packed four per byte plus an FP32 scale per row, with the quantized layers' biases in FP32. Every other tensor (convolutions, positions, norms, other biases; 1.3M parameters) is FP16.manifest.json: format, layer shapes, tie map, byte accounting, code histogram, SHA-256 of the weights, and training provenance.load_ternary.py: the standalone loader and a small transcription CLI.- Tokenizer and feature-extractor files from the base model, unchanged.
Code distribution: 26.6% -1, 52.4% 0, 21.0% +1.
How it was trained
Starting from the pretrained checkpoint, keep FP32 latent weights and quantize
each row to clamp(round(W / mean|W|), -1, 1) on every forward pass with an
identity straight-through gradient. Ramp linearly from float to fully ternary
weights over the first 25% of training, then hold. Loss is 0.5 token
cross-entropy plus 0.5 KL divergence to the frozen FP32 model. 12,000 steps at
batch 32 on LibriSpeech train-clean-100 + train-clean-360 (464 h), AdamW, peak
learning rate 1e-3 chosen on dev-clean, BF16 autocast, one RTX 5090, about 19
minutes. The same quantizer with cross-entropy alone and no ramp reached only
50% test-clean WER.
Limitations
- English read speech (LibriSpeech) only; not evaluated on conversational, accented, noisy field or domain-specific audio. test-other shows the gap widening on harder speech.
- Fine-tuned on lowercase, unpunctuated transcripts, so output is mostly lowercase with little punctuation.
- Occasionally keeps generating plausible text after the speech ends instead of stopping. Setting those utterances aside lowers WER by 0.5 points on test-clean and 0.7 on test-other (secondary diagnostic; the table above includes them).
- Activations are floating point. No speed or energy benefit is claimed for this checkpoint.
License
Apache-2.0 (see LICENSE), following the base model. This is a modified
version of openai/whisper-tiny.en; see NOTICE for attribution and the list of
changes. Trained on LibriSpeech (CC BY 4.0, Panayotov et al., 2015).
Model tree for rajb3/whisper-tiny.en-ternary
Base model
openai/whisper-tiny.enDataset used to train rajb3/whisper-tiny.en-ternary
Evaluation results
- wer on LibriSpeech test-cleantest set self-reported12.120
- wer on LibriSpeech test-othertest set self-reported28.420