Download README.md from FluidInference/phonon-2-coreml: direct link, hf CLI and curl.
- Browser
- Download file 6.64 kB
-
https://huggingface.co/FluidInference/phonon-2-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/phonon-2-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/phonon-2-coreml/resolve/main/README.md
license: cc-by-4.0
base_model: FermionResearch/Phonon-2
base_model_relation: quantized
language:
- en
pipeline_tag: automatic-speech-recognition
tags:
- coreml
- parakeet
- tdt
- phonon
- low-bit
- sparse
- speech-recognition
- apple-silicon
library_name: fluidaudio
phonon-2-coreml
Core ML build of FermionResearch/Phonon-2, Fermion Research's
quantization-aware re-training of nvidia/parakeet-tdt-0.6b-v3 for English, in which every encoder weight takes one of
five learned values per output row. Same tokenizer, 15 s window and Decoder / JointDecisionv3 contract as
parakeet-tdt-0.6b-v3-coreml; loaded in
FluidAudio via AsrModelVersion.phonon2.
Every encoder here holds the checkpoint's exact five-value weights — as fp16 palettes (iOS 18 / macOS 15
constexpr_lut_to_dense), or as a sparsity mask plus a palette over the non-zero weights (constexpr_lut_to_sparse +
constexpr_sparse_to_dense). No re-quantization noise is added, and all five produce identical transcripts. The decoder
and joint are re-exported from the checkpoint's int6 tables; preprocessor and vocabulary are the v3 ones.
Usage (FluidAudio, merged in #980)
let models = try await AsrModels.downloadAndLoad(version: .phonon2) // iOS 18+ / macOS 15+, downloads Encoder.mlmodelc
let asr = AsrManager()
try await asr.initialize(models: models)
let result = try await asr.transcribe(audioFileURL)
swift run fluidaudiocli transcribe audio.wav --model-version phonon2
swift run fluidaudiocli asr-benchmark --subset test-clean --model-version phonon2
To use another encoder file, download the model directory, rename the chosen Encoder_*.mlmodelc to Encoder.mlmodelc
and load it with AsrModels.loadLocal(from:version: .phonon2, encoderComputeUnits:).
Encoder files
One 15 s window on an M5 Pro (macOS 27); RTFx is FluidAudio asr-benchmark on full LibriSpeech test-clean with the
encoder on the Neural Engine. All five are exact and give the same transcripts.
| File | Size | Encoding | ANE | ANE RTFx | GPU |
|---|---|---|---|---|---|
Encoder.mlmodelc (default) |
321 MB | sparse mask + 6-bit palette per 8 rows | 18.6 ms | 159× | 16 ms, but ~150 s load on every launch |
Encoder_sparse-g4.mlmodelc |
246 MB | sparse mask + 4-bit palette per 4 rows | 24.3 ms | 140× | same load caveat |
Encoder_sparse-g1.mlmodelc |
176 MB | sparse mask + 2-bit palette per row | 70 ms | ~70× | same load caveat |
Encoder_lut6.mlmodelc |
470 MB | dense 6-bit palette per 8 rows | 18.6 ms | 155× | 16 ms, 0.6 s load |
Encoder_lut3.mlmodelc |
253 MB | dense 3-bit palette per row | 72 ms | 70× | 16 ms, 0.7 s load |
v3 Encoder.mlmodelc (6-bit, for reference) |
445 MB | — | 23.5 ms | 149–152× | 18 ms |
Two facts decide the choice. The Neural Engine's palette cost grows with the number of palettes, not their bit width,
so 8 rows per palette is faster than v3's own encoder while per-row palettes are 3× slower. And the GPU materializes
sparse weights at every load (150 s of CPU, not cached), while the ANE compiles them once (1 min) and caches. So:
Neural Engine (the FluidAudio default, and the only option for iOS background work) → Encoder.mlmodelc; GPU →
Encoder_lut3.mlmodelc (small) or Encoder_lut6.mlmodelc (fast everywhere); smallest possible → Encoder_sparse-g1.
Other files: Decoder.mlmodelc (RNNT prediction net, fp16, iOS 17+), JointDecisionv3.mlmodelc (single-step joint +
top-64, fp16, iOS 17+), Preprocessor.mlmodelc and parakeet_vocab.json (identical to v3), NOTICE / LICENSE
(upstream attribution, CC-BY-4.0).
Accuracy vs parakeet-tdt-0.6b-v3
Full LibriSpeech, FluidAudio asr-benchmark, M5 Pro, both models run back to back with the default (ANE) encoder. WER is
corpus-level (total edit distance over total reference words); RTFx = total audio / total processing time.
| Set (ANE) | v3 | Ultra | Phonon-2 |
|---|---|---|---|
| test-clean (2620 files) WER | 2.27 % | 2.13 % | 2.47 % |
| test-other (2939 files) WER | 4.12 % | 3.81 % | 4.62 % |
| test-clean RTFx | 149–152× | 151× | 159× |
| test-other RTFx | 138× | 142× | 146× |
v3 is the more accurate model on LibriSpeech by 0.20 / 0.50 points, reproducing the upstream card's own gap to its teacher (+0.20 / +0.79 under the Open ASR Leaderboard protocol, where Phonon-2 in turn beats v3 on AMI meetings and VoxPopuli). Phonon-2 is English only; v3 covers 25 languages. The absolute values are above the card's because FluidAudio decodes in 15 s windows with a simpler normalizer, which both models pay equally.
Conversion fidelity: on the first 100 test-clean files the Core ML transcripts differ from a NeMo fp32 full-context decode of the same checkpoint by 0.34 % WER (1.83 % vs 1.79 %).
60-minute long-form file (Earnings-22, four concatenated calls)
transcribe on the 3600 s earnings22_top4_1h.wav (M5 Pro, default ANE encoder, best of 2–3 runs, processing time
excludes model load). Reference = the concatenated Earnings-22 chunk transcripts, same normalizer.
| Model | Processing time | RTFx | WER |
|---|---|---|---|
| v3 | 10.9 s | 331× | 16.5 % |
| Ultra | 7.7 s | 469× | 13.5 % |
| Redux | 14.2 s | 254× | 14.8 % |
| Phonon-2 default (sparse, 321 MB) | 7.5 s | 478× | 17.2 % |
Phonon-2 Encoder_lut6 (470 MB) |
7.4 s | 486× | 17.2 % |
Phonon-2 Encoder_sparse-g4 (246 MB) |
9.0 s | 399× | 17.2 % |
Phonon-2 Encoder_sparse-g1 (176 MB) |
23.4 s | 154× | 17.2 % |
Phonon-2 Encoder_lut3 (253 MB) |
23.6 s | 152× | 17.2 % |
On conversational long-form audio Phonon-2 is the fastest model we ship (1.45× v3's throughput, on par with Ultra) but the least accurate of the four: Ultra and Redux both beat v3 here while Phonon-2 trails it by 0.8 points, consistent with the upstream card's Earnings-22 row (6.96 % vs its teacher's 5.85 %). All five Phonon-2 encoders produce the same transcript.
Licence and attribution
The weights are a derivative of NVIDIA's parakeet-tdt-0.6b-v3 re-trained by Fermion Research and are distributed under
CC-BY-4.0, as upstream. NOTICE reproduces Fermion Research's change list and training-data attribution. The encoders are
exact re-encodings of the upstream weights (grouped-channel palettes, or a sparsity mask plus palettes over the non-zeros);
the decoder and joint are the upstream int6 tables in fp16.