phonon-2-coreml
Core ML build of FermionResearch/Phonon-2, Fermion Research's
quantization-aware re-training of nvidia/parakeet-tdt-0.6b-v3 for English, in which every encoder weight takes one of
five learned values per output row. Same tokenizer, 15 s window and Decoder / JointDecisionv3 contract as
parakeet-tdt-0.6b-v3-coreml; loaded in
FluidAudio via AsrModelVersion.phonon2.
Every encoder here holds the checkpoint's exact five-value weights β as fp16 palettes (iOS 18 / macOS 15
constexpr_lut_to_dense), or as a sparsity mask plus a palette over the non-zero weights (constexpr_lut_to_sparse +
constexpr_sparse_to_dense). No re-quantization noise is added, and all five produce identical transcripts. The decoder
and joint are re-exported from the checkpoint's int6 tables; preprocessor and vocabulary are the v3 ones.
Usage (FluidAudio, merged in #980)
let models = try await AsrModels.downloadAndLoad(version: .phonon2) // iOS 18+ / macOS 15+, downloads Encoder.mlmodelc
let asr = AsrManager()
try await asr.initialize(models: models)
let result = try await asr.transcribe(audioFileURL)
swift run fluidaudiocli transcribe audio.wav --model-version phonon2
swift run fluidaudiocli asr-benchmark --subset test-clean --model-version phonon2
To use another encoder file, download the model directory, rename the chosen Encoder_*.mlmodelc to Encoder.mlmodelc
and load it with AsrModels.loadLocal(from:version: .phonon2, encoderComputeUnits:).
Encoder files
One 15 s window on an M5 Pro (macOS 27); RTFx is FluidAudio asr-benchmark on full LibriSpeech test-clean with the
encoder on the Neural Engine. All five are exact and give the same transcripts.
| File | Size | Encoding | ANE | ANE RTFx | GPU |
|---|---|---|---|---|---|
Encoder.mlmodelc (default) |
321 MB | sparse mask + 6-bit palette per 8 rows | 18.6 ms | 159Γ | 16 ms, but ~150 s load on every launch |
Encoder_sparse-g4.mlmodelc |
246 MB | sparse mask + 4-bit palette per 4 rows | 24.3 ms | 140Γ | same load caveat |
Encoder_sparse-g1.mlmodelc |
176 MB | sparse mask + 2-bit palette per row | 70 ms | ~70Γ | same load caveat |
Encoder_lut6.mlmodelc |
470 MB | dense 6-bit palette per 8 rows | 18.6 ms | 155Γ | 16 ms, 0.6 s load |
Encoder_lut3.mlmodelc |
253 MB | dense 3-bit palette per row | 72 ms | 70Γ | 16 ms, 0.7 s load |
v3 Encoder.mlmodelc (6-bit, for reference) |
445 MB | β | 23.5 ms | 149β152Γ | 18 ms |
Two facts decide the choice. The Neural Engine's palette cost grows with the number of palettes, not their bit width,
so 8 rows per palette is faster than v3's own encoder while per-row palettes are 3Γ slower. And the GPU materializes
sparse weights at every load (150 s of CPU, not cached), while the ANE compiles them once (1 min) and caches. So:
Neural Engine (the FluidAudio default, and the only option for iOS background work) β Encoder.mlmodelc; GPU β
Encoder_lut3.mlmodelc (small) or Encoder_lut6.mlmodelc (fast everywhere); smallest possible β Encoder_sparse-g1.
Other files: Decoder.mlmodelc (RNNT prediction net, fp16, iOS 17+), JointDecisionv3.mlmodelc (single-step joint +
top-64, fp16, iOS 17+), Preprocessor.mlmodelc and parakeet_vocab.json (identical to v3), NOTICE / LICENSE
(upstream attribution, CC-BY-4.0).
Accuracy vs parakeet-tdt-0.6b-v3
Full LibriSpeech, FluidAudio asr-benchmark, M5 Pro, both models run back to back with the default (ANE) encoder. WER is
corpus-level (total edit distance over total reference words); RTFx = total audio / total processing time.
| Set (ANE) | v3 | Ultra | Phonon-2 |
|---|---|---|---|
| test-clean (2620 files) WER | 2.27 % | 2.13 % | 2.47 % |
| test-other (2939 files) WER | 4.12 % | 3.81 % | 4.62 % |
| test-clean RTFx | 149β152Γ | 151Γ | 159Γ |
| test-other RTFx | 138Γ | 142Γ | 146Γ |
v3 is the more accurate model on LibriSpeech by 0.20 / 0.50 points, reproducing the upstream card's own gap to its teacher (+0.20 / +0.79 under the Open ASR Leaderboard protocol, where Phonon-2 in turn beats v3 on AMI meetings and VoxPopuli). Phonon-2 is English only; v3 covers 25 languages. The absolute values are above the card's because FluidAudio decodes in 15 s windows with a simpler normalizer, which both models pay equally.
Conversion fidelity: on the first 100 test-clean files the Core ML transcripts differ from a NeMo fp32 full-context decode of the same checkpoint by 0.34 % WER (1.83 % vs 1.79 %).
60-minute long-form file (Earnings-22, four concatenated calls)
transcribe on the 3600 s earnings22_top4_1h.wav (M5 Pro, default ANE encoder, best of 2β3 runs, processing time
excludes model load). Reference = the concatenated Earnings-22 chunk transcripts, same normalizer.
| Model | Processing time | RTFx | WER |
|---|---|---|---|
| v3 | 10.9 s | 331Γ | 16.5 % |
| Ultra | 7.7 s | 469Γ | 13.5 % |
| Redux | 14.2 s | 254Γ | 14.8 % |
| Phonon-2 default (sparse, 321 MB) | 7.5 s | 478Γ | 17.2 % |
Phonon-2 Encoder_lut6 (470 MB) |
7.4 s | 486Γ | 17.2 % |
Phonon-2 Encoder_sparse-g4 (246 MB) |
9.0 s | 399Γ | 17.2 % |
Phonon-2 Encoder_sparse-g1 (176 MB) |
23.4 s | 154Γ | 17.2 % |
Phonon-2 Encoder_lut3 (253 MB) |
23.6 s | 152Γ | 17.2 % |
On conversational long-form audio Phonon-2 is the fastest model we ship (1.45Γ v3's throughput, on par with Ultra) but the least accurate of the four: Ultra and Redux both beat v3 here while Phonon-2 trails it by 0.8 points, consistent with the upstream card's Earnings-22 row (6.96 % vs its teacher's 5.85 %). All five Phonon-2 encoders produce the same transcript.
Licence and attribution
The weights are a derivative of NVIDIA's parakeet-tdt-0.6b-v3 re-trained by Fermion Research and are distributed under
CC-BY-4.0, as upstream. NOTICE reproduces Fermion Research's change list and training-data attribution. The encoders are
exact re-encodings of the upstream weights (grouped-channel palettes, or a sparsity mask plus palettes over the non-zeros);
the decoder and joint are the upstream int6 tables in fp16.
- Downloads last month
- -