|
Download README.md from FluidInference/phonon-2-coreml: direct link, hf CLI and curl.
- Browser
- Download file 6.64 kB
-
https://huggingface.co/FluidInference/phonon-2-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/phonon-2-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/phonon-2-coreml/resolve/main/README.md
6.64 kB
| license: cc-by-4.0 | |
| base_model: FermionResearch/Phonon-2 | |
| base_model_relation: quantized | |
| language: [en] | |
| pipeline_tag: automatic-speech-recognition | |
| tags: [coreml, parakeet, tdt, phonon, low-bit, sparse, speech-recognition, apple-silicon] | |
| library_name: fluidaudio | |
| # phonon-2-coreml | |
| Core ML build of [FermionResearch/Phonon-2](https://huggingface.co/FermionResearch/Phonon-2), Fermion Research's | |
| quantization-aware re-training of `nvidia/parakeet-tdt-0.6b-v3` for English, in which every encoder weight takes one of | |
| five learned values per output row. Same tokenizer, 15 s window and `Decoder` / `JointDecisionv3` contract as | |
| [parakeet-tdt-0.6b-v3-coreml](https://huggingface.co/FluidInference/parakeet-tdt-0.6b-v3-coreml); loaded in | |
| [FluidAudio](https://github.com/FluidInference/FluidAudio) via `AsrModelVersion.phonon2`. | |
| Every encoder here holds the checkpoint's **exact** five-value weights β as fp16 palettes (iOS 18 / macOS 15 | |
| `constexpr_lut_to_dense`), or as a sparsity mask plus a palette over the non-zero weights (`constexpr_lut_to_sparse` + | |
| `constexpr_sparse_to_dense`). No re-quantization noise is added, and all five produce identical transcripts. The decoder | |
| and joint are re-exported from the checkpoint's int6 tables; preprocessor and vocabulary are the v3 ones. | |
| ## Usage (FluidAudio, merged in [#980](https://github.com/FluidInference/FluidAudio/pull/980)) | |
| ```swift | |
| let models = try await AsrModels.downloadAndLoad(version: .phonon2) // iOS 18+ / macOS 15+, downloads Encoder.mlmodelc | |
| let asr = AsrManager() | |
| try await asr.initialize(models: models) | |
| let result = try await asr.transcribe(audioFileURL) | |
| ``` | |
| ```bash | |
| swift run fluidaudiocli transcribe audio.wav --model-version phonon2 | |
| swift run fluidaudiocli asr-benchmark --subset test-clean --model-version phonon2 | |
| ``` | |
| To use another encoder file, download the model directory, rename the chosen `Encoder_*.mlmodelc` to `Encoder.mlmodelc` | |
| and load it with `AsrModels.loadLocal(from:version: .phonon2, encoderComputeUnits:)`. | |
| ## Encoder files | |
| One 15 s window on an M5 Pro (macOS 27); RTFx is FluidAudio `asr-benchmark` on full LibriSpeech test-clean with the | |
| encoder on the Neural Engine. All five are exact and give the same transcripts. | |
| | File | Size | Encoding | ANE | ANE RTFx | GPU | | |
| |---|---:|---|---:|---:|---| | |
| | **`Encoder.mlmodelc`** (default) | **321 MB** | sparse mask + 6-bit palette per 8 rows | **18.6 ms** | **159Γ** | 16 ms, but ~150 s load on every launch | | |
| | `Encoder_sparse-g4.mlmodelc` | 246 MB | sparse mask + 4-bit palette per 4 rows | 24.3 ms | 140Γ | same load caveat | | |
| | `Encoder_sparse-g1.mlmodelc` | 176 MB | sparse mask + 2-bit palette per row | 70 ms | ~70Γ | same load caveat | | |
| | `Encoder_lut6.mlmodelc` | 470 MB | dense 6-bit palette per 8 rows | 18.6 ms | 155Γ | 16 ms, 0.6 s load | | |
| | `Encoder_lut3.mlmodelc` | 253 MB | dense 3-bit palette per row | 72 ms | 70Γ | 16 ms, 0.7 s load | | |
| | v3 `Encoder.mlmodelc` (6-bit, for reference) | 445 MB | β | 23.5 ms | 149β152Γ | 18 ms | | |
| Two facts decide the choice. The Neural Engine's palette cost grows with the number of palettes, not their bit width, | |
| so 8 rows per palette is faster than v3's own encoder while per-row palettes are 3Γ slower. And the GPU materializes | |
| *sparse* weights at every load (~150 s of CPU, not cached), while the ANE compiles them once (~1 min) and caches. So: | |
| Neural Engine (the FluidAudio default, and the only option for iOS background work) β `Encoder.mlmodelc`; GPU β | |
| `Encoder_lut3.mlmodelc` (small) or `Encoder_lut6.mlmodelc` (fast everywhere); smallest possible β `Encoder_sparse-g1`. | |
| Other files: `Decoder.mlmodelc` (RNNT prediction net, fp16, iOS 17+), `JointDecisionv3.mlmodelc` (single-step joint + | |
| top-64, fp16, iOS 17+), `Preprocessor.mlmodelc` and `parakeet_vocab.json` (identical to v3), `NOTICE` / `LICENSE` | |
| (upstream attribution, CC-BY-4.0). | |
| ## Accuracy vs parakeet-tdt-0.6b-v3 | |
| Full LibriSpeech, FluidAudio `asr-benchmark`, M5 Pro, both models run back to back with the default (ANE) encoder. WER is | |
| corpus-level (total edit distance over total reference words); RTFx = total audio / total processing time. | |
| | Set (ANE) | v3 | Ultra | Phonon-2 | | |
| |---|---|---|---| | |
| | test-clean (2620 files) WER | 2.27 % | **2.13 %** | 2.47 % | | |
| | test-other (2939 files) WER | 4.12 % | **3.81 %** | 4.62 % | | |
| | test-clean RTFx | 149β152Γ | 151Γ | **159Γ** | | |
| | test-other RTFx | 138Γ | 142Γ | **146Γ** | | |
| v3 is the more accurate model on LibriSpeech by 0.20 / 0.50 points, reproducing the upstream card's own gap to its | |
| teacher (+0.20 / +0.79 under the Open ASR Leaderboard protocol, where Phonon-2 in turn beats v3 on AMI meetings and | |
| VoxPopuli). Phonon-2 is English only; v3 covers 25 languages. The absolute values are above the card's because | |
| FluidAudio decodes in 15 s windows with a simpler normalizer, which both models pay equally. | |
| Conversion fidelity: on the first 100 test-clean files the Core ML transcripts differ from a NeMo fp32 full-context | |
| decode of the same checkpoint by 0.34 % WER (1.83 % vs 1.79 %). | |
| ## 60-minute long-form file (Earnings-22, four concatenated calls) | |
| `transcribe` on the 3600 s `earnings22_top4_1h.wav` (M5 Pro, default ANE encoder, best of 2β3 runs, processing time | |
| excludes model load). Reference = the concatenated Earnings-22 chunk transcripts, same normalizer. | |
| | Model | Processing time | RTFx | WER | | |
| |---|---:|---:|---:| | |
| | v3 | 10.9 s | 331Γ | 16.5 % | | |
| | Ultra | 7.7 s | 469Γ | **13.5 %** | | |
| | Redux | 14.2 s | 254Γ | 14.8 % | | |
| | Phonon-2 default (sparse, 321 MB) | 7.5 s | **478Γ** | 17.2 % | | |
| | Phonon-2 `Encoder_lut6` (470 MB) | 7.4 s | 486Γ | 17.2 % | | |
| | Phonon-2 `Encoder_sparse-g4` (246 MB) | 9.0 s | 399Γ | 17.2 % | | |
| | Phonon-2 `Encoder_sparse-g1` (176 MB) | 23.4 s | 154Γ | 17.2 % | | |
| | Phonon-2 `Encoder_lut3` (253 MB) | 23.6 s | 152Γ | 17.2 % | | |
| On conversational long-form audio Phonon-2 is the fastest model we ship (1.45Γ v3's throughput, on par with Ultra) but | |
| the least accurate of the four: Ultra and Redux both beat v3 here while Phonon-2 trails it by 0.8 points, consistent | |
| with the upstream card's Earnings-22 row (6.96 % vs its teacher's 5.85 %). All five Phonon-2 encoders produce the same | |
| transcript. | |
| ## Licence and attribution | |
| The weights are a derivative of NVIDIA's parakeet-tdt-0.6b-v3 re-trained by Fermion Research and are distributed under | |
| CC-BY-4.0, as upstream. `NOTICE` reproduces Fermion Research's change list and training-data attribution. The encoders are | |
| exact re-encodings of the upstream weights (grouped-channel palettes, or a sparsity mask plus palettes over the non-zeros); | |
| the decoder and joint are the upstream int6 tables in fp16. | |