sp4ev-phoneme-ctc
A Core ML conversion of facebook/wav2vec2-lv-60-espeak-cv-ft, a multilingual wav2vec2 model fine-tuned on Common Voice to recognise phonetic labels. It is used on device by the Sp4Ev Lab app (iPad and iPhone) to place phone boundaries for a known transcription by CTC forced alignment.
What changed from the original
- Converted to a Core ML ML Program (iOS 17+) with coremltools 9.0.
- Fixed input: 10 s of mono 16 kHz audio (160,000 samples), normalised to zero mean and unit variance, zero-padded when shorter. Longer audio is processed in 10 s windows overlapping by 2 s.
- Output
logp: log-probabilities, shape (1, 499, 392), one frame per 20 ms. The 392 labels are those of the original model'svocab.json. - Weights palettized to 8 bits (k-means) for distribution, 303 MB. No retraining or fine-tuning; the weights are the original model's.
Files
PhonemeCTC.mlmodelc/— the compiled Core ML model (compute units: CPU and GPU; the Neural Engine compiler did not finish on this model in testing)vocab.json— the 392 output labels, unchanged from the original modelmanifest.json— size and SHA-256 of every file; the app checks each download against the hashes it ships with
Accuracy
Phone boundaries placed mid-way through the blank gap between phones, compared with hand-corrected boundaries on speakers not used for tuning (phone sequence taken from the reference labels):
| Language | Corpus | Boundaries | Median error | Within 20 ms | Within 50 ms |
|---|---|---|---|---|---|
| English | Buckeye | 29,231 | 14.0 ms | 64% | 91% |
| Japanese | CSJ (core) | 22,052 | 10.9 ms | 72% | 92% |
| Korean | Seoul Corpus | 25,428 | 12.8 ms | 67% | 91% |
The compressed Core ML model gives the same boundaries as the PyTorch original on the Japanese test set.
Licence
Apache License 2.0, as for the original model. Copyright of the original model belongs to its authors (Meta AI). Please cite the original work:
Xu, Q., Baevski, A., & Auli, M. (2021). Simple and Effective Zero-shot Cross-lingual Phoneme Recognition. arXiv:2109.11680.
Model tree for labphonlab/sp4ev-phoneme-ctc
Base model
facebook/wav2vec2-lv-60-espeak-cv-ft