MMS-LID 1024 (Core ML, Float16)
Core ML conversion of facebook/mms-lid-1024 for on-device speech language identification on iOS 17+ and macOS. This is the base (float16) variant: best accuracy, no quantization.
- Source: facebook/mms-lid-1024 (Wav2Vec2ForSequenceClassification, ~1B params, 1024 languages)
- Input: Raw 16 kHz mono waveform, fixed 10 seconds (160,000 samples), shape
(1, 160000)float32 - Output: Logits shape
(1, 1024);argmaxgives class index. Map index → ISO 639-3 usinglabels.jsonormms_lid_id2label.json
Contents
| File | Description |
|---|---|
mms_lid.mlpackage |
Core ML model (float16, iOS 17+) |
labels.json |
Ordered list of 1024 ISO 639-3 language codes (index = logits argmax) |
mms_lid_id2label.json |
Index → language code mapping |
Usage on iOS / macOS
- Load
mms_lid.mlpackagewith Core ML (MLModel). - Ensure audio is 16 kHz mono float32. Pad shorter than 10 s with zeros, or trim longer to 10 s (first 160,000 samples).
- Feed input:
input_values= shape(1, 160000). - Get
logitsoutput, takeargmaxalong the last dimension → predicted class index. - Look up the ISO 639-3 code in
labels.json(array index) ormms_lid_id2label.json(key as string).
Recommendations: For long audio, split into ~6 s chunks, run LID per chunk, and use majority vote. Apply a confidence threshold (e.g. softmax max < 0.7 → treat as "unknown") to reduce false positives.
Limitations
- Fixed length: Model expects exactly 10 s of audio; pad or trim accordingly.
- L2 accent: Non-native-accented speech is often misclassified as the speaker's L1 (e.g. Japanese-accented English → Japanese).
- English ↔ Hawaiian/Maori: English is sometimes misclassified as Hawaiian (haw) or Maori (mri); use chunking and/or confidence threshold to mitigate.
Mac smoke test (Core ML)
On-device smoke run: each file under INPUT/audio was resampled to 16 kHz mono float32, padded or trimmed to 160,000 samples (10 s), then passed to input_values; pred is ISO 639-3 from argmax(logits); conf is softmax mass on the predicted class (runner-side).
Note: Filenames are hints only (e.g. English.mp3 is not ground truth). Low conf or known MMS-LID confusions (e.g. English vs haw) may still appear.
Raw runner log
MMS-LID 1024 Core ML — Mac smoke test
Model: https://huggingface.co/aoiandroid/mms-lid-1024-coreml
Model dir: $PROJECT_ROOT/Log/mms_lid_1024_coreml_mac_test/model_repo
Audio dir: $PROJECT_ROOT/INPUT/audio
Compiled temp: /var/folders/ky/nmbswxzs0s79wdxndfw1y6wh0000gn/T/model_repo.mlmodelc
Compute: MLComputeUnits(rawValue: 2)
Input: input_values Output: logits
Labels: 1024
Host: ams-macbook-air.local macOS: Version 26.3.1 (a) (Build 25D771280a)
English.mp3 pcm_samples=9054841 pred=haw conf=0.2392 max_logit=7.5547 time_ms=1227.0
Euskara.mp3 pcm_samples=1865769 pred=hin conf=0.4018 max_logit=7.9141 time_ms=451.6
Guaraní.mp3 pcm_samples=1682285 pred=grn conf=0.9993 max_logit=14.8125 time_ms=427.8
Yorùbá.mp3 pcm_samples=1067049 pred=haw conf=0.4383 max_logit=7.9219 time_ms=404.7
afrikaasns.mp3 pcm_samples=2387800 pred=nld conf=0.9994 max_logit=15.0156 time_ms=483.1
arabic.mp3 pcm_samples=2060120 pred=ara conf=0.9989 max_logit=14.3047 time_ms=470.2
bengali.m4a pcm_samples=7836432 pred=ben conf=0.9986 max_logit=14.5312 time_ms=616.0
chinese.mp3 pcm_samples=12904245 pred=cmn conf=0.9993 max_logit=14.3516 time_ms=1400.2
isiZulu.mp3 pcm_samples=1396819 pred=heb conf=0.5570 max_logit=8.1641 time_ms=460.3
kiswahili.mp3 pcm_samples=1888757 pred=swh conf=0.9989 max_logit=14.2734 time_ms=503.7
korean.mp3 pcm_samples=2364395 pred=kor conf=0.9995 max_logit=15.1953 time_ms=475.1
russinan.m4a pcm_samples=15431029 pred=rus conf=0.2830 max_logit=7.8672 time_ms=926.3
test.mp3 pcm_samples=274560 pred=jpn conf=0.9984 max_logit=14.5234 time_ms=361.3
日本語.mp3 pcm_samples=1798234 pred=jpn conf=0.9984 max_logit=14.5547 time_ms=527.0
License
CC-BY-NC-4.0 (inherited from facebook/mms-lid-1024).
Citation
@article{pratap2023mms,
title={Scaling Speech Technology to 1,000+ Languages},
author={Pratap, Vineel and others},
journal={arXiv preprint arXiv:2305.13516},
year={2023}
}
- Downloads last month
- 3