File size: 3,090 Bytes
ff5f59d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d8edfe6
ff5f59d
 
d8edfe6
 
ff5f59d
 
 
 
d8edfe6
ff5f59d
d8edfe6
 
 
 
 
 
 
ff5f59d
d8edfe6
ff5f59d
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
---
license: apache-2.0
language:
- zh
tags:
- coreml
- apple-neural-engine
- input-method
- zhuyin
- bopomofo
base_model: Luigi/sloth-ime-models
---

# McBopomofoLM models (Core ML, Apple Neural Engine)

Core ML conversions of the SlothE models used by
[McBopomofoLM](https://github.com/workfunction/McBopomofoLM) (branch `lm`), a fork of
McBopomofo that uses the models to improve candidate selection and to suggest corrections for mistyped syllables.
Both models run on the Apple Neural Engine only (`computeUnits = CPU_AND_NE`).

**Revisions:** `main` and tag `v3.1.2` match McBopomofoLM v3.1.x. Tag `v2.1.1` is the runtime for McBopomofoLM v2.1.1, which has no next-token functions.

## Contents

| Path | What |
|---|---|
| `runtime/` | The exact runtime bundle shipped in McBopomofoLM v3.1.2 (`Contents/Resources/SlothE`): compiled `.mlmodelc`, embedding table, vocabularies, the next-token first-char table `dec_next_first.tsv`, and `runtime-manifest.txt`. The app checks the size and SHA-256 of each file against the manifest before it loads that file. |
| `mlpackage/enc25m_multi_pal2_emb.mlpackage` | SlothE-T 25M encoder, 2-bit palettized. The ternary weights make this lossless. Multifunction package with functions L8/L16/L32/L64/L256. The embedding lookup runs in the caller, from `runtime/enc25m_embed_f16.bin`. |
| `mlpackage/dec_mf_tn_fp16.mlpackage` | SlothE decoder `pred_q35_60m` (Qwen3.5 architecture), fp16, one weight file shared by two function families. **t16/t32/t64/t96** (batch 3) return per-token log-probabilities of given texts; they are bit-identical to v2.1.1's decoder. **n16/n32/n64/n96** (batch 1; input `ids` right-padded and `last` = index of the last real token) return the full next-token log-softmax over the 16k vocabulary. Gated DeltaNet runs in fixed 16-token chunks. |
| `scripts/` | Conversion and bundling scripts, research-grade. They import helpers from the author's research workspace, which is not included, so they document the conversion and do not run on their own. `MCBPMF_LM_WORK` sets the workspace root. `build_next.py` builds the n* functions; `build_dec_tn.py` merges them with the t* functions. |

Measured on an M2 (ANE):
- encoder: 100% of ops on the ANE, about 0.8 ms per call;
- decoder t*: 86–95% on the ANE, 2.6–7.9 ms per call depending on length;
- decoder n*: 83–94% on the ANE, 2.3–3.1 ms per call.

Against HF fp32 on CPU, the n* functions agree on the top-1 token for 95.5% of 288 prefixes; every mismatch is a near-tie. The largest log-prob difference over the top-50 tokens is 0.31 nats at the median and 0.73 at P95, from fp16 on the ANE. The first compile of all functions takes about 3 minutes.

## Sources and licenses

- Weights: [Luigi/sloth-ime-models](https://huggingface.co/Luigi/sloth-ime-models), Apache-2.0.
- `char2id.tsv`: converted from `enc/char2id.json` in
  [Luigi/slothing-web](https://huggingface.co/spaces/Luigi/slothing-web), Apache-2.0.
- Variant classes: derived from OpenCC `TWVariants.txt` and `HKVariants.txt`, Apache-2.0.
- Conversion: Apache-2.0 (`LICENSE`). See `NOTICE.txt`.