McBopomofoLM v3.1.2 models: decoder with next-token functions (n16-n96), dec_next_first.tsv
d8edfe6 verified |
Download README.md from workfunction/McBopomofoLM-models: direct link, hf CLI and curl.
- Browser
- Download file 3.09 kB
-
https://huggingface.co/workfunction/McBopomofoLM-models/resolve/main/README.md
- Command line
-
hf download hf://workfunction/McBopomofoLM-models/README.md
-
curl -L -o README.md https://huggingface.co/workfunction/McBopomofoLM-models/resolve/main/README.md
3.09 kB
metadata
license: apache-2.0
language:
- zh
tags:
- coreml
- apple-neural-engine
- input-method
- zhuyin
- bopomofo
base_model: Luigi/sloth-ime-models
McBopomofoLM models (Core ML, Apple Neural Engine)
Core ML conversions of the SlothE models used by
McBopomofoLM (branch lm), a fork of
McBopomofo that uses the models to improve candidate selection and to suggest corrections for mistyped syllables.
Both models run on the Apple Neural Engine only (computeUnits = CPU_AND_NE).
Revisions: main and tag v3.1.2 match McBopomofoLM v3.1.x. Tag v2.1.1 is the runtime for McBopomofoLM v2.1.1, which has no next-token functions.
Contents
| Path | What |
|---|---|
runtime/ |
The exact runtime bundle shipped in McBopomofoLM v3.1.2 (Contents/Resources/SlothE): compiled .mlmodelc, embedding table, vocabularies, the next-token first-char table dec_next_first.tsv, and runtime-manifest.txt. The app checks the size and SHA-256 of each file against the manifest before it loads that file. |
mlpackage/enc25m_multi_pal2_emb.mlpackage |
SlothE-T 25M encoder, 2-bit palettized. The ternary weights make this lossless. Multifunction package with functions L8/L16/L32/L64/L256. The embedding lookup runs in the caller, from runtime/enc25m_embed_f16.bin. |
mlpackage/dec_mf_tn_fp16.mlpackage |
SlothE decoder pred_q35_60m (Qwen3.5 architecture), fp16, one weight file shared by two function families. t16/t32/t64/t96 (batch 3) return per-token log-probabilities of given texts; they are bit-identical to v2.1.1's decoder. n16/n32/n64/n96 (batch 1; input ids right-padded and last = index of the last real token) return the full next-token log-softmax over the 16k vocabulary. Gated DeltaNet runs in fixed 16-token chunks. |
scripts/ |
Conversion and bundling scripts, research-grade. They import helpers from the author's research workspace, which is not included, so they document the conversion and do not run on their own. MCBPMF_LM_WORK sets the workspace root. build_next.py builds the n* functions; build_dec_tn.py merges them with the t* functions. |
Measured on an M2 (ANE):
- encoder: 100% of ops on the ANE, about 0.8 ms per call;
- decoder t*: 86–95% on the ANE, 2.6–7.9 ms per call depending on length;
- decoder n*: 83–94% on the ANE, 2.3–3.1 ms per call.
Against HF fp32 on CPU, the n* functions agree on the top-1 token for 95.5% of 288 prefixes; every mismatch is a near-tie. The largest log-prob difference over the top-50 tokens is 0.31 nats at the median and 0.73 at P95, from fp16 on the ANE. The first compile of all functions takes about 3 minutes.
Sources and licenses
- Weights: Luigi/sloth-ime-models, Apache-2.0.
char2id.tsv: converted fromenc/char2id.jsonin Luigi/slothing-web, Apache-2.0.- Variant classes: derived from OpenCC
TWVariants.txtandHKVariants.txt, Apache-2.0. - Conversion: Apache-2.0 (
LICENSE). SeeNOTICE.txt.