mindtype-tagger-v1.1
MindType's own grammar and spelling correction model, trained from scratch by Mindscend for an on-device iOS keyboard. Text never leaves the phone.
- 26.2M parameters, encoder-only transformer (8 layers, d=384), own 16k byte-level BPE tokenizer.
- Predicts one edit label per word (KEEP / DELETE / REPLACE_x / APPEND_x / case) in a single pass, so it is fast, tiny, and cannot invent words or rewrite correct text.
- Core ML package included (8-bit weights) for iOS 18+.
Results (MindType evaluation suite, 4,356 held-out cases)
| Set | n | F0.5 | Exact | False-positive rate | Lost protected spans |
|---|---|---|---|---|---|
| casual_clean | 118 | 0.000 | 0.898 | 0.102 | 0.000 |
| casual_noisy | 109 | 0.912 | 0.780 | 0.000 | 0.000 |
| coedit | 579 | 0.434 | 0.104 | 0.000 | 0.000 |
| formal_clean | 1000 | 0.000 | 0.963 | 0.037 | 0.000 |
| formal_noisy | 1000 | 0.882 | 0.684 | 0.000 | 0.000 |
| golden | 50 | 0.927 | 0.880 | 0.000 | 0.000 |
| preserve | 1000 | 0.761 | 0.662 | 0.000 | 0.000 |
| preserve_clean | 500 | 0.000 | 0.792 | 0.208 | 0.000 |
| overall | 4356 | 0.742 | 0.689 | 0.094 | 0.000 |
With the plausibility guard used in the MindType app (a fix is kept only if it is a known kind of fix or a small spelling difference): F0.5 0.748, exact 68.3%, false positives 8.65%.
Compared with fine-tuned small LLMs (same suite)
| Model | Size | F0.5 โ | Exact โ | False positives โ | Lost names/numbers โ |
|---|---|---|---|---|---|
| Gemma 3 270M instruct, zero-shot | 270M | 0.122 | 0.2% | 99.6% | 3.9% |
| Gemma 3 270M + LoRA r32 | 270M | 0.571 | 57.4% | 23.1% | 2.9% |
| SmolLM2 360M + LoRA r32 | 360M | 0.663 | 63.1% | 20.1% | 0.3% |
| mindtype-tagger-v1.1 (keep bias 2, guard) | 26M | 0.748 | 68.3% | 8.65% | 0.0% |
Files
| File | What it is |
|---|---|
model.safetensors |
PyTorch weights (fp32): encoder + per-token classification head |
config.json |
architecture (layers, width, heads, vocabulary and label counts) |
tokenizer.json |
16k byte-level BPE tokenizer (Hugging Face tokenizers format) |
labels.json |
the 15,000 edit labels, index โ label |
plausible_pairs.json |
the word pairs the plausibility guard accepts |
coreml/MindTypeTagger.mlpackage |
Core ML model (8-bit weights, 26 MB), inputs padded to 16/32/64/96/160 tokens |
metrics.json, log.jsonl, pretrain-log.jsonl, run.json |
evaluation results and training logs |
Training
- Pretraining: masked language modeling on ~20M sentences from FineWeb-Edu (ODC-BY).
- Correction: 6M (noisy, clean) pairs: synthetic phone-typing errors, CoEdIT GEC (Apache-2.0), and casual messages written by a local open-weight model (Qwen2.5-14B-Instruct) and filtered for correctness. ~30% are already-correct text, so the model learns to leave good text alone.
Hardware: one RTX 3080 Ti laptop GPU: 1 hour of pretraining, 2.5 hours of correction training.
How it works
Text is split into words and punctuation; each word gets a label from labels.json: $KEEP, $DELETE,
$REPLACE_<word>, $APPEND_<word> (insert after), or a case change. The corrected sentence is the labels applied
to the words, and the model is run again (up to 3 passes) on its own output. On iPhone (Core ML, CPU) one sentence
takes about 4 ms.
Inference settings used for the results: KEEP bias 2.0 (added to the $KEEP logit; higher = fewer, surer fixes).
Names, numbers, emails and URLs are never changed.
Limitations
English only, sentence-level, experimental. Fixes what's there; it doesn't reword or change tone. Can still change correct text sometimes (8.65% of correct sentences in our suite).
License
Apache-2.0 (weights, tokenizer, Core ML package). Training data: FineWeb-Edu (ODC-BY), CoEdIT (Apache-2.0).
- Downloads last month
- -