mindtype-tagger-v1.1

MindType's own grammar and spelling correction model, trained from scratch by Mindscend for an on-device iOS keyboard. Text never leaves the phone.

  • 26.2M parameters, encoder-only transformer (8 layers, d=384), own 16k byte-level BPE tokenizer.
  • Predicts one edit label per word (KEEP / DELETE / REPLACE_x / APPEND_x / case) in a single pass, so it is fast, tiny, and cannot invent words or rewrite correct text.
  • Core ML package included (8-bit weights) for iOS 18+.

Results (MindType evaluation suite, 4,356 held-out cases)

Set n F0.5 Exact False-positive rate Lost protected spans
casual_clean 118 0.000 0.898 0.102 0.000
casual_noisy 109 0.912 0.780 0.000 0.000
coedit 579 0.434 0.104 0.000 0.000
formal_clean 1000 0.000 0.963 0.037 0.000
formal_noisy 1000 0.882 0.684 0.000 0.000
golden 50 0.927 0.880 0.000 0.000
preserve 1000 0.761 0.662 0.000 0.000
preserve_clean 500 0.000 0.792 0.208 0.000
overall 4356 0.742 0.689 0.094 0.000

With the plausibility guard used in the MindType app (a fix is kept only if it is a known kind of fix or a small spelling difference): F0.5 0.748, exact 68.3%, false positives 8.65%.

Compared with fine-tuned small LLMs (same suite)

Model Size F0.5 โ†‘ Exact โ†‘ False positives โ†“ Lost names/numbers โ†“
Gemma 3 270M instruct, zero-shot 270M 0.122 0.2% 99.6% 3.9%
Gemma 3 270M + LoRA r32 270M 0.571 57.4% 23.1% 2.9%
SmolLM2 360M + LoRA r32 360M 0.663 63.1% 20.1% 0.3%
mindtype-tagger-v1.1 (keep bias 2, guard) 26M 0.748 68.3% 8.65% 0.0%

Files

File What it is
model.safetensors PyTorch weights (fp32): encoder + per-token classification head
config.json architecture (layers, width, heads, vocabulary and label counts)
tokenizer.json 16k byte-level BPE tokenizer (Hugging Face tokenizers format)
labels.json the 15,000 edit labels, index โ†’ label
plausible_pairs.json the word pairs the plausibility guard accepts
coreml/MindTypeTagger.mlpackage Core ML model (8-bit weights, 26 MB), inputs padded to 16/32/64/96/160 tokens
metrics.json, log.jsonl, pretrain-log.jsonl, run.json evaluation results and training logs

Training

  1. Pretraining: masked language modeling on ~20M sentences from FineWeb-Edu (ODC-BY).
  2. Correction: 6M (noisy, clean) pairs: synthetic phone-typing errors, CoEdIT GEC (Apache-2.0), and casual messages written by a local open-weight model (Qwen2.5-14B-Instruct) and filtered for correctness. ~30% are already-correct text, so the model learns to leave good text alone.

Hardware: one RTX 3080 Ti laptop GPU: 1 hour of pretraining, 2.5 hours of correction training.

How it works

Text is split into words and punctuation; each word gets a label from labels.json: $KEEP, $DELETE, $REPLACE_<word>, $APPEND_<word> (insert after), or a case change. The corrected sentence is the labels applied to the words, and the model is run again (up to 3 passes) on its own output. On iPhone (Core ML, CPU) one sentence takes about 4 ms.

Inference settings used for the results: KEEP bias 2.0 (added to the $KEEP logit; higher = fewer, surer fixes). Names, numbers, emails and URLs are never changed.

Limitations

English only, sentence-level, experimental. Fixes what's there; it doesn't reword or change tone. Can still change correct text sometimes (8.65% of correct sentences in our suite).

License

Apache-2.0 (weights, tokenizer, Core ML package). Training data: FineWeb-Edu (ODC-BY), CoEdIT (Apache-2.0).

Downloads last month
-
Safetensors
Model size
26.2M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support