Instructions to use divergentlabs/masker-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use divergentlabs/masker-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="divergentlabs/masker-mini")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("divergentlabs/masker-mini") model = AutoModelForTokenClassification.from_pretrained("divergentlabs/masker-mini", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Masker Mini (v3)
Masker-mini is the on-device sibling of masker. It is a 6-layer distilled PII detector for 23 European languages that fits in 18 MB (4-bit Core ML) or 22 MB (4-bit ONNX).
It runs fully offline, so no PII leaves the device.
| artifact | format | size | token agreement vs fp32 | strict F1 |
|---|---|---|---|---|
model.safetensors |
PyTorch fp16 | 71 MB | — | 0.975 |
onnx/model_fp16.onnx |
ONNX fp16 (portable) | 71 MB | 100% | 0.975 |
onnx/model_int4.onnx |
ONNX 4-bit (ONNX Runtime) | 22 MB | 99.95% | 0.975 |
coreml/masker_mini_4bit.mlpackage |
Core ML 4-bit (Apple NE) | 18 MB | 99.93% | 0.974 |
Per-artifact columns are measured on a 2,000-document openpii validation sample.
What's new in v3
- Much better on real text. On out-of-distribution person names (real court
judgments and Dutch news), names F1 goes from 0.774 → 0.865. Leak recall
goes from 0.868 → 0.922. It beats
nationaldesignstudio/rampart(0.690) anddesert-ant-labs/redact(0.626) on every metric we measure. - Far fewer false names in Dutch. On CoNLL-2002 NL, the share of other named entities (companies, places) wrongly tagged as persons drops from 16.4% → 7.0%.
- One change caused it: the student now starts from pretrained weights. v2's student started from random weights, so everything it knew about language came from the distillation data. v3 starts from the bottom six layers of Multilingual-MiniLM, moved onto masker's 64K tokenizer. The teacher, the data and the recipe are otherwise identical to v2.
- Drop-in replacement for v2. The architecture,
config.json, label map, tokenizer files (byte-identical), inputs and outputs are all the same. You only swap the weights. v2 is kept under thev2tag.
Model type & training
Masker Mini is a 6-layer BERT-architecture token classifier (hidden size 384, 12 heads, 35.4M parameters), a MiniLM-class encoder. It is trained by knowledge distillation (soft-label KL + hard-label CE, 3 epochs, lr 1e-4) from masker v2.
Recipe:
- Teacher: masker v2 (mDeBERTa-v3, trained on openpii plus name-form augmentation plus 293K sentences of human-annotated real-world NER in 14 languages). It is pruned to the 63,896-piece vocabulary and healed with a 0.25-epoch fine-tune, the same as for v2.
- Student initialisation: layers 0–5 of
microsoft/Multilingual-MiniLM-L12-H384, with position embeddings, type embeddings and embedding LayerNorm. Its word embeddings are transplanted onto the teacher's tokenizer by piece string: 32,209 pieces copy exactly, and the other 31,687 get the mean of the MiniLM sub-pieces that spell them. The student and teacher share one tokenizer, so distillation stays exact per token. - Distillation on the same 1.44M-row mix as v2.
Three deployment artifacts are provided:
- ONNX fp16 (71 MB): portable, runs anywhere via ONNX Runtime (Android / iOS / web / server). It is numerically identical to the PyTorch model (100% token agreement).
- ONNX 4-bit (22 MB): weight-only block quantization (
MatMulNBits+GatherBlockQuantized, block 32, symmetric) of the Linear layers and the embedding table. Needs ONNX Runtime ≥ 1.18 for the 4-bit ops. It is 99.95% token-faithful to fp32. - Core ML, 4-bit palettized (18 MB): weight-only k-means palettization for the Apple Neural Engine (iOS 18 / macOS 15+). It is 99.93% token-faithful, with Δ strict F1 = −0.001.
It emits the same 12 entity types as
masker (48 BIOES labels + O):
GIVEN_NAME, SURNAME, CITY, STREET_NAME, BUILDING_NUMBER, ZIP_CODE,
PHONE, EMAIL, CREDIT_CARD, DATE, AGE, GOVERNMENT_ID. It slots into
the same rules-layer pipeline for structured PII.
Usage
For a plain PyTorch / Transformers quick start, the snippet on the
masker card runs unchanged. Just
point it at divergentlabs/masker-mini.
This repo exists mainly for the two on-device builds:
ONNX Runtime (portable for Android, iOS, web, server)
import onnxruntime as ort, numpy as np
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
repo = "divergentlabs/masker-mini"
tok = AutoTokenizer.from_pretrained(repo)
sess = ort.InferenceSession(hf_hub_download(repo, "onnx/model_fp16.onnx"))
text = "Sanne de Groot woont in Utrecht."
feed = {k: v.astype(np.int64) for k, v in tok(text, return_tensors="np").items()}
logits = sess.run(None, feed)[0] # [1, seq_len, 49] BIOES logits -> argmax + decode
For the smallest portable build, swap in onnx/model_int4.onnx (22 MB, 4-bit).
It has the same inputs and outputs and needs ONNX Runtime ≥ 1.18.
Swift (iOS / macOS): swift-masker is a Swift package that wraps the Core ML build with tokenization, the BIOES→span decode, the structured-PII rules layer, and reversible masking. It bundles its own copy of the model. v3 uses the same tokenizer and interface as v2, so the bundled model can be swapped without code changes.
// .package(url: "https://github.com/divergentlabsxyz/swift-masker.git", from: "0.1.0")
let masker = try await Masker()
let masked = try await masker.mask(userText) // "[EMAIL_1] lives in [CITY_1]"
let final = masker.unmask(reply, with: masked.session)
To use the Core ML model directly, add
coreml/masker_mini_4bit.mlpackage to an Xcode target. Its inputs are
input_ids, attention_mask and token_type_ids (Int32, fixed length 256),
and its output is logits ([1, 256, 49], fp16). Tokenize with this repo's
spm.model / tokenizer.json.
Evaluation
Out-of-distribution: person names on real text
These scores come from real documents that none of the models below trained on. Metrics are person-level (given name + surname merged into one span) and character-exact. Every row was measured with the identical evaluation suite.
- TAB: the Text Anonymization Benchmark, 127 English ECHR court judgments.
- CoNLL-2002 NL: Dutch news, test split. Used for evaluation only.
| model | names F1 (mean) | TAB overlap F1 | TAB leak recall | CoNLL NL overlap F1 | CoNLL NL leak recall | CoNLL NL false-person rate |
|---|---|---|---|---|---|---|
| masker v2, 64K (the teacher) | 0.931 | 0.946 | 0.923 | 0.915 | 0.957 | 2.7% |
| masker-mini v3 | 0.865 | 0.887 | 0.912 | 0.843 | 0.931 | 7.0% |
| masker-mini v2 | 0.774 | 0.833 | 0.840 | 0.716 | 0.897 | 16.4% |
| nationaldesignstudio/rampart | 0.690 | 0.769 | 0.757 | 0.611 | 0.743 | 11.0% |
| desert-ant-labs/redact | 0.626 | 0.498 | 0.361 | 0.755 | 0.794 | 7.3% |
How the competitors were run:
- redact v0.4.0: through its own SDK pipeline (
@desert-ant-labs/redact@3.4.0, min confidence 0.6, plus its rules). - rampart (commit
b1993e4): through its own SDK (@nationaldesignstudio/rampart@0.1.3;createGuard().detect()with its default policy, the spansprotect()redacts). On TAB, 43% of its person spans include a title ("Mr X"), and 14% are paragraph-sized. That puts its TAB strict F1 at 0.04, and the paragraph spans make its overlap and leak figures above flatter it.
masker models are scored on the neural head alone, without the rules layer. Leak recall is the fraction of gold name characters that get redacted. The false-person rate is the share of the set's other named entities that are wrongly tagged as a person.
In-distribution: openpii-1m validation
These are span-level, boundary-exact scores on synthetic openpii text, the same distribution as most of the training data. Treat them as an upper bound. They are from the fp32 model on 8,000 validation documents.
Overall: strict F1 0.971, typed F1 0.991, leak-safe recall 0.999 (v2: 0.968 / 0.989 / 0.998).
Per-type strict F1 (v2 in brackets)
| entity | F1 | entity | F1 |
|---|---|---|---|
| DATE | 1.000 (0.999) | CITY | 0.989 (0.987) |
| 0.999 (1.000) | STREET_NAME | 0.989 (0.987) | |
| GOVERNMENT_ID | 0.998 (0.997) | BUILDING_NUMBER | 0.989 (0.986) |
| CREDIT_CARD | 0.997 (0.996) | AGE | 0.977 (0.970) |
| PHONE | 0.997 (0.997) | GIVEN_NAME | 0.905 (0.898) |
| ZIP_CODE | 0.995 (0.994) | SURNAME | 0.894 (0.887) |
Limitations & biases
- Some Dutch over-redaction remains. On CoNLL-2002 NL, 7.0% of the other named entities are still tagged as persons, against 2.7% for the teacher. Brand names built from a person's name are the typical case: "Albert Heijn" (a supermarket chain) is tagged as a person. For Dutch documents where over-redaction matters, use masker.
- Dutch tussenvoegsel surnames may be mistyped on short inputs. v2 sometimes
typed the surname of "Sanne de Groot" or "Pieter van den Berg" as
STREET_NAMEorCITY. We did not reproduce this with v3, but have not ruled it out. When it happens, the characters are still redacted; only the placeholder type is wrong. - This is a compressed model. It trails the full-size masker (v2) by about
0.9 points strict F1 in-distribution and by 7.3 points names F1 on real text. The
loss lands almost entirely on
GIVEN_NAME/SURNAME. Structured types (email, phone, IDs, cards, dates) stay ≥ 0.99. - The out-of-distribution sets cover only English legal text and Dutch news. Measure on your own text, and design for residual leakage rather than assuming full coverage.
Credits & attribution
Initialised from Multilingual-MiniLM and distilled from masker:
- Student initialisation:
microsoft/Multilingual-MiniLM-L12-H384by Microsoft, MIT License. masker-mini v3 starts from its bottom six layers. MiniLM: Wang et al., 2020 (arXiv:2002.10957).
masker is itself a derivative of:
- Backbone / tokenizer lineage:
microsoft/mdeberta-v3-baseby Microsoft, MIT License. masker-mini uses a pruned version of mDeBERTa's SentencePiece tokenizer. DeBERTaV3: He, Gao & Chen, 2021 (arXiv:2111.09543). - Synthetic training data:
ai4privacy/pii-masking-openpii-1mby Ai4Privacy, CC-BY-4.0. - Real-text training data (person annotations, used for masker v2 and as
distillation text):
- Universal NER: en_ewt, da_ddt, pt_bosque, sk_snk, sr_set, sv_talbanken. Mayhew et al., 2024. CC-BY-SA-4.0
- GermEval 2014 (de) and AnCora (es), via the OpenNER standardized collection. CC-BY-4.0
- SUC-X 3.0 (sv), CC-BY-4.0
- KPWr (pl), CC-BY-3.0
- EstNER (et), CC-BY-4.0
- RONEC (ro), MIT
- hr500k (hr), CC-BY-SA-4.0
- Turku NER corpus (fi), CC-BY-SA-4.0
- WikiNER-fr-gold (fr), CC-BY-4.0
- Evaluation only: TAB (MIT) and CoNLL-2002 Dutch (Tjong Kim Sang, 2002).
Please retain these credits in downstream use.
License
Licensed under the Apache License, Version 2.0. See LICENSE and NOTICE.
The underlying components keep their own terms, listed in Credits above.
v2 (the v2 tag) is
also Apache-2.0. Versions before v2 (the
v1 tag) were released
under the Offchain Studio Source License, Version 1.0, and remain under those terms.
Citation
@software{masker_mini,
title = {masker-mini: on-device PII detection for 23 European languages},
year = {2026},
version = {3},
note = {6-layer distillation of masker (mDeBERTa-v3) into a Multilingual-MiniLM student; ONNX + Core ML},
url = {https://huggingface.co/divergentlabs/masker-mini}
}
- Downloads last month
- 361
Model tree for divergentlabs/masker-mini
Datasets used to train divergentlabs/masker-mini
bltlab/open-ner-standardized
ai4privacy/pii-masking-openpii-1m
Collection including divergentlabs/masker-mini
Papers for divergentlabs/masker-mini
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Evaluation results
- Strict span F1 on pii-masking-openpii-1m (validation)self-reported0.971
- Typed F1 on pii-masking-openpii-1m (validation)self-reported0.991
- Leak-safe recall on pii-masking-openpii-1m (validation)self-reported0.999
- Person overlap F1 on TAB (ECHR court judgments, EN, test)self-reported0.887
- Person leak recall on TAB (ECHR court judgments, EN, test)self-reported0.912
- Person overlap F1 on CoNLL-2002 (Dutch, test)self-reported0.843
- Person leak recall on CoNLL-2002 (Dutch, test)self-reported0.931