--- license: apache-2.0 language: - en tags: - token-classification - named-entity-recognition - pii - phi - redaction - de-identification - ner - bert - knowledge-distillation - on-device pipeline_tag: token-classification --- # phi-redact — compact English PII/PHI redaction model **phi-redact** is an 11M-parameter BERT-mini token classifier for detecting PII and PHI in English text. It covers 29 entity types with BIOES labels and is designed for on-device deployment (ONNX, Core ML, LiteRT) where data must not leave the device. It was knowledge-distilled from a 66M-parameter DistilBERT teacher and trained on ~157k documents spanning clinical notes, government benefits text, student essays, and Wikipedia-style prose. > **Not a certified de-identification solution.** This model assists redaction > workflows. It is **not** certified under HIPAA Expert Determination or Safe > Harbor methods. Evaluate it on your own data before relying on it in > compliance-sensitive contexts. English only. --- ## Model details | Property | Value | |---|---| | Architecture | BERT-mini (google/bert_uncased_L-4_H-256_A-4), 4L/256H/4A | | Parameters | 11.13M | | Labels | 121 BIOES tags (30 entity types × 4 positions + O) | | Max sequence length | 256 tokens | | Vocabulary | 30,522 WordPiece (bert-base-uncased) | | fp32 size | 44.5 MB | | fp16 size | 22.3 MB | | Est. 4-bit Core ML | ~5.6 MB | | Est. int8 LiteRT | ~11.2 MB | | Training method | Knowledge distillation (α=0.7) from distilbert-base-uncased teacher | | Language | English only | --- ## Entity types (30) `NAME_PATIENT` · `NAME_FAMILY` · `NAME_PROVIDER` · `DATE` · `ADDRESS` · `PHONE` · `EMAIL` · `SSN` · `MRN` · `HEALTHPLAN` · `ACCOUNT` · `IDNUM` · `LICENSE` · `EMPLOYER` · `PROFESSION` · `AGE` · `CITY` · `STATE` · `COUNTRY` · `ZIP` · `URL` · `IP` · `DEVICE` · `ROUTING_NUMBER` · `IBAN` · `CREDITCARD` · `HOSPITAL` · `FAX` · `PASSPORT` · `VIN` > **Note on TAX_ID:** included in the schema design but not in any commercial-safe > training source for v7 (ai4privacy excluded). TAX_ID predictions will be absent. > Fix queued for v8 (add EIN templates to `gen_synth.py`). --- ## Benchmark results (v7) Strict span-level scoring: a prediction counts only if both the character boundaries and the label match exactly. ### Standard benchmarks | Benchmark | Recall | F1 | Notes | |---|---|---|---| | WikiANN (EN, name recall) | **84.3%** | 0.775 | CC-BY-SA 3.0; direct comparison with Desert Ant ✓ | | FewNERD (EN, name recall) | **85.9%** | 0.782 | CC-BY-SA 4.0; Wikipedia-prose names | | LA essays (Micro F1 / Recall) | 0.757 / **65.9%** | — | CC-BY 4.0; informal student prose | | govsynth (government benefits) | **100%** | 1.000 | Synthetic; primary deployment domain | | **Composite (5-benchmark recall avg)** | **86.1%** | — | See note below | Composite = (WikiANN + FewNERD + LA essays recall + govsynth + MultiNERD†) / 5. †MultiNERD (CC-BY-NC-SA 4.0) is included for completeness: v7 recall = 94.3%, within −0.6 pt of Desert Ant's 94.9% on the same dataset. ### Nemotron-PII structured benchmark (CC-BY 4.0) First valid run — 13,134 entities across 25 entity types in 2,000 test documents. MICRO F1 = **0.860**, Recall = 0.843. Selected per-entity F1 scores: | Entity | F1 | Entity | F1 | |---|---|---|---| | NAME_PATIENT | 0.944 | SSN | 0.908 | | NAME_FAMILY | 0.937 | EMAIL | 0.905 | | MRN | 0.951 | CREDITCARD | 0.905 | | ADDRESS | 0.949 | IDNUM | 0.908 | | LICENSE | 0.920 | DATE | 0.859 | | COUNTRY | 0.919 | URL | 0.828 | | CITY | 0.904 | DEVICE | 0.791 | | ROUTING_NUMBER | 0.876 | **PROFESSION** | **0.494** | **Known weak spot — PROFESSION (F1=0.494, Precision=0.413):** the model over-fires on job-title nouns in non-PII context ("She saw a doctor" vs "Occupation: Doctor"). Not a concern for most PHI/PII redaction use cases where PROFESSION is rarely the sensitive entity, but flag it if your data is occupation-heavy. **URL in informal prose:** Nemotron URL F1=0.828 (formal text) but LA essays URL recall=0.318 (informal prose). The model handles URLs well in structured documents; it misses them in casual first-person writing. ### Comparison to Desert Ant Redact v0.4.0 (English) | Metric | phi-redact v7 | Desert Ant (EN) | |---|---|---| | WikiANN name recall | **84.3%** | 69.5% | | MultiNERD name recall | 94.3% | **94.9%** | | Composite (5 benchmarks) | **86.1%** | **86.5%** | | Parameters | **11.13M** | 23M | | Est. 4-bit on-device size | **~5.6 MB** | 11.6 MB | | Languages | English only | 27 languages | | License | **Apache-2.0** | Source-available | phi-redact v7 is essentially tied with Desert Ant's English composite at half the parameters and half the on-device footprint. ### Inference latency Measured 2026-10-03 against the live demo (corytrimm.com/phi-redact): the same 15-template typing-length corpus as `evals/latency-bench/bench.py`, 15 templates × 10 cycles = 150 redactions, timed from the page's own per-redaction readout — TypeScript rules scan + ONNX detect + merge + placeholders, end to end, onnxruntime-web WASM (Chromium on Linux). | Input | p50 | p95 | p99 | |---|---|---|---| | Chat-length (150 samples) | **6.0 ms** | 8.0 ms | 9.0 ms | | Long text (3,959 chars, 139 entities) | 572 ms | — | — | Model load: 0.5 s cached, 4.4 s cold (11.3 MB int8 ONNX + runtime download). Caveats: the model was cached for this run; browser version and hardware details were not obtainable from the measurement setup (Linux VM, not a phone); readout resolution is whole milliseconds. Not a head-to-head with any vendor's published number — different hardware, corpus, and setup. Methodology travels with the number; see `evals/latency-bench/`. ### Synthetic redaction eval (164 cases) `evals/synthetic-164/` carries a 164-case synthetic redaction test set, the scoring harness, and v7's results: **0.812 full-value recall** (term recall 0.615, precision 0.698, clean-case FP rate 0.120). Perfect 1.000 on BANK_ACCOUNT, CREDIT_CARD, DRIVERS_LICENSE, EMAIL, PHONE, ROUTING_NUMBER; weak on NAME (0.347) and ADDRESS (0.333). See the directory README for methodology and the taxonomy note. --- ## Training data All sources are commercial-safe (CC-BY 4.0 or CC-BY-SA 4.0). No real patient data was used. | Source | Size | License | Entity coverage | |---|---|---|---| | Synthetic clinical notes (gen_synth.py) | 20k | Synthetic (ours) | Full 29-type schema, clinical PHI | | Learning Agency Lab PII (Kaggle) | ~5,400 | CC-BY 4.0 | Names, emails, URLs in student essays | | govsynth (synthetic govt. benefits) | ~20k | Synthetic (ours) | SNAP/Medicaid/WIC case management PII | | essay_names.jsonl (synthetic prose) | 10k | Synthetic (ours) | Names in biography, news, legal, informal prose | | FewNERD train split (DFKI-SLT/few-nerd) | 50k | CC-BY-SA 4.0 | Wikipedia person entities | | WikiANN train split (EN) | ~20k | CC-BY-SA 3.0 | Wikipedia person, location, org | | WNUT-17 train split | ~3.4k | CC-BY 4.0 | Emerging entities in social media | | Nemotron-PII train split (nvidia) | 50k | CC-BY 4.0 | 30 PII/PHI types in synthetic documents | **Excluded:** ai4privacy/pii-masking-400k (non-commercial license — not used in training or evaluation). MultiNERD (CC-BY-NC-SA 4.0 — evaluation only, not in training corpus). --- ## Training procedure **Teacher** (distilbert-base-uncased, 66M params): ``` epochs 10, batch 32, lr 3e-5, cosine LR, label-smoothing 0.05, device A100 ``` **Student** (this model — BERT-mini, 11.13M params): ``` epochs 60, batch 64, lr 5e-5, cosine LR, label-smoothing 0.05, distill-alpha 0.7, device A100 SXM4-40GB (~6h) ``` Knowledge distillation loss = 0.7 × KL(teacher logits ∥ student logits) + 0.3 × cross-entropy. --- ## How to use ```python from transformers import AutoTokenizer, AutoModelForTokenClassification import torch tokenizer = AutoTokenizer.from_pretrained("CDT5058/phi-redact") model = AutoModelForTokenClassification.from_pretrained("CDT5058/phi-redact") model.eval() text = "Patient Maria Garcia, DOB 1985-03-22, MRN 4471829, called 415-555-0189." enc = tokenizer(text, return_tensors="pt", truncation=True, max_length=256) with torch.no_grad(): logits = model(**enc).logits # [1, seq, 117] pred_ids = logits.argmax(-1)[0].tolist() # [seq] id2label = model.config.id2label # Decode BIOES tags to character spans tokens = tokenizer.convert_ids_to_tokens(enc["input_ids"][0]) offsets = enc.encodings[0].offsets # (char_start, char_end) per token spans = [] span_start, span_label = None, None for i, tag_id in enumerate(pred_ids): tag = id2label[tag_id] prefix, *label_parts = tag.split("-") if "-" in tag else (tag, []) label = "-".join(label_parts) if prefix == "S": spans.append((offsets[i][0], offsets[i][1], label)) elif prefix == "B": span_start, span_label = offsets[i][0], label elif prefix == "E" and span_start is not None: spans.append((span_start, offsets[i][1], span_label)) span_start = None elif prefix in ("O", "B"): span_start = None # reset on unexpected transition for start, end, label in spans: print(f"[{start}:{end}] {label!r:20s} {text[start:end]!r}") ``` For production use, layer the deterministic `rules.py` (Luhn, ABA routing checksum, SSN, IBAN, regex patterns) on top — it wins on structured identifiers regardless of model confidence. --- ## Limitations - **English only.** The bert-base-uncased WordPiece vocabulary has no meaningful cross-lingual representations. French, Spanish, or other languages will produce silent misses — not safe degradation. - **PROFESSION over-fires** (F1=0.494 on Nemotron). Treat PROFESSION predictions with lower confidence or disable the entity type if not needed. - **URL in informal prose** (recall=0.318 on student essays vs 0.828 on formal text). Check URL recall on your specific document type. - **Trained on synthetic data.** Real-world typos, OCR noise, and unusual formatting will degrade performance relative to the benchmarks above. - **Not a certified de-identification system.** Does not satisfy HIPAA Expert Determination or Safe Harbor. Use as a detection aid alongside human review. --- ## License Apache-2.0. Training data licenses: CC-BY 4.0 (Nemotron-PII, LA essays, WNUT-17), CC-BY-SA 4.0 (FewNERD), CC-BY-SA 3.0 (WikiANN). CC-BY-SA ShareAlike requirements apply to the datasets, not the model weights, but consult your legal counsel for commercial deployment. --- *Card generated 2026-09-25 · Repo: CDT5058/phi-redact*