|
Download README.md from CDT5058/phi-redact: direct link, hf CLI and curl.
- Browser
- Download file 10.6 kB
-
https://huggingface.co/CDT5058/phi-redact/resolve/main/README.md
- Command line
-
hf download hf://CDT5058/phi-redact/README.md
-
curl -L -o README.md https://huggingface.co/CDT5058/phi-redact/resolve/main/README.md
10.6 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| tags: | |
| - token-classification | |
| - named-entity-recognition | |
| - pii | |
| - phi | |
| - redaction | |
| - de-identification | |
| - ner | |
| - bert | |
| - knowledge-distillation | |
| - on-device | |
| pipeline_tag: token-classification | |
| # phi-redact — compact English PII/PHI redaction model | |
| **phi-redact** is an 11M-parameter BERT-mini token classifier for detecting PII | |
| and PHI in English text. It covers 29 entity types with BIOES labels and is | |
| designed for on-device deployment (ONNX, Core ML, LiteRT) where data must not | |
| leave the device. | |
| It was knowledge-distilled from a 66M-parameter DistilBERT teacher and trained | |
| on ~157k documents spanning clinical notes, government benefits text, student | |
| essays, and Wikipedia-style prose. | |
| > **Not a certified de-identification solution.** This model assists redaction | |
| > workflows. It is **not** certified under HIPAA Expert Determination or Safe | |
| > Harbor methods. Evaluate it on your own data before relying on it in | |
| > compliance-sensitive contexts. English only. | |
| --- | |
| ## Model details | |
| | Property | Value | | |
| |---|---| | |
| | Architecture | BERT-mini (google/bert_uncased_L-4_H-256_A-4), 4L/256H/4A | | |
| | Parameters | 11.13M | | |
| | Labels | 121 BIOES tags (30 entity types × 4 positions + O) | | |
| | Max sequence length | 256 tokens | | |
| | Vocabulary | 30,522 WordPiece (bert-base-uncased) | | |
| | fp32 size | 44.5 MB | | |
| | fp16 size | 22.3 MB | | |
| | Est. 4-bit Core ML | ~5.6 MB | | |
| | Est. int8 LiteRT | ~11.2 MB | | |
| | Training method | Knowledge distillation (α=0.7) from distilbert-base-uncased teacher | | |
| | Language | English only | | |
| --- | |
| ## Entity types (30) | |
| `NAME_PATIENT` · `NAME_FAMILY` · `NAME_PROVIDER` · `DATE` · `ADDRESS` · | |
| `PHONE` · `EMAIL` · `SSN` · `MRN` · `HEALTHPLAN` · `ACCOUNT` · `IDNUM` · | |
| `LICENSE` · `EMPLOYER` · `PROFESSION` · `AGE` · `CITY` · `STATE` · `COUNTRY` · | |
| `ZIP` · `URL` · `IP` · `DEVICE` · `ROUTING_NUMBER` · `IBAN` · `CREDITCARD` · | |
| `HOSPITAL` · `FAX` · `PASSPORT` · `VIN` | |
| > **Note on TAX_ID:** included in the schema design but not in any commercial-safe | |
| > training source for v7 (ai4privacy excluded). TAX_ID predictions will be absent. | |
| > Fix queued for v8 (add EIN templates to `gen_synth.py`). | |
| --- | |
| ## Benchmark results (v7) | |
| Strict span-level scoring: a prediction counts only if both the character | |
| boundaries and the label match exactly. | |
| ### Standard benchmarks | |
| | Benchmark | Recall | F1 | Notes | | |
| |---|---|---|---| | |
| | WikiANN (EN, name recall) | **84.3%** | 0.775 | CC-BY-SA 3.0; direct comparison with Desert Ant ✓ | | |
| | FewNERD (EN, name recall) | **85.9%** | 0.782 | CC-BY-SA 4.0; Wikipedia-prose names | | |
| | LA essays (Micro F1 / Recall) | 0.757 / **65.9%** | — | CC-BY 4.0; informal student prose | | |
| | govsynth (government benefits) | **100%** | 1.000 | Synthetic; primary deployment domain | | |
| | **Composite (5-benchmark recall avg)** | **86.1%** | — | See note below | | |
| Composite = (WikiANN + FewNERD + LA essays recall + govsynth + MultiNERD†) / 5. | |
| †MultiNERD (CC-BY-NC-SA 4.0) is included for completeness: v7 recall = 94.3%, | |
| within −0.6 pt of Desert Ant's 94.9% on the same dataset. | |
| ### Nemotron-PII structured benchmark (CC-BY 4.0) | |
| First valid run — 13,134 entities across 25 entity types in 2,000 test | |
| documents. MICRO F1 = **0.860**, Recall = 0.843. | |
| Selected per-entity F1 scores: | |
| | Entity | F1 | Entity | F1 | | |
| |---|---|---|---| | |
| | NAME_PATIENT | 0.944 | SSN | 0.908 | | |
| | NAME_FAMILY | 0.937 | EMAIL | 0.905 | | |
| | MRN | 0.951 | CREDITCARD | 0.905 | | |
| | ADDRESS | 0.949 | IDNUM | 0.908 | | |
| | LICENSE | 0.920 | DATE | 0.859 | | |
| | COUNTRY | 0.919 | URL | 0.828 | | |
| | CITY | 0.904 | DEVICE | 0.791 | | |
| | ROUTING_NUMBER | 0.876 | **PROFESSION** | **0.494** | | |
| **Known weak spot — PROFESSION (F1=0.494, Precision=0.413):** the model | |
| over-fires on job-title nouns in non-PII context ("She saw a doctor" vs | |
| "Occupation: Doctor"). Not a concern for most PHI/PII redaction use cases where | |
| PROFESSION is rarely the sensitive entity, but flag it if your data is | |
| occupation-heavy. | |
| **URL in informal prose:** Nemotron URL F1=0.828 (formal text) but LA essays | |
| URL recall=0.318 (informal prose). The model handles URLs well in structured | |
| documents; it misses them in casual first-person writing. | |
| ### Comparison to Desert Ant Redact v0.4.0 (English) | |
| | Metric | phi-redact v7 | Desert Ant (EN) | | |
| |---|---|---| | |
| | WikiANN name recall | **84.3%** | 69.5% | | |
| | MultiNERD name recall | 94.3% | **94.9%** | | |
| | Composite (5 benchmarks) | **86.1%** | **86.5%** | | |
| | Parameters | **11.13M** | 23M | | |
| | Est. 4-bit on-device size | **~5.6 MB** | 11.6 MB | | |
| | Languages | English only | 27 languages | | |
| | License | **Apache-2.0** | Source-available | | |
| phi-redact v7 is essentially tied with Desert Ant's English composite at half | |
| the parameters and half the on-device footprint. | |
| ### Inference latency | |
| Measured 2026-10-03 against the live demo (corytrimm.com/phi-redact): the | |
| same 15-template typing-length corpus as `evals/latency-bench/bench.py`, | |
| 15 templates × 10 cycles = 150 redactions, timed from the page's own | |
| per-redaction readout — TypeScript rules scan + ONNX detect + merge + | |
| placeholders, end to end, onnxruntime-web WASM (Chromium on Linux). | |
| | Input | p50 | p95 | p99 | | |
| |---|---|---|---| | |
| | Chat-length (150 samples) | **6.0 ms** | 8.0 ms | 9.0 ms | | |
| | Long text (3,959 chars, 139 entities) | 572 ms | — | — | | |
| Model load: 0.5 s cached, 4.4 s cold (11.3 MB int8 ONNX + runtime download). | |
| Caveats: the model was cached for this run; browser version and hardware | |
| details were not obtainable from the measurement setup (Linux VM, not a | |
| phone); readout resolution is whole milliseconds. Not a head-to-head with | |
| any vendor's published number — different hardware, corpus, and setup. | |
| Methodology travels with the number; see `evals/latency-bench/`. | |
| ### Synthetic redaction eval (164 cases) | |
| `evals/synthetic-164/` carries a 164-case synthetic redaction test set, the | |
| scoring harness, and v7's results: **0.812 full-value recall** (term recall | |
| 0.615, precision 0.698, clean-case FP rate 0.120). Perfect 1.000 on | |
| BANK_ACCOUNT, CREDIT_CARD, DRIVERS_LICENSE, EMAIL, PHONE, ROUTING_NUMBER; | |
| weak on NAME (0.347) and ADDRESS (0.333). See the directory README for | |
| methodology and the taxonomy note. | |
| --- | |
| ## Training data | |
| All sources are commercial-safe (CC-BY 4.0 or CC-BY-SA 4.0). No real patient | |
| data was used. | |
| | Source | Size | License | Entity coverage | | |
| |---|---|---|---| | |
| | Synthetic clinical notes (gen_synth.py) | 20k | Synthetic (ours) | Full 29-type schema, clinical PHI | | |
| | Learning Agency Lab PII (Kaggle) | ~5,400 | CC-BY 4.0 | Names, emails, URLs in student essays | | |
| | govsynth (synthetic govt. benefits) | ~20k | Synthetic (ours) | SNAP/Medicaid/WIC case management PII | | |
| | essay_names.jsonl (synthetic prose) | 10k | Synthetic (ours) | Names in biography, news, legal, informal prose | | |
| | FewNERD train split (DFKI-SLT/few-nerd) | 50k | CC-BY-SA 4.0 | Wikipedia person entities | | |
| | WikiANN train split (EN) | ~20k | CC-BY-SA 3.0 | Wikipedia person, location, org | | |
| | WNUT-17 train split | ~3.4k | CC-BY 4.0 | Emerging entities in social media | | |
| | Nemotron-PII train split (nvidia) | 50k | CC-BY 4.0 | 30 PII/PHI types in synthetic documents | | |
| **Excluded:** ai4privacy/pii-masking-400k (non-commercial license — not used in | |
| training or evaluation). MultiNERD (CC-BY-NC-SA 4.0 — evaluation only, not in | |
| training corpus). | |
| --- | |
| ## Training procedure | |
| **Teacher** (distilbert-base-uncased, 66M params): | |
| ``` | |
| epochs 10, batch 32, lr 3e-5, cosine LR, label-smoothing 0.05, device A100 | |
| ``` | |
| **Student** (this model — BERT-mini, 11.13M params): | |
| ``` | |
| epochs 60, batch 64, lr 5e-5, cosine LR, label-smoothing 0.05, | |
| distill-alpha 0.7, device A100 SXM4-40GB (~6h) | |
| ``` | |
| Knowledge distillation loss = 0.7 × KL(teacher logits ∥ student logits) + 0.3 × cross-entropy. | |
| --- | |
| ## How to use | |
| ```python | |
| from transformers import AutoTokenizer, AutoModelForTokenClassification | |
| import torch | |
| tokenizer = AutoTokenizer.from_pretrained("CDT5058/phi-redact") | |
| model = AutoModelForTokenClassification.from_pretrained("CDT5058/phi-redact") | |
| model.eval() | |
| text = "Patient Maria Garcia, DOB 1985-03-22, MRN 4471829, called 415-555-0189." | |
| enc = tokenizer(text, return_tensors="pt", truncation=True, max_length=256) | |
| with torch.no_grad(): | |
| logits = model(**enc).logits # [1, seq, 117] | |
| pred_ids = logits.argmax(-1)[0].tolist() # [seq] | |
| id2label = model.config.id2label | |
| # Decode BIOES tags to character spans | |
| tokens = tokenizer.convert_ids_to_tokens(enc["input_ids"][0]) | |
| offsets = enc.encodings[0].offsets # (char_start, char_end) per token | |
| spans = [] | |
| span_start, span_label = None, None | |
| for i, tag_id in enumerate(pred_ids): | |
| tag = id2label[tag_id] | |
| prefix, *label_parts = tag.split("-") if "-" in tag else (tag, []) | |
| label = "-".join(label_parts) | |
| if prefix == "S": | |
| spans.append((offsets[i][0], offsets[i][1], label)) | |
| elif prefix == "B": | |
| span_start, span_label = offsets[i][0], label | |
| elif prefix == "E" and span_start is not None: | |
| spans.append((span_start, offsets[i][1], span_label)) | |
| span_start = None | |
| elif prefix in ("O", "B"): | |
| span_start = None # reset on unexpected transition | |
| for start, end, label in spans: | |
| print(f"[{start}:{end}] {label!r:20s} {text[start:end]!r}") | |
| ``` | |
| For production use, layer the deterministic `rules.py` (Luhn, ABA routing | |
| checksum, SSN, IBAN, regex patterns) on top — it wins on structured identifiers | |
| regardless of model confidence. | |
| --- | |
| ## Limitations | |
| - **English only.** The bert-base-uncased WordPiece vocabulary has no meaningful | |
| cross-lingual representations. French, Spanish, or other languages will produce | |
| silent misses — not safe degradation. | |
| - **PROFESSION over-fires** (F1=0.494 on Nemotron). Treat PROFESSION predictions | |
| with lower confidence or disable the entity type if not needed. | |
| - **URL in informal prose** (recall=0.318 on student essays vs 0.828 on formal | |
| text). Check URL recall on your specific document type. | |
| - **Trained on synthetic data.** Real-world typos, OCR noise, and unusual | |
| formatting will degrade performance relative to the benchmarks above. | |
| - **Not a certified de-identification system.** Does not satisfy HIPAA Expert | |
| Determination or Safe Harbor. Use as a detection aid alongside human review. | |
| --- | |
| ## License | |
| Apache-2.0. Training data licenses: CC-BY 4.0 (Nemotron-PII, LA essays, | |
| WNUT-17), CC-BY-SA 4.0 (FewNERD), CC-BY-SA 3.0 (WikiANN). CC-BY-SA ShareAlike | |
| requirements apply to the datasets, not the model weights, but consult your | |
| legal counsel for commercial deployment. | |
| --- | |
| *Card generated 2026-09-25 · Repo: CDT5058/phi-redact* | |