PrivacyGuard-NER: Lightweight PII Detection Model 🛡️
PrivacyGuard-NER is a compact, high-efficiency token classification model fine-tuned from distilbert-base-uncased for detecting 16 categories of Personally Identifiable Information (PII) with empirically verified ~17ms latency on standard CPU hardware.
Model Details
- Model Type: Token Classification (
DistilBertForTokenClassification) - Base Architecture:
distilbert-base-uncased(~66M parameters) - Model Size: 253.95 MB (safetensors format)
- Number of Labels: 33 (BIO tagging scheme for 16 entity classes + 'O')
- Primary Language: English (specialized for Indian and international PII entities, names, addresses, and ID cards)
- License: Apache 2.0
- Intended Use: Real-time PII detection, redaction, and sanitization for LLM inputs/outputs, RAG knowledge bases, and customer support databases.
Detected Entity Types
| Entity Label | Description |
|---|---|
PERSON |
Full names, given names, family names |
EMAIL |
Personal and corporate email addresses |
PHONE |
International, national, and local telephone numbers |
ADDRESS |
Street addresses, residential locations |
LOCATION |
Cities, regions, states, countries |
ORGANIZATION |
Companies, institutions, government bodies |
DATE_OF_BIRTH |
Birth dates across multiple standard date formats |
CREDIT_CARD |
Credit and debit card numbers |
BANK_ACCOUNT |
Bank account numbers and IBANs |
PAN |
Indian Permanent Account Number |
AADHAAR |
Indian 12-digit Unique Identification Number |
PASSPORT |
International passport identification codes |
IP_ADDRESS |
IPv4 network addresses |
USERNAME |
Account handles, system user identifiers |
URL |
Web URLs, domains, endpoints |
EMPLOYEE_ID |
Organizational staff and badge IDs |
Quick Usage
Using the privacyguard Python Library
from privacyguard import PrivacyGuard
# Load model
guard = PrivacyGuard.from_pretrained("VrajGoti/privacyguard-ner")
# Detect
text = "Contact Alice at alice@company.com or phone +1-555-0199"
result = guard.detect(text)
for ent in result.entities:
print(f"{ent.label}: '{ent.text}' ({ent.confidence:.2%})")
# Redact
clean_text = guard.redact(text, mode="label").redacted_text
print(clean_text)
# Contact [PERSON] at [EMAIL] or phone [PHONE]
Using Hugging Face pipeline
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
model_id = "VrajGoti/privacyguard-ner"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)
ner_pipeline = pipeline(
"token-classification",
model=model,
tokenizer=tokenizer,
aggregation_strategy="simple"
)
results = ner_pipeline("Please email john.doe@secure.io")
print(results)
Training Data & Methodology
The model was trained on a deterministically generated, balanced dataset of 3,000 synthetic text samples with strict boundary validation:
- Train: 2,100 samples
- Validation: 450 samples
- Test: 450 samples
Label Strategy:
- BIO (Begin-Inside-Outside) tagging scheme.
- Subword alignment with first-token classification and -100 masking on subsequent subword tokens to prevent bias during cross-entropy loss calculation.
Evaluation Results
Quantitative Test Metrics (Held-Out Test Set: 450 samples, 955 entities)
- Overall Precision:
0.9990 - Overall Recall:
0.9990 - Overall F1-Score:
0.9990
| Entity Type | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
AADHAAR |
1.0000 | 1.0000 | 1.0000 | 51 |
ADDRESS |
1.0000 | 1.0000 | 1.0000 | 40 |
BANK_ACCOUNT |
1.0000 | 1.0000 | 1.0000 | 34 |
CREDIT_CARD |
1.0000 | 1.0000 | 1.0000 | 36 |
DATE_OF_BIRTH |
1.0000 | 1.0000 | 1.0000 | 34 |
EMAIL |
1.0000 | 1.0000 | 1.0000 | 108 |
EMPLOYEE_ID |
1.0000 | 1.0000 | 1.0000 | 31 |
IP_ADDRESS |
1.0000 | 1.0000 | 1.0000 | 44 |
LOCATION |
1.0000 | 1.0000 | 1.0000 | 49 |
ORGANIZATION |
1.0000 | 1.0000 | 1.0000 | 73 |
PAN |
1.0000 | 1.0000 | 1.0000 | 55 |
PASSPORT |
1.0000 | 1.0000 | 1.0000 | 42 |
PERSON |
1.0000 | 1.0000 | 1.0000 | 209 |
PHONE |
1.0000 | 1.0000 | 1.0000 | 79 |
URL |
1.0000 | 1.0000 | 1.0000 | 29 |
USERNAME |
0.9756 | 0.9756 | 0.9756 | 41 |
Empirical CPU Benchmarks (Hardware: Intel x86_64, Windows, PyTorch 2.12 CPU)
- Single-Text Latency (Mean):
17.02 ms(p50:16.28 ms, min:14.02 ms) - Batched Latency (Batch of 8):
9.88 ms / text(79.06 msper batch) - Throughput:
58.74 texts / second - ONNX Runtime Latency:
17.82 ms(verified exact parity with PyTorch logits)
Limitations & Ethical Considerations
- Synthetic Training Data: The training dataset uses synthetic templates to eliminate privacy risks. Performance on niche domain-specific jargon may require light fine-tuning.
- Language Scope: This checkpoint is validated on English sentences (containing international & Indian identifiers). For multilingual (Hindi, Hinglish, Tamil, Telugu), a fine-tuned
xlm-roberta-basemodel is on the roadmap. - Context Dependency: Highly ambiguous acronyms or short strings might occasionally yield false positives or negatives depending on sentence context.
- Security: While PrivacyGuard-NER achieves high precision and recall, no single automated NER system guarantees 100% PII removal. Multi-layered defense (the included hybrid regex rules for structured tokens like PAN, Aadhaar, Credit Card) is recommended for critical compliance.
- Downloads last month
- -