PrivacyGuard-NER: Lightweight PII Detection Model 🛡️

PrivacyGuard-NER is a compact, high-efficiency token classification model fine-tuned from distilbert-base-uncased for detecting 16 categories of Personally Identifiable Information (PII) with empirically verified ~17ms latency on standard CPU hardware.

Model Details

  • Model Type: Token Classification (DistilBertForTokenClassification)
  • Base Architecture: distilbert-base-uncased (~66M parameters)
  • Model Size: 253.95 MB (safetensors format)
  • Number of Labels: 33 (BIO tagging scheme for 16 entity classes + 'O')
  • Primary Language: English (specialized for Indian and international PII entities, names, addresses, and ID cards)
  • License: Apache 2.0
  • Intended Use: Real-time PII detection, redaction, and sanitization for LLM inputs/outputs, RAG knowledge bases, and customer support databases.

Detected Entity Types

Entity Label Description
PERSON Full names, given names, family names
EMAIL Personal and corporate email addresses
PHONE International, national, and local telephone numbers
ADDRESS Street addresses, residential locations
LOCATION Cities, regions, states, countries
ORGANIZATION Companies, institutions, government bodies
DATE_OF_BIRTH Birth dates across multiple standard date formats
CREDIT_CARD Credit and debit card numbers
BANK_ACCOUNT Bank account numbers and IBANs
PAN Indian Permanent Account Number
AADHAAR Indian 12-digit Unique Identification Number
PASSPORT International passport identification codes
IP_ADDRESS IPv4 network addresses
USERNAME Account handles, system user identifiers
URL Web URLs, domains, endpoints
EMPLOYEE_ID Organizational staff and badge IDs

Quick Usage

Using the privacyguard Python Library

from privacyguard import PrivacyGuard

# Load model
guard = PrivacyGuard.from_pretrained("VrajGoti/privacyguard-ner")

# Detect
text = "Contact Alice at alice@company.com or phone +1-555-0199"
result = guard.detect(text)

for ent in result.entities:
    print(f"{ent.label}: '{ent.text}' ({ent.confidence:.2%})")

# Redact
clean_text = guard.redact(text, mode="label").redacted_text
print(clean_text)
# Contact [PERSON] at [EMAIL] or phone [PHONE]

Using Hugging Face pipeline

from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline

model_id = "VrajGoti/privacyguard-ner"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)

ner_pipeline = pipeline(
    "token-classification",
    model=model,
    tokenizer=tokenizer,
    aggregation_strategy="simple"
)

results = ner_pipeline("Please email john.doe@secure.io")
print(results)

Training Data & Methodology

The model was trained on a deterministically generated, balanced dataset of 3,000 synthetic text samples with strict boundary validation:

  • Train: 2,100 samples
  • Validation: 450 samples
  • Test: 450 samples

Label Strategy:

  • BIO (Begin-Inside-Outside) tagging scheme.
  • Subword alignment with first-token classification and -100 masking on subsequent subword tokens to prevent bias during cross-entropy loss calculation.

Evaluation Results

Quantitative Test Metrics (Held-Out Test Set: 450 samples, 955 entities)

  • Overall Precision: 0.9990
  • Overall Recall: 0.9990
  • Overall F1-Score: 0.9990
Entity Type Precision Recall F1-Score Support
AADHAAR 1.0000 1.0000 1.0000 51
ADDRESS 1.0000 1.0000 1.0000 40
BANK_ACCOUNT 1.0000 1.0000 1.0000 34
CREDIT_CARD 1.0000 1.0000 1.0000 36
DATE_OF_BIRTH 1.0000 1.0000 1.0000 34
EMAIL 1.0000 1.0000 1.0000 108
EMPLOYEE_ID 1.0000 1.0000 1.0000 31
IP_ADDRESS 1.0000 1.0000 1.0000 44
LOCATION 1.0000 1.0000 1.0000 49
ORGANIZATION 1.0000 1.0000 1.0000 73
PAN 1.0000 1.0000 1.0000 55
PASSPORT 1.0000 1.0000 1.0000 42
PERSON 1.0000 1.0000 1.0000 209
PHONE 1.0000 1.0000 1.0000 79
URL 1.0000 1.0000 1.0000 29
USERNAME 0.9756 0.9756 0.9756 41

Empirical CPU Benchmarks (Hardware: Intel x86_64, Windows, PyTorch 2.12 CPU)

  • Single-Text Latency (Mean): 17.02 ms (p50: 16.28 ms, min: 14.02 ms)
  • Batched Latency (Batch of 8): 9.88 ms / text (79.06 ms per batch)
  • Throughput: 58.74 texts / second
  • ONNX Runtime Latency: 17.82 ms (verified exact parity with PyTorch logits)

Limitations & Ethical Considerations

  1. Synthetic Training Data: The training dataset uses synthetic templates to eliminate privacy risks. Performance on niche domain-specific jargon may require light fine-tuning.
  2. Language Scope: This checkpoint is validated on English sentences (containing international & Indian identifiers). For multilingual (Hindi, Hinglish, Tamil, Telugu), a fine-tuned xlm-roberta-base model is on the roadmap.
  3. Context Dependency: Highly ambiguous acronyms or short strings might occasionally yield false positives or negatives depending on sentence context.
  4. Security: While PrivacyGuard-NER achieves high precision and recall, no single automated NER system guarantees 100% PII removal. Multi-layered defense (the included hybrid regex rules for structured tokens like PAN, Aadhaar, Credit Card) is recommended for critical compliance.
Downloads last month
-
Safetensors
Model size
66.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support