pii
redaction
nnue

piitag

A per-byte PII tagger in the NNUE shape: hashed byte n-grams feed two int16 accumulators, then a small int8 MLP scores every byte. It is the model the Rust crate cynch-piitag embeds; the crate pins this repository's commit and the file's sha256 in its weights.lock.

piitag.piitag is the .piitag model file, format version 4: an 8-byte magic, a little-endian u32 header length, a flat JSON header, then 64-byte-aligned little-endian arrays (the two n-gram tables and their biases, the MLP, the type head and the sorted u64 name-dictionary hashes).

Training data and attribution

Trained on nvidia/Nemotron-PII at revision b70ffaf5ff39e079776134c5bf4381f00a9fd1ed, by NVIDIA, licensed CC-BY-4.0. These weights are learned from that dataset; this card is its attribution. Test metrics are on its held-out test split.

The name dictionary is built from public-domain US government data: the Social Security Administration's baby names (via hadley/data-baby-names) and the Census 2000 surname table (via fivethirtyeight/data).

Results (held-out test split)

The threshold threshold_q = -12238 was chosen on validation data for 0.980 byte recall.

model span leak rate fully missed byte precision byte recall PR-AUC
piitag 0.0104 0.0066 0.830 0.979 0.981
linear baseline 0.0170 0.0132 0.644 0.978 0.959

A span leaks when any of its bytes is left unredacted. The integer Rust inference matches the Python reference exactly: 0 mismatches on 2000 test buffers.

Entity types

Each span is named by a type head (one of 34 types); a gold span's type is right when the type summed over its blocks matches. Accuracy 0.970 over 466210 test spans, 0.954 averaged over types.

type test spans accuracy
account_number 16698 0.944
api_key 4667 0.971
bank_routing_number 8354 0.949
biometric_identifier 11379 0.979
certificate_license_number 3002 0.943
coordinate 7677 0.990
credit_debit_card 12936 0.982
customer_id 20502 0.964
cvv 4843 0.942
date_of_birth 18079 0.994
device_identifier 2507 0.927
email 53930 0.997
employee_id 8878 0.960
fax_number 6442 0.868
health_plan_beneficiary_number 10561 0.972
http_cookie 5014 0.991
ipv4 6078 0.996
ipv6 3139 0.999
license_plate 4083 0.971
mac_address 4250 0.996
medical_record_number 11098 0.972
national_id 2848 0.926
password 6986 0.971
person_name 143624 0.981
phone_number 23963 0.961
pin 6456 0.901
postcode 6280 0.938
ssn 6062 0.961
street_address 16875 0.981
swift_bic 5559 0.990
tax_id 1302 0.871
unique_id 1777 0.793
user_name 15800 0.864
vehicle_identifier 4561 0.989

Spec

{"k": 8, "ctx_orders": [1, 2, 3], "ctx_l": 32, "ctx_r": 32, "log2_ctx": 13, "loc_offsets": [-2, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9], "loc_orders": [1, 2], "log2_loc": 12, "h1": 32, "h2": 32, "h3": 32, "dilate": 1}

Dictionary names: 152167.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train scottmas/piitag