|
Download README.md from purebyte/pii: direct link, hf CLI and curl.
- Browser
- Download file 14.6 kB
-
https://huggingface.co/purebyte/pii/resolve/main/README.md
- Command line
-
hf download hf://purebyte/pii/README.md
-
curl -L -o README.md https://huggingface.co/purebyte/pii/resolve/main/README.md
14.6 kB
| license: other | |
| license_name: purebyte-model-license | |
| license_link: https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/MODEL_LICENSE.md | |
| pipeline_tag: token-classification | |
| tags: | |
| - pii | |
| - privacy | |
| - redaction | |
| - token-classification | |
| - byte-level | |
| - state-space-model | |
| - ternary | |
| - cpu | |
| # Model card: pii | |
| > The model file is `purebyte-pii-1.0.0.gguf` in this repository; its SHA-256 is next to it. It runs with the PureByte runtime ([github.com/purebyte-ai/purebyte](https://github.com/purebyte-ai/purebyte)): `pip install purebyte`, then `purebyte models pull pii`, which downloads the same file from the runtime's release and checks its checksum. | |
| An **AI Specialist** that finds personal data in any text and masks it: names, contact details, identity and account | |
| numbers, dates, places, companies, credentials and the like, all as one type, `PII`. It marks byte spans; the runtime | |
| writes a copy of the input in which only those spans are replaced by markers such as `[PII_1]`. It runs on ordinary CPUs | |
| and never sends anything anywhere. | |
| It is one of the three example specialists of PureByte 1.0, which show what one byte-level architecture does on different | |
| jobs; the architecture, the runtime and the training stack are described in the [README](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/README.md). | |
| Version 1.0.0 is the first release. Its recipe, data scripts and pre-registered criterion are in the training | |
| repository, [purebyte-train](https://github.com/purebyte-ai/purebyte-train): | |
| [specialists/pii](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/pii/README.md). | |
| | | | | |
| |---|---| | |
| | Version | 1.0.0 | | |
| | Task | Masking of personal data as byte spans of one type, `PII` (label-agnostic), for redaction by copy | | |
| | Input | Any file up to 64 MiB, in windows of 512 bytes (stride 384); inputs of any length are analyzed, even very short ones (the file declares `purebyte.window.min = 1`) | | |
| | Output | A redacted copy (`purebyte redact`): the input byte for byte with each span replaced by `[PII_n]`, the same value always the same marker; a report of the spans (offsets, markers, confidence, never the values); or findings with masked snippets (`purebyte scan`). See [spec/OUTPUT.md](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/spec/OUTPUT.md), section 9 | | |
| | Size | 4.5 MiB (4,764,864 bytes, one file; the n-gram memory is stored at 4 bits) | | |
| | Architecture | Byte-level state-space model with ternary weights (four `ssm_v2` blocks of width 256, read left to right) and hashed n-gram memory tables (byte 2-, 3- and 4-grams; 3 x 2^15 x 64), a span-tagging head over one type gated by a window head with max pooling ([docs/architecture.md](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/docs/architecture.md)); 8,268,128 parameters (8.3 M) in the network: the byte embedding, the blocks, the final norm, and the n-gram tables with their projection, 6.3 M of them in the tables. The two heads add 34,439 (8,302,567 in all); the per-group scales of the ternary weights are derived, not counted | | |
| | Default decision rule | One model, at the operating point stored in its file (`--bias` moves it) | | |
| | Post-processing | The `redact` profile: consistent markers, and the copy property (every byte outside the spans is in the output, unchanged) checked before anything is written | | |
| | Runtime | `purebyte` 1.0.0 or newer, on an x86-64 or arm64 CPU (other little-endian CPUs through the portable build); no GPU, no network | | |
| | License | [PureByte Model License](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/MODEL_LICENSE.md) (the code of the runtime is Apache-2.0) | | |
| | Files | `purebyte-pii-1.0.0.gguf`; checksum in [models.json](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/models/models.json) | | |
| | Recipe | [specialists/pii](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/pii/README.md) in purebyte-train: the data scripts, the recipe and the evaluation, to rebuild or improve it | | |
| ```bash | |
| purebyte models pull pii | |
| purebyte redact --model pii ticket.txt > ticket.redacted.txt | |
| purebyte redact --model pii --report ticket.report.json --map ticket.map.json ticket.txt > ticket.redacted.txt | |
| purebyte redact --restore ticket.redacted.txt --map ticket.map.json # the original again (the map holds the values) | |
| ``` | |
| ## Intended use | |
| - Redacting logs, support tickets, chat transcripts, e-mails, documents and records before they are shared, stored, | |
| indexed or sent to a hosted model. | |
| - A local privacy filter in front of prompts and tool calls, and small decisions through the HTTP API. | |
| - Pseudonymized copies with a local map to restore the original. The map holds the values: treat it as a secret. | |
| ## Out of scope | |
| - **Telling kinds of personal data apart.** Every span is `PII`: an e-mail address and a name get the same kind of | |
| marker. An experimental typed recipe exists in | |
| [specialists/pii](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/pii/README.md#experimental-typed-spans) | |
| of purebyte-train; no typed model is released. | |
| - **Keeping part of the personal data.** Its policy is the benchmark's and its training data's: companies, | |
| occupations, dates, cities and URLs are masked too. If they must stay, this model is the wrong tool. | |
| - **Proving that a text holds no personal data**, or anonymization in a legal sense. It misses a share of what it looks | |
| for, and what it leaves (context, rare details) can still identify a person. | |
| - **Checking values.** It never validates or looks up anything. | |
| - **Credentials in source code.** It masks the passwords and keys its data label, but [secrets-code](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/models/secrets-code.md) is | |
| the specialist for code and configuration. | |
| - **Personal data in images, audio or binary formats.** | |
| ## What it masks | |
| Every span its training data label as personal data, as one type: | |
| | Family | Examples | | |
| |---|---| | |
| | People | given names, surnames, full names, titles, gender, age, nationality, occupation, education | | |
| | Contact | e-mail addresses, phone and fax numbers, street addresses, building numbers, postcodes, cities, states, countries, coordinates | | |
| | Identifiers | national identity, passport, driver's license, tax and social security numbers, customer, employee, account, medical record, device and vehicle identifiers, license plates | | |
| | Finance | payment card numbers and codes, IBAN, BBAN, SWIFT/BIC, routing numbers, PINs | | |
| | Online | usernames, passwords, API keys, URLs, IP and MAC addresses, cookies | | |
| | Dates and organizations | dates and times, dates of birth, company names | | |
| ## Evaluation | |
| On the [PII Masking Benchmark](https://huggingface.co/datasets/piimb/pii-masking-benchmark) (revision `7d797b9f`, | |
| CC BY-NC 4.0, evaluation only): its `sentences` subset (150,022 sentences in six tasks), character-level F2, | |
| label-agnostic, the mean of the six tasks, the metric of its leaderboard. Measured with the `purebyte` CLI on the | |
| released file at its default operating point, and on the file of each of the recipe's six seeds | |
| ([benchmarks/piimb](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/benchmarks/piimb/README.md)); the two models below them are the leaderboard's, and masking every | |
| character is a reference point. | |
| | Model | ai4privacy-en | ai4privacy-multi | nemotron-pii | gretel | privy | mapa-eur-lex | **Mean F2** | Mean precision | | |
| |---|---:|---:|---:|---:|---:|---:|---:|---:| | |
| | **pii** 1.0.0 (8.3 M parameters) | 0.9569 | 0.9498 | 0.8918 | 0.9463 | 0.5360 | 0.3325 | **0.7689** | 0.773 | | |
| | pii 1.0.0 recipe, mean of its six seeds [range] | 0.9624 | 0.9569 | 0.9003 | 0.9522 | 0.5183 | 0.3744 | 0.7774 [0.7689-0.7853] | 0.765 | | |
| | openai/privacy-filter (1.4 B parameters, about 50 M active) | 0.8426 | 0.8052 | 0.6068 | 0.8687 | 0.5362 | 0.3129 | 0.6621 | 0.751 | | |
| | Leaderboard's first: OpenMed/privacy-filter-multilingual-v2 (1.4 B, about 50 M active) | 0.9754 | 0.9701 | 0.9070 | 0.9727 | 0.8176 | 0.3728 | 0.8359 | 0.832 | | |
| | Masking every character | 0.5883 | 0.5867 | 0.4300 | 0.6443 | 0.2679 | 0.1364 | 0.4423 | 0.157 | | |
| The pre-registered criterion (the mean of six seeds above 0.6621 mean F2 and 0.8426 on ai4privacy-en, with a mean | |
| precision of at least 0.70) is met at the default operating point: 0.7774, 0.9624 and 0.765. As first written, it | |
| named the "PIIMB mode" point instead; we moved it to the default point a day later, after the pool-only mix of the | |
| recipe had missed the precision floor at "PIIMB mode" and before the released recipe was scored. At "PIIMB mode" the | |
| released recipe misses the floor too (0.686): by the criterion as first written, the verdict is a fail | |
| ([details](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/benchmarks/RESULTS.md#pii-on-the-pii-masking-benchmark)). The released file is the seed that the rule of | |
| our training stack chose on the training pool's own `select` documents, before the benchmark was run; on the benchmark | |
| it is the lowest of the six, and its own figures are the first row. Without the 1,018 benchmark sentences that share a | |
| 64-byte block with the training data it scores 0.7686. The recipe variant that ships is not the one that scores best | |
| here ([Training data](#training-data)). Every seed, the "PIIMB mode" operating point and the conditions: | |
| [benchmarks/RESULTS.md](https://github.com/purebyte-ai/purebyte/blob/pii-1.0.0/benchmarks/RESULTS.md#pii-on-the-pii-masking-benchmark). | |
| **How to read these figures.** Four of the six tasks (ai4privacy-en, ai4privacy-multi, nemotron-pii, gretel) are | |
| sampled from the test or validation splits of datasets whose TRAIN splits the model learned from: same generators, | |
| other documents. privy | |
| (protocol traces: JSON, XML, SQL) and mapa-eur-lex (legal texts in 21 languages) resemble nothing it learned from. F2 | |
| (尾 = 2) weights recall above precision, and masking everything scores 0.4423, which is why the criterion also requires | |
| a mean precision of at least 0.70. | |
| ## Known limitations | |
| - **Its training data are synthetic documents** of four generators, in 23 languages. Real documents, other languages | |
| and other genres (protocol traces, legal texts, source code) are harder: 0.536 on privy (protocol traces) and 0.333 | |
| on mapa-eur-lex (legal texts) above. | |
| - **Short inputs.** The window head that gates the spans was trained on 512-byte windows; on a single short sentence it | |
| says "nothing here" more often. On the benchmark's sentences, switching the gate off and raising the bias from -0.5 to | |
| 2.0 (the "PIIMB mode" point, chosen on the pool's own `select` documents cut into sentences) lifts the mean F2 of the | |
| released seed from 0.769 to 0.819, but its mean precision falls from 0.773 to 0.680, below the recipe's floor of 0.70: | |
| the file keeps the default point (measured in PyTorch: the CLI cannot switch the gate off). | |
| - **Boundaries are read left to right** (a causal model): a span can start a little late or end a little early. | |
| - **Over-masking by design**: companies, occupations, dates and places are masked wherever its data label them. | |
| - **Confidence values are not calibrated yet.** | |
| - **The n-gram memory is stored at 4 bits**, which changes the spans of a small share of windows against the model | |
| before export: 2 of the 40 windows that the `verify` stage of the training stack compares (the engine gives the same | |
| spans as Python on all 40). | |
| ## Training data | |
| 192,000 windows of 512 bytes. 80 % are cut from documents of the TRAIN splits of four public datasets of synthetic | |
| personal data, rebuilt from pinned downloads by the scripts in | |
| [data/](https://github.com/purebyte-ai/purebyte-train/blob/main/data/README.md#the-pii-pool) of purebyte-train: Ai4Privacy | |
| OpenPII-1M (23 languages), NVIDIA Nemotron-PII, Gretel PII masking EN and Gretel synthetic PII finance (7 languages); | |
| 387,391 documents after a hard overlap check that drops any document sharing a sentence, a line or a masked template | |
| with the benchmark or the other evaluation sets. 20 % are the specialist's generated **structured forms**: records | |
| (JSON, YAML, CSV, SQL, XML, `.env`...), log lines and code with personal data in them, all of it one type, and | |
| look-alike values (order ids, timestamps, postcodes, company names, URLs...) taught neither way. No value belongs to | |
| anyone. No test or validation split, and nothing of the benchmark, was used for training or tuning; the seed that | |
| ships is the best of six on the pool's own `select` documents. | |
| **Why this mix.** The recipe was also trained on the pool alone (three seeds of each mix, default operating point). On | |
| the benchmark, the pool-only mix scored a mean F2 of 0.786 and the structured-forms mix 0.779: a tie by the | |
| pre-registered rule (a difference under 0.01), and by that rule a tie keeps the mix measured first, the pool-only one. | |
| On a test set of 2,000 generated documents (code, structured records, logs, prose and e-mails, built from held-out | |
| value lists, formats and code), the pool-only mix left 26 % of the personal data unmasked (38 % in code, 28 % in | |
| structured records, 21 % in logs, 21 % in e-mails, 27 % in prose) and the structured-forms mix 11 % (1.3 %, 1.9 %, | |
| 0.04 %, 19 % and 27 %). The release follows the product, against the tie rule: the structured-forms mix. That test set | |
| comes from the same generator as the structured-forms windows, so it favors them; the benchmark does not, and the | |
| model's figures under [Evaluation](#evaluation) are those of the structured-forms mix, the lower of the two there: the | |
| choice inflates nothing on the benchmark. | |
| The code windows (4 % of all) are cut from a corpus of public repositories that the data scripts rebuild; the released | |
| weights were trained with an earlier corpus, so the code windows of a retraining differ. The rest rebuilds bit for bit | |
| ([specialists/pii](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/pii/README.md#status) in | |
| purebyte-train). | |
| Attribution: Ai4Privacy / Ai Suisse SA OpenPII-1M (CC BY 4.0) 路 NVIDIA Nemotron-PII (CC BY 4.0) 路 Gretel.ai datasets | |
| (Apache 2.0). The generated windows also draw on: INE, www.ine.es (CC BY 4.0) 路 Tatoeba, tatoeba.org (CC BY 2.0 FR) 路 US | |
| Census Bureau (public domain) 路 Faker (MIT) 路 the credential formats of gitleaks and trufflehog. | |
| ## Responsible use | |
| Redaction lowers the risk of sharing a document; it does not remove it. Review what matters before it leaves your | |
| hands, keep the restore map as private as the original, and do not use the model to find personal data in material you | |
| are not entitled to process. | |
| ## Versions | |
| | Version | Date | Changes | | |
| |---|---|---| | |
| | 1.0.0 | 2026-09-24 | First public release | | |