Instructions to use SlayerLab/NERGAL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/NERGAL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="SlayerLab/NERGAL")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("SlayerLab/NERGAL") model = AutoModelForTokenClassification.from_pretrained("SlayerLab/NERGAL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
1.1.2: email boundary rules
Browse filesRules only; same weights, threshold and API. An address match ends before a glued www.
host or at a capital glued onto a lowercase domain ending; @-mention lists are not
addresses. Card basis moves to amended gold (gold_check_v2), 1.1.1 restated on it.
841-dev union FP 133 -> 80, rules FP 108 -> 24, whole 326/354 unchanged. Full
Dynaword: 165 documents change, no mask dropped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- CHANGELOG.md +22 -0
- README.md +17 -13
- hybrid.json +9 -6
- nergal.py +2 -2
- scrub_pii.py +36 -3
- test_nergal.py +16 -4
CHANGELOG.md
CHANGED
|
@@ -8,6 +8,28 @@ Semver for this island:
|
|
| 8 |
|
| 9 |
Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
|
| 10 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
## 1.1.1
|
| 12 |
|
| 13 |
Same weights, threshold and API. Rules SHA `ad5c51f7…`. The rules file is the frozen and scored candidate `7412261c…` with comment tags removed; its syntax tree is identical.
|
|
|
|
| 8 |
|
| 9 |
Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
|
| 10 |
|
| 11 |
+
## 1.1.2
|
| 12 |
+
|
| 13 |
+
Same weights, threshold and API. Rules SHA `08faef84…`. The rules file is the frozen and scored candidate `08faef84…`.
|
| 14 |
+
|
| 15 |
+
- **Rules (email only):** an address match ends where glued text begins, before a `www.` host glued onto the domain (`jan@firma.plwww.…`) or at a capital glued onto a lowercase domain ending (`jan@firma.plKontakt`). A cut is kept only when what remains is a complete address. A match right after another `@` is dropped when a space comes before the match's own `@` (a list of @-mentions). Without the space it keeps its mask: a handle (`@jan@firma.social`), a label glued on with `@` or glued addresses can hold a real one.
|
| 16 |
+
- **Gold:** one 841-dev email span ran on into a glued URL host, unlike the other seven addresses in its passage and the labelling policy. It was reviewed and trimmed by 10 characters (`gold_check_v2`, `hybrid.json` `eval.gold_amendments`). Numbers from 1.1.2 on use the amended gold; earlier rows use the original. On the amended gold, 1.1.1 has rules FP 108, union FP 133, char P 97.77% and char R 96.49%.
|
| 17 |
+
|
| 18 |
+
841-dev (amended gold): whole 326/354 and residual 23 unchanged; union FP 133 → 80, rules FP 108 → 24; exact-span and per-label numbers unchanged. All changes are in one passage. Its 24 remaining rules FP are capitalised words glued before a local part, which the 1.0.1 prefix trim cuts only partly. Web development set (41,204 passages): 13 passages change, 0 characters added, and a review ruled all dropped runs out of scope. Human test v2 email benchmark half: rules FP 89 → 65, union FP 162 → 138; not independent, since that half surfaced the cases. Cue-less phone test (150 passages): one passage changes, false characters 95 → 78, phones unchanged. Full Dynaword (2,797,400 documents): 165 documents change, almost all EUR-Lex. There are 553 cuts: 521 at a capital and 32 at a glued `www.`. No cut frees an `@`, and 26 cuts free a glued `Tel`/`Fax` label, whose phone number the rules now mask. No mask is dropped. A first candidate also dropped a match after an `@` inside a word (`a@b@firma.pl`). On full Dynaword that lost real addresses in glued lists, so it was removed.
|
| 19 |
+
|
| 20 |
+
Known limits: a lowercase word glued onto the domain (`jan@firma.plkontakt`) stays masked with the address, because the TLD list does not separate the two; text glued before the local part (a `www.` host or a word) is not cut.
|
| 21 |
+
|
| 22 |
+
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 23 |
+
|---|---:|---:|---:|---:|---:|---:|---|
|
| 24 |
+
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
|
| 25 |
+
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
|
| 26 |
+
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
|
| 27 |
+
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
|
| 28 |
+
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
|
| 29 |
+
| 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
|
| 30 |
+
| 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
|
| 31 |
+
| 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules |
|
| 32 |
+
|
| 33 |
## 1.1.1
|
| 34 |
|
| 35 |
Same weights, threshold and API. Rules SHA `ad5c51f7…`. The rules file is the frozen and scored candidate `7412261c…` with comment tags removed; its syntax tree is identical.
|
README.md
CHANGED
|
@@ -13,7 +13,7 @@ tags:
|
|
| 13 |
- hybrid
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# NERGAL 1.1.
|
| 17 |
|
| 18 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 19 |
|
|
@@ -23,8 +23,8 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
|
|
| 23 |
|
| 24 |
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
|
| 25 |
|
| 26 |
-
- **Version:** `1.1.
|
| 27 |
-
- **Ground:** `scrub_pii` regex (SHA256 `
|
| 28 |
- **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
|
| 29 |
- **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
|
| 30 |
- **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
|
|
@@ -62,13 +62,13 @@ NERGAL is not designed to remove:
|
|
| 62 |
|
| 63 |
These are intended exclusions; false positives can still mask some of this content.
|
| 64 |
|
| 65 |
-
### Known gaps in 1.1.
|
| 66 |
|
| 67 |
-
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.1.
|
| 68 |
|
| 69 |
## Versions
|
| 70 |
|
| 71 |
-
841-dev, union at 0.95, 354 gold spans. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table.
|
| 72 |
|
| 73 |
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 74 |
|---|---:|---:|---:|---:|---:|---:|---|
|
|
@@ -77,7 +77,9 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
|
|
| 77 |
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
|
| 78 |
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
|
| 79 |
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
|
| 80 |
-
|
|
|
|
|
|
|
|
| 81 |
|
| 82 |
## Cue-less phone test
|
| 83 |
|
|
@@ -91,12 +93,14 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
|
|
| 91 |
| Passages with false masks /150 | 3 | 4 |
|
| 92 |
| False characters | 83 | 95 |
|
| 93 |
|
| 94 |
-
1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer.
|
| 95 |
|
| 96 |
## 841-dev
|
| 97 |
|
| 98 |
Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
|
| 99 |
|
|
|
|
|
|
|
| 100 |
The files contain real identifiers, so they are not released with the weights.
|
| 101 |
|
| 102 |
## Why XLM-R
|
|
@@ -127,15 +131,15 @@ Same 841-dev split, threshold **0.95**. **Naked** is the transformer alone. **
|
|
| 127 |
|
| 128 |
| System | Mode | Whole /354 | Residual | False chars | Char P | Char R |
|
| 129 |
|---|---|---:|---:|---:|---:|---:|
|
| 130 |
-
| Regex (`scrub_pii`) | rules | 263 | 68 |
|
| 131 |
| GLiNER 2.5-multi zero-shot | naked | 81 | 188 | 970 | 56.98% | 21.22% |
|
| 132 |
| GLiNER 2.5-multi zero-shot | ∪ regex | 273 | 61 | 1,068 | 83.33% | 88.16% |
|
| 133 |
| Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
|
| 134 |
| Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
|
| 135 |
-
| XLM-R epoch 5 | naked | 298 | 42 |
|
| 136 |
-
| **NERGAL 1.1.
|
| 137 |
|
| 138 |
-
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. The two GLiNER ∪ regex rows use the 1.1.0 rules.
|
| 139 |
|
| 140 |
Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
|
| 141 |
|
|
@@ -170,6 +174,6 @@ masked, counts = nergal.scrub(text)
|
|
| 170 |
|
| 171 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 172 |
|
| 173 |
-
`hybrid.json` records version `1.1.
|
| 174 |
|
| 175 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
|
|
|
| 13 |
- hybrid
|
| 14 |
---
|
| 15 |
|
| 16 |
+
# NERGAL 1.1.2
|
| 17 |
|
| 18 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 19 |
|
|
|
|
| 23 |
|
| 24 |
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
|
| 25 |
|
| 26 |
+
- **Version:** `1.1.2` (`hybrid.json`, `CHANGELOG.md`)
|
| 27 |
+
- **Ground:** `scrub_pii` regex (SHA256 `08faef84…`)
|
| 28 |
- **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
|
| 29 |
- **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
|
| 30 |
- **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
|
|
|
|
| 62 |
|
| 63 |
These are intended exclusions; false positives can still mask some of this content.
|
| 64 |
|
| 65 |
+
### Known gaps in 1.1.2
|
| 66 |
|
| 67 |
+
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.1.2 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it.
|
| 68 |
|
| 69 |
## Versions
|
| 70 |
|
| 71 |
+
841-dev, union at 0.95, 354 gold spans. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table. From 1.1.2 on, one 841-dev span is amended by a reviewed gold check (`gold_check_v2`, 10 characters trimmed); 1.1.1 is restated on it, and earlier rows use the original gold.
|
| 72 |
|
| 73 |
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 74 |
|---|---:|---:|---:|---:|---:|---:|---|
|
|
|
|
| 77 |
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
|
| 78 |
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
|
| 79 |
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
|
| 80 |
+
| 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
|
| 81 |
+
| 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
|
| 82 |
+
| **1.1.2** | **326** | **23** | **24** | **80** | **98.65%** | **96.49%** | Email boundary rules |
|
| 83 |
|
| 84 |
## Cue-less phone test
|
| 85 |
|
|
|
|
| 93 |
| Passages with false masks /150 | 3 | 4 |
|
| 94 |
| False characters | 83 | 95 |
|
| 95 |
|
| 96 |
+
1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged.
|
| 97 |
|
| 98 |
## 841-dev
|
| 99 |
|
| 100 |
Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
|
| 101 |
|
| 102 |
+
One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold.
|
| 103 |
+
|
| 104 |
The files contain real identifiers, so they are not released with the weights.
|
| 105 |
|
| 106 |
## Why XLM-R
|
|
|
|
| 131 |
|
| 132 |
| System | Mode | Whole /354 | Residual | False chars | Char P | Char R |
|
| 133 |
|---|---|---:|---:|---:|---:|---:|
|
| 134 |
+
| Regex (`scrub_pii`) | rules | 263 | 68 | 24 | 99.54% | 85.76% |
|
| 135 |
| GLiNER 2.5-multi zero-shot | naked | 81 | 188 | 970 | 56.98% | 21.22% |
|
| 136 |
| GLiNER 2.5-multi zero-shot | ∪ regex | 273 | 61 | 1,068 | 83.33% | 88.16% |
|
| 137 |
| Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
|
| 138 |
| Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
|
| 139 |
+
| XLM-R epoch 5 | naked | 298 | 42 | 67 | 98.78% | 89.36% |
|
| 140 |
+
| **NERGAL 1.1.2** | **∪ regex** | 326 | 23 | 80 | 98.65% | 96.49% |
|
| 141 |
|
| 142 |
+
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
|
| 143 |
|
| 144 |
Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
|
| 145 |
|
|
|
|
| 174 |
|
| 175 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 176 |
|
| 177 |
+
`hybrid.json` records version `1.1.2`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block (with its `gold_amendments`), and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
|
| 178 |
|
| 179 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
hybrid.json
CHANGED
|
@@ -1,21 +1,24 @@
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
-
"version": "1.1.
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
| 8 |
"threshold": 0.95,
|
| 9 |
-
"rules_sha256": "
|
| 10 |
"eval": {
|
| 11 |
"split": "841-dev",
|
|
|
|
|
|
|
|
|
|
| 12 |
"gold_entities": 354,
|
| 13 |
"whole_entities": 326,
|
| 14 |
"residual_passages": 23,
|
| 15 |
-
"union_fp":
|
| 16 |
-
"rules_fp":
|
| 17 |
-
"character_precision": 0.
|
| 18 |
-
"character_recall": 0.
|
| 19 |
"exact_precision": 0.8641,
|
| 20 |
"exact_recall": 0.8983,
|
| 21 |
"exact_f1": 0.8809,
|
|
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
+
"version": "1.1.2",
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
| 8 |
"threshold": 0.95,
|
| 9 |
+
"rules_sha256": "08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb",
|
| 10 |
"eval": {
|
| 11 |
"split": "841-dev",
|
| 12 |
+
"gold_amendments": [
|
| 13 |
+
"gold_check_v2"
|
| 14 |
+
],
|
| 15 |
"gold_entities": 354,
|
| 16 |
"whole_entities": 326,
|
| 17 |
"residual_passages": 23,
|
| 18 |
+
"union_fp": 80,
|
| 19 |
+
"rules_fp": 24,
|
| 20 |
+
"character_precision": 0.9865,
|
| 21 |
+
"character_recall": 0.9649,
|
| 22 |
"exact_precision": 0.8641,
|
| 23 |
"exact_recall": 0.8983,
|
| 24 |
"exact_f1": 0.8809,
|
nergal.py
CHANGED
|
@@ -17,13 +17,13 @@ import scrub_pii
|
|
| 17 |
from scrub_pii import PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
-
VERSION = '1.1.
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
| 24 |
LABELS = ['phone', 'pii']
|
| 25 |
THRESHOLD = 0.95
|
| 26 |
-
RULES_SHA = '
|
| 27 |
DTYPES = ('float32', 'float16')
|
| 28 |
|
| 29 |
|
|
|
|
| 17 |
from scrub_pii import PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
+
VERSION = '1.1.2'
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
| 24 |
LABELS = ['phone', 'pii']
|
| 25 |
THRESHOLD = 0.95
|
| 26 |
+
RULES_SHA = '08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb'
|
| 27 |
DTYPES = ('float32', 'float16')
|
| 28 |
|
| 29 |
|
scrub_pii.py
CHANGED
|
@@ -14,6 +14,10 @@ Wrapped email domains, local parts hyphenated across one line break and small
|
|
| 14 |
extraction gaps around @/hyphens are supported. A title-case alphabetic prefix
|
| 15 |
of 5–11 letters is dropped from the redaction when the remainder is already a
|
| 16 |
complete lowercase-local email and the same passage has at least two such glues.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
Explicit phone extensions and terminal suffix ranges are included; room numbers are not. Labelled numeric
|
| 18 |
PINs (including URL pin= values) map to [PII] before phone detection.
|
| 19 |
Contact/helpline headings cover consecutive descriptive phone-list entries;
|
|
@@ -61,6 +65,7 @@ _EMAIL_WRAP_RE = re.compile(_EMAIL_START + r"[ \t]{0,3}(?:" + _EMAIL_DOMAIN
|
|
| 61 |
# Damaged contact fields can retain only the local part and @.
|
| 62 |
_EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
|
| 63 |
_EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
|
|
|
|
| 64 |
# A layout wrap may follow any hyphen; a blank line still ends the address.
|
| 65 |
_EDELIVERY_WRAP = r"-(?:[ \t]*\r?\n[ \t]*)?"
|
| 66 |
_EDELIVERY_RE = re.compile(r"\bAE:PL" + _EDELIVERY_WRAP + r"\d{5}" + _EDELIVERY_WRAP + r"\d{5}"
|
|
@@ -584,6 +589,30 @@ def _trim_glued_email_prefix(text: str, start: int, end: int) -> int:
|
|
| 584 |
return start
|
| 585 |
|
| 586 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 587 |
def _glued_email_trim_starts(text: str) -> dict[int, int]:
|
| 588 |
"""Map match starts to trimmed starts when a passage has two or more glues.
|
| 589 |
|
|
@@ -593,6 +622,8 @@ def _glued_email_trim_starts(text: str) -> dict[int, int]:
|
|
| 593 |
starts = {}
|
| 594 |
for pattern in (_EMAIL_WRAP_RE, _EMAIL_RE):
|
| 595 |
for m in pattern.finditer(text):
|
|
|
|
|
|
|
| 596 |
new = _trim_glued_email_prefix(text, m.start(), m.end())
|
| 597 |
if new != m.start():
|
| 598 |
starts.setdefault(m.start(), new)
|
|
@@ -610,12 +641,14 @@ def _replace_checked(text: str, pattern: re.Pattern, tag: str, ok,
|
|
| 610 |
return m.group(0)
|
| 611 |
if pattern is _PESEL_RE and re.search(r'[/?&][^\s]*\Z', text[max(0, m.start()-500):m.start()]):
|
| 612 |
return m.group(0) # Bare URL path/query digits are not personal identifiers.
|
| 613 |
-
|
| 614 |
-
|
|
|
|
| 615 |
return m.group(0)
|
| 616 |
start = trims.get(m.start(), m.start())
|
|
|
|
| 617 |
n += 1
|
| 618 |
-
return text[m.start():start] + mask(start,
|
| 619 |
|
| 620 |
return pattern.sub(_sub, text), n
|
| 621 |
|
|
|
|
| 14 |
extraction gaps around @/hyphens are supported. A title-case alphabetic prefix
|
| 15 |
of 5–11 letters is dropped from the redaction when the remainder is already a
|
| 16 |
complete lowercase-local email and the same passage has at least two such glues.
|
| 17 |
+
An email ends before a URL host glued onto its domain ("…plwww.…") and at a
|
| 18 |
+
capital glued onto a lowercase TLD ("…plKontakt"); an all-lowercase glued word
|
| 19 |
+
is kept. A match right after another "@" is not an email when that "@" is
|
| 20 |
+
inside a token or the match has a gap before its own "@" (a mention list).
|
| 21 |
Explicit phone extensions and terminal suffix ranges are included; room numbers are not. Labelled numeric
|
| 22 |
PINs (including URL pin= values) map to [PII] before phone detection.
|
| 23 |
Contact/helpline headings cover consecutive descriptive phone-list entries;
|
|
|
|
| 65 |
# Damaged contact fields can retain only the local part and @.
|
| 66 |
_EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
|
| 67 |
_EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
|
| 68 |
+
_GLUED_WWW_RE = re.compile(r"(?<=[^\W\d_]{2})www\.", re.I)
|
| 69 |
# A layout wrap may follow any hyphen; a blank line still ends the address.
|
| 70 |
_EDELIVERY_WRAP = r"-(?:[ \t]*\r?\n[ \t]*)?"
|
| 71 |
_EDELIVERY_RE = re.compile(r"\bAE:PL" + _EDELIVERY_WRAP + r"\d{5}" + _EDELIVERY_WRAP + r"\d{5}"
|
|
|
|
| 589 |
return start
|
| 590 |
|
| 591 |
|
| 592 |
+
def _email_length(frag: str) -> int:
|
| 593 |
+
"""Length of the address in an email match that runs into glued text.
|
| 594 |
+
|
| 595 |
+
Each cut is kept only when what remains is still a complete address. A URL
|
| 596 |
+
host glued onto a domain label ("…plwww.…") ends the address before "www";
|
| 597 |
+
a capital glued onto a lowercase TLD ("…plKontakt") ends it at the capital.
|
| 598 |
+
An all-lowercase glued word ("…plkontakt") is kept; its boundary
|
| 599 |
+
needs a TLD list, and the IANA registry did not separate those cases.
|
| 600 |
+
"""
|
| 601 |
+
www = _GLUED_WWW_RE.search(frag, frag.rindex('@'))
|
| 602 |
+
if www and (_EMAIL_RE.fullmatch(frag[:www.start()]) or _EMAIL_WRAP_RE.fullmatch(frag[:www.start()])):
|
| 603 |
+
frag = frag[:www.start()]
|
| 604 |
+
tld = re.search(r"[^\W\d_]+\Z", frag).group(0)
|
| 605 |
+
cut = next((i for i, c in enumerate(tld) if c.isupper()), 0)
|
| 606 |
+
return len(frag) - len(tld) + cut if cut >= 2 and tld[:cut].islower() else len(frag)
|
| 607 |
+
|
| 608 |
+
|
| 609 |
+
def _after_at(text: str, start: int, frag: str) -> bool:
|
| 610 |
+
"""A match right after another '@' with a gap before its own '@' ("@jan @firma.pl") is a list of mentions, not an
|
| 611 |
+
address. Without a gap it keeps the baseline mask: a handle ("@jan@firma.social"), a label glued on with '@'
|
| 612 |
+
("kontakt@jan7@wp.pl") or glued addresses ("jan@firma.comanna@firma.pl") can hold a real one."""
|
| 613 |
+
return start > 0 and text[start - 1] == '@' and bool(re.search(r"[ \t]@", frag))
|
| 614 |
+
|
| 615 |
+
|
| 616 |
def _glued_email_trim_starts(text: str) -> dict[int, int]:
|
| 617 |
"""Map match starts to trimmed starts when a passage has two or more glues.
|
| 618 |
|
|
|
|
| 622 |
starts = {}
|
| 623 |
for pattern in (_EMAIL_WRAP_RE, _EMAIL_RE):
|
| 624 |
for m in pattern.finditer(text):
|
| 625 |
+
if _after_at(text, m.start(), m.group(0)):
|
| 626 |
+
continue
|
| 627 |
new = _trim_glued_email_prefix(text, m.start(), m.end())
|
| 628 |
if new != m.start():
|
| 629 |
starts.setdefault(m.start(), new)
|
|
|
|
| 641 |
return m.group(0)
|
| 642 |
if pattern is _PESEL_RE and re.search(r'[/?&][^\s]*\Z', text[max(0, m.start()-500):m.start()]):
|
| 643 |
return m.group(0) # Bare URL path/query digits are not personal identifiers.
|
| 644 |
+
email = pattern in (_EMAIL_RE, _EMAIL_WRAP_RE, _EMAIL_FRAGMENT_RE)
|
| 645 |
+
if not ok(m.group(0)) or (email and _after_at(text, m.start(), m.group(0))) or (
|
| 646 |
+
not email and _is_amount(text, m.start(), m.end())):
|
| 647 |
return m.group(0)
|
| 648 |
start = trims.get(m.start(), m.start())
|
| 649 |
+
end = m.start() + _email_length(m.group(0)) if pattern in (_EMAIL_RE, _EMAIL_WRAP_RE) else m.end()
|
| 650 |
n += 1
|
| 651 |
+
return text[m.start():start] + mask(start, end, tag) + text[end:m.end()]
|
| 652 |
|
| 653 |
return pattern.sub(_sub, text), n
|
| 654 |
|
test_nergal.py
CHANGED
|
@@ -5,7 +5,7 @@ import unittest
|
|
| 5 |
from pathlib import Path
|
| 6 |
|
| 7 |
HERE = Path(__file__).resolve().parent
|
| 8 |
-
RULES_SHA = '
|
| 9 |
|
| 10 |
|
| 11 |
class NergalTests(unittest.TestCase):
|
|
@@ -13,10 +13,10 @@ class NergalTests(unittest.TestCase):
|
|
| 13 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 14 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 15 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 16 |
-
self.assertEqual(VERSION, '1.1.
|
| 17 |
self.assertEqual(card['version'], VERSION)
|
| 18 |
-
self.assertEqual(card['eval']['union_fp'],
|
| 19 |
-
self.assertEqual(card['eval']['rules_fp'],
|
| 20 |
self.assertEqual(GAPS, card['gaps'])
|
| 21 |
self.assertEqual(GAP_IDS, card['gap_ids'])
|
| 22 |
self.assertEqual(THRESHOLD, card['threshold'])
|
|
@@ -75,6 +75,18 @@ class NergalTests(unittest.TestCase):
|
|
| 75 |
with self.subTest(text=text):
|
| 76 |
self.assertEqual([text[s['start']:s['end']] for s in rules(text) if s['label'] == 'phone'], masked)
|
| 77 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
def test_union_keeps_regex_and_adds_model_spans(self):
|
| 79 |
from nergal import apply_union, scrub_spans
|
| 80 |
text = 'Ring 000000000 then extra.'
|
|
|
|
| 5 |
from pathlib import Path
|
| 6 |
|
| 7 |
HERE = Path(__file__).resolve().parent
|
| 8 |
+
RULES_SHA = '08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb'
|
| 9 |
|
| 10 |
|
| 11 |
class NergalTests(unittest.TestCase):
|
|
|
|
| 13 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 14 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 15 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 16 |
+
self.assertEqual(VERSION, '1.1.2')
|
| 17 |
self.assertEqual(card['version'], VERSION)
|
| 18 |
+
self.assertEqual(card['eval']['union_fp'], 80)
|
| 19 |
+
self.assertEqual(card['eval']['rules_fp'], 24)
|
| 20 |
self.assertEqual(GAPS, card['gaps'])
|
| 21 |
self.assertEqual(GAP_IDS, card['gap_ids'])
|
| 22 |
self.assertEqual(THRESHOLD, card['threshold'])
|
|
|
|
| 75 |
with self.subTest(text=text):
|
| 76 |
self.assertEqual([text[s['start']:s['end']] for s in rules(text) if s['label'] == 'phone'], masked)
|
| 77 |
|
| 78 |
+
def test_email_ends_at_glued_text_and_mention_lists_are_not_addresses(self):
|
| 79 |
+
from nergal import rules
|
| 80 |
+
for text, masked in (('kontakt@firma.plKontakt', ['kontakt@firma.pl']), # capital glued onto the TLD
|
| 81 |
+
('jan@firma.plwww.firma.pl', ['jan@firma.pl']), # URL host glued onto the TLD
|
| 82 |
+
('jan@firma.plkontakt', ['jan@firma.plkontakt']), # all-lowercase glue: known limit
|
| 83 |
+
('BIURO@FIRMA.PL', ['BIURO@FIRMA.PL']),
|
| 84 |
+
('kontakt@jan7@wp.pl', ['jan7@wp.pl']), # word glued on with '@'
|
| 85 |
+
('Dzięki @kasia @firma.pl @tomek', []), # list of mentions
|
| 86 |
+
('Obserwuj @jan@firma.social', ['jan@firma.social'])): # handle keeps the mask
|
| 87 |
+
with self.subTest(text=text):
|
| 88 |
+
self.assertEqual([text[s['start']:s['end']] for s in rules(text)], masked)
|
| 89 |
+
|
| 90 |
def test_union_keeps_regex_and_adds_model_spans(self):
|
| 91 |
from nergal import apply_union, scrub_spans
|
| 92 |
text = 'Ring 000000000 then extra.'
|