ppuzio Claude Opus 5.5 commited on
Commit
e0caa8b
·
1 Parent(s): 4464d47

1.1.2: email boundary rules

Browse files

Rules only; same weights, threshold and API. An address match ends before a glued www.
host or at a capital glued onto a lowercase domain ending; @-mention lists are not
addresses. Card basis moves to amended gold (gold_check_v2), 1.1.1 restated on it.
841-dev union FP 133 -> 80, rules FP 108 -> 24, whole 326/354 unchanged. Full
Dynaword: 165 documents change, no mask dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Files changed (6) hide show
  1. CHANGELOG.md +22 -0
  2. README.md +17 -13
  3. hybrid.json +9 -6
  4. nergal.py +2 -2
  5. scrub_pii.py +36 -3
  6. test_nergal.py +16 -4
CHANGELOG.md CHANGED
@@ -8,6 +8,28 @@ Semver for this island:
8
 
9
  Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11
  ## 1.1.1
12
 
13
  Same weights, threshold and API. Rules SHA `ad5c51f7…`. The rules file is the frozen and scored candidate `7412261c…` with comment tags removed; its syntax tree is identical.
 
8
 
9
  Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
10
 
11
+ ## 1.1.2
12
+
13
+ Same weights, threshold and API. Rules SHA `08faef84…`. The rules file is the frozen and scored candidate `08faef84…`.
14
+
15
+ - **Rules (email only):** an address match ends where glued text begins, before a `www.` host glued onto the domain (`jan@firma.plwww.…`) or at a capital glued onto a lowercase domain ending (`jan@firma.plKontakt`). A cut is kept only when what remains is a complete address. A match right after another `@` is dropped when a space comes before the match's own `@` (a list of @-mentions). Without the space it keeps its mask: a handle (`@jan@firma.social`), a label glued on with `@` or glued addresses can hold a real one.
16
+ - **Gold:** one 841-dev email span ran on into a glued URL host, unlike the other seven addresses in its passage and the labelling policy. It was reviewed and trimmed by 10 characters (`gold_check_v2`, `hybrid.json` `eval.gold_amendments`). Numbers from 1.1.2 on use the amended gold; earlier rows use the original. On the amended gold, 1.1.1 has rules FP 108, union FP 133, char P 97.77% and char R 96.49%.
17
+
18
+ 841-dev (amended gold): whole 326/354 and residual 23 unchanged; union FP 133 → 80, rules FP 108 → 24; exact-span and per-label numbers unchanged. All changes are in one passage. Its 24 remaining rules FP are capitalised words glued before a local part, which the 1.0.1 prefix trim cuts only partly. Web development set (41,204 passages): 13 passages change, 0 characters added, and a review ruled all dropped runs out of scope. Human test v2 email benchmark half: rules FP 89 → 65, union FP 162 → 138; not independent, since that half surfaced the cases. Cue-less phone test (150 passages): one passage changes, false characters 95 → 78, phones unchanged. Full Dynaword (2,797,400 documents): 165 documents change, almost all EUR-Lex. There are 553 cuts: 521 at a capital and 32 at a glued `www.`. No cut frees an `@`, and 26 cuts free a glued `Tel`/`Fax` label, whose phone number the rules now mask. No mask is dropped. A first candidate also dropped a match after an `@` inside a word (`a@b@firma.pl`). On full Dynaword that lost real addresses in glued lists, so it was removed.
19
+
20
+ Known limits: a lowercase word glued onto the domain (`jan@firma.plkontakt`) stays masked with the address, because the TLD list does not separate the two; text glued before the local part (a `www.` host or a word) is not cut.
21
+
22
+ | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
23
+ |---|---:|---:|---:|---:|---:|---:|---|
24
+ | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
25
+ | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
26
+ | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
27
+ | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
28
+ | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
29
+ | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
30
+ | 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
31
+ | 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules |
32
+
33
  ## 1.1.1
34
 
35
  Same weights, threshold and API. Rules SHA `ad5c51f7…`. The rules file is the frozen and scored candidate `7412261c…` with comment tags removed; its syntax tree is identical.
README.md CHANGED
@@ -13,7 +13,7 @@ tags:
13
  - hybrid
14
  ---
15
 
16
- # NERGAL 1.1.1
17
 
18
  **Named Entity Recognition with Grounded Additive Labels**
19
 
@@ -23,8 +23,8 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
23
 
24
  Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
25
 
26
- - **Version:** `1.1.1` (`hybrid.json`, `CHANGELOG.md`)
27
- - **Ground:** `scrub_pii` regex (SHA256 `ad5c51f7…`)
28
  - **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
29
  - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
30
  - **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
@@ -62,13 +62,13 @@ NERGAL is not designed to remove:
62
 
63
  These are intended exclusions; false positives can still mask some of this content.
64
 
65
- ### Known gaps in 1.1.1
66
 
67
- Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.1.1 model. Do not rely on it to remove them consistently.
68
 
69
  ## Versions
70
 
71
- 841-dev, union at 0.95, 354 gold spans. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table.
72
 
73
  | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
74
  |---|---:|---:|---:|---:|---:|---:|---|
@@ -77,7 +77,9 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
77
  | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
78
  | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
79
  | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
80
- | **1.1.1** | **326** | **23** | **98** | **123** | **97.94%** | **96.50%** | Grouped national phones without a label |
 
 
81
 
82
  ## Cue-less phone test
83
 
@@ -91,12 +93,14 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
91
  | Passages with false masks /150 | 3 | 4 |
92
  | False characters | 83 | 95 |
93
 
94
- 1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer.
95
 
96
  ## 841-dev
97
 
98
  Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
99
 
 
 
100
  The files contain real identifiers, so they are not released with the weights.
101
 
102
  ## Why XLM-R
@@ -127,15 +131,15 @@ Same 841-dev split, threshold **0.95**. **Naked** is the transformer alone. **
127
 
128
  | System | Mode | Whole /354 | Residual | False chars | Char P | Char R |
129
  |---|---|---:|---:|---:|---:|---:|
130
- | Regex (`scrub_pii`) | rules | 263 | 68 | 98 | 98.15% | 85.78% |
131
  | GLiNER 2.5-multi zero-shot | naked | 81 | 188 | 970 | 56.98% | 21.22% |
132
  | GLiNER 2.5-multi zero-shot | ∪ regex | 273 | 61 | 1,068 | 83.33% | 88.16% |
133
  | Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
134
  | Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
135
- | XLM-R epoch 5 | naked | 298 | 42 | 57 | 98.96% | 89.38% |
136
- | **NERGAL 1.1.1** | **∪ regex** | 326 | 23 | 123 | 97.94% | 96.50% |
137
 
138
- Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. The two GLiNER ∪ regex rows use the 1.1.0 rules.
139
 
140
  Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
141
 
@@ -170,6 +174,6 @@ masked, counts = nergal.scrub(text)
170
 
171
  For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
172
 
173
- `hybrid.json` records version `1.1.1`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block, and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
174
 
175
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
 
13
  - hybrid
14
  ---
15
 
16
+ # NERGAL 1.1.2
17
 
18
  **Named Entity Recognition with Grounded Additive Labels**
19
 
 
23
 
24
  Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
25
 
26
+ - **Version:** `1.1.2` (`hybrid.json`, `CHANGELOG.md`)
27
+ - **Ground:** `scrub_pii` regex (SHA256 `08faef84…`)
28
  - **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
29
  - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
30
  - **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
 
62
 
63
  These are intended exclusions; false positives can still mask some of this content.
64
 
65
+ ### Known gaps in 1.1.2
66
 
67
+ Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.1.2 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it.
68
 
69
  ## Versions
70
 
71
+ 841-dev, union at 0.95, 354 gold spans. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table. From 1.1.2 on, one 841-dev span is amended by a reviewed gold check (`gold_check_v2`, 10 characters trimmed); 1.1.1 is restated on it, and earlier rows use the original gold.
72
 
73
  | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
74
  |---|---:|---:|---:|---:|---:|---:|---|
 
77
  | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
78
  | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
79
  | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
80
+ | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
81
+ | 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
82
+ | **1.1.2** | **326** | **23** | **24** | **80** | **98.65%** | **96.49%** | Email boundary rules |
83
 
84
  ## Cue-less phone test
85
 
 
93
  | Passages with false masks /150 | 3 | 4 |
94
  | False characters | 83 | 95 |
95
 
96
+ 1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged.
97
 
98
  ## 841-dev
99
 
100
  Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
101
 
102
+ One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold.
103
+
104
  The files contain real identifiers, so they are not released with the weights.
105
 
106
  ## Why XLM-R
 
131
 
132
  | System | Mode | Whole /354 | Residual | False chars | Char P | Char R |
133
  |---|---|---:|---:|---:|---:|---:|
134
+ | Regex (`scrub_pii`) | rules | 263 | 68 | 24 | 99.54% | 85.76% |
135
  | GLiNER 2.5-multi zero-shot | naked | 81 | 188 | 970 | 56.98% | 21.22% |
136
  | GLiNER 2.5-multi zero-shot | ∪ regex | 273 | 61 | 1,068 | 83.33% | 88.16% |
137
  | Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
138
  | Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
139
+ | XLM-R epoch 5 | naked | 298 | 42 | 67 | 98.78% | 89.36% |
140
+ | **NERGAL 1.1.2** | **∪ regex** | 326 | 23 | 80 | 98.65% | 96.49% |
141
 
142
+ Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
143
 
144
  Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
145
 
 
174
 
175
  For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
176
 
177
+ `hybrid.json` records version `1.1.2`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block (with its `gold_amendments`), and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
178
 
179
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
hybrid.json CHANGED
@@ -1,21 +1,24 @@
1
  {
2
  "full_name": "Named Entity Recognition with Grounded Additive Labels",
3
  "hub_id": "SlayerLab/NERGAL",
4
- "version": "1.1.1",
5
  "mode": "rules_union",
6
  "epoch": 5,
7
  "seed": 202609160,
8
  "threshold": 0.95,
9
- "rules_sha256": "ad5c51f7c712dd2b7af40a337b703aeec211a57170a87c262bf9bd9386ba1925",
10
  "eval": {
11
  "split": "841-dev",
 
 
 
12
  "gold_entities": 354,
13
  "whole_entities": 326,
14
  "residual_passages": 23,
15
- "union_fp": 123,
16
- "rules_fp": 98,
17
- "character_precision": 0.9794,
18
- "character_recall": 0.965,
19
  "exact_precision": 0.8641,
20
  "exact_recall": 0.8983,
21
  "exact_f1": 0.8809,
 
1
  {
2
  "full_name": "Named Entity Recognition with Grounded Additive Labels",
3
  "hub_id": "SlayerLab/NERGAL",
4
+ "version": "1.1.2",
5
  "mode": "rules_union",
6
  "epoch": 5,
7
  "seed": 202609160,
8
  "threshold": 0.95,
9
+ "rules_sha256": "08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb",
10
  "eval": {
11
  "split": "841-dev",
12
+ "gold_amendments": [
13
+ "gold_check_v2"
14
+ ],
15
  "gold_entities": 354,
16
  "whole_entities": 326,
17
  "residual_passages": 23,
18
+ "union_fp": 80,
19
+ "rules_fp": 24,
20
+ "character_precision": 0.9865,
21
+ "character_recall": 0.9649,
22
  "exact_precision": 0.8641,
23
  "exact_recall": 0.8983,
24
  "exact_f1": 0.8809,
nergal.py CHANGED
@@ -17,13 +17,13 @@ import scrub_pii
17
  from scrub_pii import PHONE_TAG, PII_TAG
18
 
19
  HUB_ID = 'SlayerLab/NERGAL'
20
- VERSION = '1.1.1'
21
  GAPS = ['[PII_SPACE]', '[PII_BREAK]']
22
  GAP_IDS = [250002, 250003]
23
  BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
24
  LABELS = ['phone', 'pii']
25
  THRESHOLD = 0.95
26
- RULES_SHA = 'ad5c51f7c712dd2b7af40a337b703aeec211a57170a87c262bf9bd9386ba1925'
27
  DTYPES = ('float32', 'float16')
28
 
29
 
 
17
  from scrub_pii import PHONE_TAG, PII_TAG
18
 
19
  HUB_ID = 'SlayerLab/NERGAL'
20
+ VERSION = '1.1.2'
21
  GAPS = ['[PII_SPACE]', '[PII_BREAK]']
22
  GAP_IDS = [250002, 250003]
23
  BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
24
  LABELS = ['phone', 'pii']
25
  THRESHOLD = 0.95
26
+ RULES_SHA = '08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb'
27
  DTYPES = ('float32', 'float16')
28
 
29
 
scrub_pii.py CHANGED
@@ -14,6 +14,10 @@ Wrapped email domains, local parts hyphenated across one line break and small
14
  extraction gaps around @/hyphens are supported. A title-case alphabetic prefix
15
  of 5–11 letters is dropped from the redaction when the remainder is already a
16
  complete lowercase-local email and the same passage has at least two such glues.
 
 
 
 
17
  Explicit phone extensions and terminal suffix ranges are included; room numbers are not. Labelled numeric
18
  PINs (including URL pin= values) map to [PII] before phone detection.
19
  Contact/helpline headings cover consecutive descriptive phone-list entries;
@@ -61,6 +65,7 @@ _EMAIL_WRAP_RE = re.compile(_EMAIL_START + r"[ \t]{0,3}(?:" + _EMAIL_DOMAIN
61
  # Damaged contact fields can retain only the local part and @.
62
  _EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
63
  _EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
 
64
  # A layout wrap may follow any hyphen; a blank line still ends the address.
65
  _EDELIVERY_WRAP = r"-(?:[ \t]*\r?\n[ \t]*)?"
66
  _EDELIVERY_RE = re.compile(r"\bAE:PL" + _EDELIVERY_WRAP + r"\d{5}" + _EDELIVERY_WRAP + r"\d{5}"
@@ -584,6 +589,30 @@ def _trim_glued_email_prefix(text: str, start: int, end: int) -> int:
584
  return start
585
 
586
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
587
  def _glued_email_trim_starts(text: str) -> dict[int, int]:
588
  """Map match starts to trimmed starts when a passage has two or more glues.
589
 
@@ -593,6 +622,8 @@ def _glued_email_trim_starts(text: str) -> dict[int, int]:
593
  starts = {}
594
  for pattern in (_EMAIL_WRAP_RE, _EMAIL_RE):
595
  for m in pattern.finditer(text):
 
 
596
  new = _trim_glued_email_prefix(text, m.start(), m.end())
597
  if new != m.start():
598
  starts.setdefault(m.start(), new)
@@ -610,12 +641,14 @@ def _replace_checked(text: str, pattern: re.Pattern, tag: str, ok,
610
  return m.group(0)
611
  if pattern is _PESEL_RE and re.search(r'[/?&][^\s]*\Z', text[max(0, m.start()-500):m.start()]):
612
  return m.group(0) # Bare URL path/query digits are not personal identifiers.
613
- if not ok(m.group(0)) or (pattern not in (_EMAIL_RE, _EMAIL_WRAP_RE, _EMAIL_FRAGMENT_RE)
614
- and _is_amount(text, m.start(), m.end())):
 
615
  return m.group(0)
616
  start = trims.get(m.start(), m.start())
 
617
  n += 1
618
- return text[m.start():start] + mask(start, m.end(), tag)
619
 
620
  return pattern.sub(_sub, text), n
621
 
 
14
  extraction gaps around @/hyphens are supported. A title-case alphabetic prefix
15
  of 5–11 letters is dropped from the redaction when the remainder is already a
16
  complete lowercase-local email and the same passage has at least two such glues.
17
+ An email ends before a URL host glued onto its domain ("…plwww.…") and at a
18
+ capital glued onto a lowercase TLD ("…plKontakt"); an all-lowercase glued word
19
+ is kept. A match right after another "@" is not an email when that "@" is
20
+ inside a token or the match has a gap before its own "@" (a mention list).
21
  Explicit phone extensions and terminal suffix ranges are included; room numbers are not. Labelled numeric
22
  PINs (including URL pin= values) map to [PII] before phone detection.
23
  Contact/helpline headings cover consecutive descriptive phone-list entries;
 
65
  # Damaged contact fields can retain only the local part and @.
66
  _EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
67
  _EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
68
+ _GLUED_WWW_RE = re.compile(r"(?<=[^\W\d_]{2})www\.", re.I)
69
  # A layout wrap may follow any hyphen; a blank line still ends the address.
70
  _EDELIVERY_WRAP = r"-(?:[ \t]*\r?\n[ \t]*)?"
71
  _EDELIVERY_RE = re.compile(r"\bAE:PL" + _EDELIVERY_WRAP + r"\d{5}" + _EDELIVERY_WRAP + r"\d{5}"
 
589
  return start
590
 
591
 
592
+ def _email_length(frag: str) -> int:
593
+ """Length of the address in an email match that runs into glued text.
594
+
595
+ Each cut is kept only when what remains is still a complete address. A URL
596
+ host glued onto a domain label ("…plwww.…") ends the address before "www";
597
+ a capital glued onto a lowercase TLD ("…plKontakt") ends it at the capital.
598
+ An all-lowercase glued word ("…plkontakt") is kept; its boundary
599
+ needs a TLD list, and the IANA registry did not separate those cases.
600
+ """
601
+ www = _GLUED_WWW_RE.search(frag, frag.rindex('@'))
602
+ if www and (_EMAIL_RE.fullmatch(frag[:www.start()]) or _EMAIL_WRAP_RE.fullmatch(frag[:www.start()])):
603
+ frag = frag[:www.start()]
604
+ tld = re.search(r"[^\W\d_]+\Z", frag).group(0)
605
+ cut = next((i for i, c in enumerate(tld) if c.isupper()), 0)
606
+ return len(frag) - len(tld) + cut if cut >= 2 and tld[:cut].islower() else len(frag)
607
+
608
+
609
+ def _after_at(text: str, start: int, frag: str) -> bool:
610
+ """A match right after another '@' with a gap before its own '@' ("@jan @firma.pl") is a list of mentions, not an
611
+ address. Without a gap it keeps the baseline mask: a handle ("@jan@firma.social"), a label glued on with '@'
612
+ ("kontakt@jan7@wp.pl") or glued addresses ("jan@firma.comanna@firma.pl") can hold a real one."""
613
+ return start > 0 and text[start - 1] == '@' and bool(re.search(r"[ \t]@", frag))
614
+
615
+
616
  def _glued_email_trim_starts(text: str) -> dict[int, int]:
617
  """Map match starts to trimmed starts when a passage has two or more glues.
618
 
 
622
  starts = {}
623
  for pattern in (_EMAIL_WRAP_RE, _EMAIL_RE):
624
  for m in pattern.finditer(text):
625
+ if _after_at(text, m.start(), m.group(0)):
626
+ continue
627
  new = _trim_glued_email_prefix(text, m.start(), m.end())
628
  if new != m.start():
629
  starts.setdefault(m.start(), new)
 
641
  return m.group(0)
642
  if pattern is _PESEL_RE and re.search(r'[/?&][^\s]*\Z', text[max(0, m.start()-500):m.start()]):
643
  return m.group(0) # Bare URL path/query digits are not personal identifiers.
644
+ email = pattern in (_EMAIL_RE, _EMAIL_WRAP_RE, _EMAIL_FRAGMENT_RE)
645
+ if not ok(m.group(0)) or (email and _after_at(text, m.start(), m.group(0))) or (
646
+ not email and _is_amount(text, m.start(), m.end())):
647
  return m.group(0)
648
  start = trims.get(m.start(), m.start())
649
+ end = m.start() + _email_length(m.group(0)) if pattern in (_EMAIL_RE, _EMAIL_WRAP_RE) else m.end()
650
  n += 1
651
+ return text[m.start():start] + mask(start, end, tag) + text[end:m.end()]
652
 
653
  return pattern.sub(_sub, text), n
654
 
test_nergal.py CHANGED
@@ -5,7 +5,7 @@ import unittest
5
  from pathlib import Path
6
 
7
  HERE = Path(__file__).resolve().parent
8
- RULES_SHA = 'ad5c51f7c712dd2b7af40a337b703aeec211a57170a87c262bf9bd9386ba1925'
9
 
10
 
11
  class NergalTests(unittest.TestCase):
@@ -13,10 +13,10 @@ class NergalTests(unittest.TestCase):
13
  from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
14
  card = json.loads((HERE / 'hybrid.json').read_text())
15
  self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
16
- self.assertEqual(VERSION, '1.1.1')
17
  self.assertEqual(card['version'], VERSION)
18
- self.assertEqual(card['eval']['union_fp'], 123)
19
- self.assertEqual(card['eval']['rules_fp'], 98)
20
  self.assertEqual(GAPS, card['gaps'])
21
  self.assertEqual(GAP_IDS, card['gap_ids'])
22
  self.assertEqual(THRESHOLD, card['threshold'])
@@ -75,6 +75,18 @@ class NergalTests(unittest.TestCase):
75
  with self.subTest(text=text):
76
  self.assertEqual([text[s['start']:s['end']] for s in rules(text) if s['label'] == 'phone'], masked)
77
 
 
 
 
 
 
 
 
 
 
 
 
 
78
  def test_union_keeps_regex_and_adds_model_spans(self):
79
  from nergal import apply_union, scrub_spans
80
  text = 'Ring 000000000 then extra.'
 
5
  from pathlib import Path
6
 
7
  HERE = Path(__file__).resolve().parent
8
+ RULES_SHA = '08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb'
9
 
10
 
11
  class NergalTests(unittest.TestCase):
 
13
  from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
14
  card = json.loads((HERE / 'hybrid.json').read_text())
15
  self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
16
+ self.assertEqual(VERSION, '1.1.2')
17
  self.assertEqual(card['version'], VERSION)
18
+ self.assertEqual(card['eval']['union_fp'], 80)
19
+ self.assertEqual(card['eval']['rules_fp'], 24)
20
  self.assertEqual(GAPS, card['gaps'])
21
  self.assertEqual(GAP_IDS, card['gap_ids'])
22
  self.assertEqual(THRESHOLD, card['threshold'])
 
75
  with self.subTest(text=text):
76
  self.assertEqual([text[s['start']:s['end']] for s in rules(text) if s['label'] == 'phone'], masked)
77
 
78
+ def test_email_ends_at_glued_text_and_mention_lists_are_not_addresses(self):
79
+ from nergal import rules
80
+ for text, masked in (('kontakt@firma.plKontakt', ['kontakt@firma.pl']), # capital glued onto the TLD
81
+ ('jan@firma.plwww.firma.pl', ['jan@firma.pl']), # URL host glued onto the TLD
82
+ ('jan@firma.plkontakt', ['jan@firma.plkontakt']), # all-lowercase glue: known limit
83
+ ('BIURO@FIRMA.PL', ['BIURO@FIRMA.PL']),
84
+ ('kontakt@jan7@wp.pl', ['jan7@wp.pl']), # word glued on with '@'
85
+ ('Dzięki @kasia @firma.pl @tomek', []), # list of mentions
86
+ ('Obserwuj @jan@firma.social', ['jan@firma.social'])): # handle keeps the mask
87
+ with self.subTest(text=text):
88
+ self.assertEqual([text[s['start']:s['end']] for s in rules(text)], masked)
89
+
90
  def test_union_keeps_regex_and_adds_model_spans(self):
91
  from nergal import apply_union, scrub_spans
92
  text = 'Ring 000000000 then extra.'