ppuzio Claude Opus 5.5 commited on
Commit
bbe063e
·
1 Parent(s): bbd1d02

Card: 315-span comparison, condensed

Browse files

Comparison table rescored on the 315-span 841-dev gold with every model
through the 1.2.0 recipe (spans at 0.95, phone policy v3 cut, 1.2.0 rules);
stored spans, no inference. The fine-tuned GLiNER union is now the more
precise union (39 vs 80 false characters) and is stated as such. Removed
the /354 version table (per-version numbers stay in CHANGELOG.md), the
selection figures and the historical seed table. Weights, code and
hybrid.json unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

README.md CHANGED
@@ -23,11 +23,10 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
23
 
24
  Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked.
25
 
26
- - **Version:** `1.2.0` (`hybrid.json`, `CHANGELOG.md`)
27
- - **Ground:** `scrub_pii` regex (SHA256 `b238d5b8…`)
28
- - **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
29
- - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
30
- - **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
31
 
32
  ## What NERGAL detects — and what it does not
33
 
@@ -64,105 +63,48 @@ These are intended exclusions; false positives can still mask some of this conte
64
 
65
  ### Known gaps in 1.2.0
66
 
67
- Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.2.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (`12 / 345 / 678`) is left as text.
68
 
69
  ## Versions
70
 
71
- 841-dev, union at 0.95. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table. From 1.1.2 on, one 841-dev span is amended by a reviewed gold check (`gold_check_v2`, 10 characters trimmed); 1.1.1 is restated on it, and earlier rows use the original gold. From 1.2.0 the gold is restated to phone policy v3 (315 spans, see 841-dev below); 1.1.2 is restated on it, and the two tables are not comparable.
72
 
73
  | Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
74
  |---|---:|---:|---:|---:|---:|---:|---|
75
- | 1.1.2, restated gold | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Restated on phone policy v3 and `gold_check_v4` |
76
  | **1.2.0** | **303** | **9** | **24** | **80** | **98.59%** | **97.83%** | Phone policy v3 at mask time, rules and model spans |
77
 
78
- | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
79
- |---|---:|---:|---:|---:|---:|---:|---|
80
- | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
81
- | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
82
- | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
83
- | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
84
- | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
85
- | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
86
- | 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
87
- | 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules |
88
 
89
  ## Cue-less phone test
90
 
91
- 841-dev holds few phones without a label, so it barely moves for 1.1.1. This targeted test does: 150 web passages selected for phones with no contact label, labelled and frozen before any prediction, and scored once (2026-09-29). Union at 0.95, same model spans for both.
92
-
93
- | | 1.1.0 | **1.1.1** |
94
- |---|---:|---:|
95
- | Phones wholly masked /111 | 37 | **66** |
96
- | All values wholly masked /178 | 98 | **127** |
97
- | Positive passages fully covered /70 (95% CI) | 24 (0.23–0.47) | **37 (0.41–0.65)** |
98
- | Passages with false masks /150 | 3 | 4 |
99
- | False characters | 83 | 95 |
100
-
101
- 1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged. Half of this set is now a benchmark on gold restated to phone policy v3 (70 passages): 1.2.0 and 1.1.2 both wholly mask 62 of 83 values there, with 33 false characters each.
102
 
103
  ## 841-dev
104
 
105
- Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
106
-
107
- One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold. From 1.2.0 the gold is restated to phone policy v3 (55 passages; short and emergency numbers dropped, joined numbers split) and 5 more passages are amended by a reviewed check (`gold_check_v4`): **315 spans** (130 phone, 185 other PII). The tables below that are marked /354 use the earlier gold.
108
-
109
- The files contain real identifiers, so they are not released with the weights.
110
-
111
- ## Why XLM-R
112
-
113
- GLiNER, HerBERT-large, and XLM-R-large were trained on the same split and unioned with the same regex. Plot: diagnostic threshold 0.50; selection used the full threshold grid. GLiNER covers more at 0.50 and then dumps precision. XLM-R is the architecture we kept.
114
-
115
- ![Primary three-model curves](figures/primary-three-model-curves.png)
116
-
117
- ## Why epoch 5
118
-
119
- Fresh XLM-R, seven epochs. 133 epoch/threshold combinations. Epoch 5 at 0.95 was the only point that both beat the historical GLiNER∪regex incumbent on coverage and introduced zero new false-mask characters. Epoch 7 covers more PII (334/354) but adds 15 new false characters.
120
-
121
- ![Seven-epoch XLM-R curves](figures/xlmr-seven-epoch-curves.png)
122
-
123
- | Epoch | Covered @ 0.95 /354 | Residual passages | False characters | New false vs incumbent |
124
- |---|---:|---:|---:|---:|
125
- | 1 | 272 | 59 | 175 | 42 |
126
- | 2 | 291 | 49 | 141 | 8 |
127
- | 3 | 317 | 30 | 140 | 7 |
128
- | 4 | 320 | 28 | 143 | 10 |
129
- | **5** | **323** | **25** | **133** | **0** |
130
- | 6 | 328 | 20 | 147 | 14 |
131
- | 7 | 334 | 16 | 148 | 15 |
132
 
133
  ## Compared with other systems
134
 
135
- Same 841-dev split, threshold **0.95**. **Naked** is the transformer alone. **∪ regex** is that model unioned with the current rules. Character scores are gold vs masked characters.
136
 
137
- | System | Mode | Whole /354 | Residual | False chars | Char P | Char R |
138
  |---|---|---:|---:|---:|---:|---:|
139
- | Regex (`scrub_pii`) | rules | 263 | 68 | 24 | 99.54% | 85.76% |
140
- | GLiNER 2.5-multi zero-shot | naked | 81 | 188 | 970 | 56.98% | 21.22% |
141
- | GLiNER 2.5-multi zero-shot | ∪ regex | 273 | 61 | 1,068 | 83.33% | 88.16% |
142
- | Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
143
- | Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
144
- | XLM-R epoch 5 | naked | 298 | 42 | 67 | 98.78% | 89.36% |
145
- | **NERGAL 1.1.2** | **∪ regex** | 326 | 23 | 80 | 98.65% | 96.49% |
146
-
147
- Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL 1.1.2 is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. 1.2.0 is scored only on the restated gold (/315): 124/130 phone, 179/185 other PII, union FP 80, exact-span precision 91.05%, recall 93.65%, F1 92.33%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
148
-
149
- Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
150
-
151
- ## Extra seeds
152
-
153
- Historical seed-comparison results, before the 1.0.2 parser fix.
154
 
155
- | Seed | Whole /354 | False chars | New false vs historical union |
156
- |---|---:|---:|---:|
157
- | 202609160 (selected weights) | 323 | 123 | 0 |
158
- | 202609161 | 322 | 134 | 1 |
159
- | 202609162 | 316 | 151 | 18 |
160
 
161
- Seed 160 is the published weights. 161 and 162 were confirmation runs of the same recipe.
162
 
163
  ## Load
164
 
165
- This repo is the PII island: `scrub_pii.py` plus `nergal.py`. `pipeline("token-classification")` will not match. Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[Telefon]` / `[PII]`. Text that already holds `[PII]` / `[Telefon]` is scrubbed as usual, but an identifier with a placeholder inside it or right before it can be missed.
166
 
167
  ```python
168
  from pathlib import Path
@@ -179,6 +121,6 @@ masked, counts = nergal.scrub(text)
179
 
180
  For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
181
 
182
- `hybrid.json` records version `1.2.0`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block (with its `gold_amendments` and `gold_restatements`), and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
183
 
184
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
 
23
 
24
  Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked.
25
 
26
+ - **Ground:** `scrub_pii.py` rules
27
+ - **Additive labels:** XLM-RoBERTa-large token classifier (epoch 5 of 7), BIO tags `phone` / `pii`, threshold 0.95
28
+ - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes
29
+ - **Changes:** `CHANGELOG.md`
 
30
 
31
  ## What NERGAL detects — and what it does not
32
 
 
63
 
64
  ### Known gaps in 1.2.0
65
 
66
+ Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.2.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (`12 / 345 / 678`) is left as text. An identifier with a `[PII]` / `[Telefon]` placeholder inside it or right before it can be missed.
67
 
68
  ## Versions
69
 
70
+ 841-dev (below), union at 0.95. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. 1.2.0 restated the gold to phone policy v3 and rescored 1.1.2 on it. The restatement and the mask-time cut apply the same policy, so the gain measures agreement with it, not an independent test.
71
 
72
  | Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
73
  |---|---:|---:|---:|---:|---:|---:|---|
74
+ | 1.1.2 | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Email boundary rules |
75
  | **1.2.0** | **303** | **9** | **24** | **80** | **98.59%** | **97.83%** | Phone policy v3 at mask time, rules and model spans |
76
 
77
+ 1.0.0–1.1.2 were scored on the earlier 354-span gold (whole 323 → 326, union FP 133 → 80); per-version numbers are in `CHANGELOG.md`.
 
 
 
 
 
 
 
 
 
78
 
79
  ## Cue-less phone test
80
 
81
+ 841-dev holds few phones without a contact label. This targeted set does: 150 web passages selected for them, labelled and frozen before any prediction. 1.1.1 added the grouped national forms and raised phones wholly masked from 37 to 66 of 111, losing none. Its benchmark half, restated to phone policy v3 (70 passages): 1.2.0 wholly masks 62 of 83 values, with 33 false characters. The set is enriched by selection and has a single reviewer, so it says nothing about how common such phones are.
 
 
 
 
 
 
 
 
 
 
82
 
83
  ## 841-dev
84
 
85
+ 841 development passages, 215 with gold PII, **315 spans** (130 phone, 185 other PII). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review. Since 1.2.0 the gold follows phone policy v3 (short and emergency numbers dropped, joined numbers split) and includes reviewed corrections; before, it held 354 spans. The files contain real identifiers, so they are not released with the weights.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
86
 
87
  ## Compared with other systems
88
 
89
+ Same split and gold. Every model runs through the 1.2.0 recipe: its spans at 0.95, phone spans cut by phone policy v3, unioned with the 1.2.0 rules. **Naked** is the model alone. **Residual** is passages with any gold character left unmasked. Character scores are gold vs masked characters.
90
 
91
+ | System | Mode | Whole /315 | Residual | False chars | Char P | Char R |
92
  |---|---|---:|---:|---:|---:|---:|
93
+ | Regex (`scrub_pii`) | rules | 268 | 28 | 24 | 99.53% | 89.85% |
94
+ | GLiNER 2.5-multi zero-shot | naked | 84 | 151 | 451 | 72.96% | 21.33% |
95
+ | GLiNER 2.5-multi zero-shot | ∪ regex | 275 | 25 | 475 | 91.68% | 91.73% |
96
+ | Fine-tuned GLiNER (previous incumbent) | naked | 263 | 28 | 15 | 99.70% | 88.80% |
97
+ | Fine-tuned GLiNER (previous incumbent) | ∪ regex | 299 | 12 | 39 | 99.30% | 97.18% |
98
+ | XLM-R epoch 5 | naked | 275 | 28 | 67 | 98.72% | 90.25% |
99
+ | **NERGAL 1.2.0** | **∪ regex** | **303** | **9** | **80** | **98.59%** | **97.83%** |
 
 
 
 
 
 
 
 
100
 
101
+ Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive, especially on non-phone PII (7/185 whole, naked). The fine-tuned GLiNER that NERGAL replaced is the more precise union (39 false characters vs 80) and the stronger phone model (128/130 whole, naked, vs 115), but it masks 4 fewer values whole and leaves 12 residual passages vs 9; XLM-R is the stronger model on other PII (160/185 vs 135). NERGAL 1.2.0 wholly masks 124/130 phones and 179/185 other PII; exact-span precision 91.05%, recall 93.65%, F1 92.33%.
 
 
 
 
102
 
103
+ XLM-R and epoch 5 were chosen in September 2026, on the earlier gold and before the phone-policy cut. Fine-tuned GLiNER, HerBERT-large and XLM-R-large were trained on the same split; only XLM-R passed the content-preservation gate, and epoch 5 at 0.95 (seed `202609160`) was the only one of 133 epoch/threshold points that covered more than the fine-tuned GLiNER union of that time while adding no false-mask characters it did not already make. Under the 1.2.0 rules, gold and phone cut, that no longer holds: the GLiNER union makes fewer false masks (table above).
104
 
105
  ## Load
106
 
107
+ Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[Telefon]` / `[PII]`.
108
 
109
  ```python
110
  from pathlib import Path
 
121
 
122
  For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
123
 
124
+ `hybrid.json` records the version, threshold, 841-dev `eval` block and weight hashes. `test_nergal.py` is synthetic (no corpus text): `python -m unittest test_nergal`.
125
 
126
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
figures/primary-three-model-curves.png DELETED

Git LFS Details

  • SHA256: c84b632606f398fc49b2289cc13d22c5ec6481cb6fcf147fdf0a315e46053deb
  • Pointer size: 131 Bytes
  • Size of remote file: 177 kB
figures/xlmr-seven-epoch-curves.png DELETED

Git LFS Details

  • SHA256: 3ccce9f3bdf6ee636f08b9fa3277a4582370e3ed193ec3bbb4ad1792819b7eb6
  • Pointer size: 131 Bytes
  • Size of remote file: 210 kB