Instructions to use SlayerLab/NERGAL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/NERGAL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="SlayerLab/NERGAL")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("SlayerLab/NERGAL") model = AutoModelForTokenClassification.from_pretrained("SlayerLab/NERGAL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Card: 315-span comparison, condensed
Browse filesComparison table rescored on the 315-span 841-dev gold with every model
through the 1.2.0 recipe (spans at 0.95, phone policy v3 cut, 1.2.0 rules);
stored spans, no inference. The fine-tuned GLiNER union is now the more
precise union (39 vs 80 false characters) and is stated as such. Removed
the /354 version table (per-version numbers stay in CHANGELOG.md), the
selection figures and the historical seed table. Weights, code and
hybrid.json unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- README.md +23 -81
- figures/primary-three-model-curves.png +0 -3
- figures/xlmr-seven-epoch-curves.png +0 -3
README.md
CHANGED
|
@@ -23,11 +23,10 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
|
|
| 23 |
|
| 24 |
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked.
|
| 25 |
|
| 26 |
-
- **
|
| 27 |
-
- **
|
| 28 |
-
- **
|
| 29 |
-
- **
|
| 30 |
-
- **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
|
| 31 |
|
| 32 |
## What NERGAL detects — and what it does not
|
| 33 |
|
|
@@ -64,105 +63,48 @@ These are intended exclusions; false positives can still mask some of this conte
|
|
| 64 |
|
| 65 |
### Known gaps in 1.2.0
|
| 66 |
|
| 67 |
-
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.2.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (`12 / 345 / 678`) is left as text.
|
| 68 |
|
| 69 |
## Versions
|
| 70 |
|
| 71 |
-
841-dev, union at 0.95. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major.
|
| 72 |
|
| 73 |
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 74 |
|---|---:|---:|---:|---:|---:|---:|---|
|
| 75 |
-
| 1.1.2
|
| 76 |
| **1.2.0** | **303** | **9** | **24** | **80** | **98.59%** | **97.83%** | Phone policy v3 at mask time, rules and model spans |
|
| 77 |
|
| 78 |
-
|
| 79 |
-
|---|---:|---:|---:|---:|---:|---:|---|
|
| 80 |
-
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
|
| 81 |
-
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
|
| 82 |
-
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
|
| 83 |
-
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
|
| 84 |
-
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
|
| 85 |
-
| 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
|
| 86 |
-
| 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
|
| 87 |
-
| 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules |
|
| 88 |
|
| 89 |
## Cue-less phone test
|
| 90 |
|
| 91 |
-
841-dev holds few phones without a
|
| 92 |
-
|
| 93 |
-
| | 1.1.0 | **1.1.1** |
|
| 94 |
-
|---|---:|---:|
|
| 95 |
-
| Phones wholly masked /111 | 37 | **66** |
|
| 96 |
-
| All values wholly masked /178 | 98 | **127** |
|
| 97 |
-
| Positive passages fully covered /70 (95% CI) | 24 (0.23–0.47) | **37 (0.41–0.65)** |
|
| 98 |
-
| Passages with false masks /150 | 3 | 4 |
|
| 99 |
-
| False characters | 83 | 95 |
|
| 100 |
-
|
| 101 |
-
1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged. Half of this set is now a benchmark on gold restated to phone policy v3 (70 passages): 1.2.0 and 1.1.2 both wholly mask 62 of 83 values there, with 33 false characters each.
|
| 102 |
|
| 103 |
## 841-dev
|
| 104 |
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold. From 1.2.0 the gold is restated to phone policy v3 (55 passages; short and emergency numbers dropped, joined numbers split) and 5 more passages are amended by a reviewed check (`gold_check_v4`): **315 spans** (130 phone, 185 other PII). The tables below that are marked /354 use the earlier gold.
|
| 108 |
-
|
| 109 |
-
The files contain real identifiers, so they are not released with the weights.
|
| 110 |
-
|
| 111 |
-
## Why XLM-R
|
| 112 |
-
|
| 113 |
-
GLiNER, HerBERT-large, and XLM-R-large were trained on the same split and unioned with the same regex. Plot: diagnostic threshold 0.50; selection used the full threshold grid. GLiNER covers more at 0.50 and then dumps precision. XLM-R is the architecture we kept.
|
| 114 |
-
|
| 115 |
-

|
| 116 |
-
|
| 117 |
-
## Why epoch 5
|
| 118 |
-
|
| 119 |
-
Fresh XLM-R, seven epochs. 133 epoch/threshold combinations. Epoch 5 at 0.95 was the only point that both beat the historical GLiNER∪regex incumbent on coverage and introduced zero new false-mask characters. Epoch 7 covers more PII (334/354) but adds 15 new false characters.
|
| 120 |
-
|
| 121 |
-

|
| 122 |
-
|
| 123 |
-
| Epoch | Covered @ 0.95 /354 | Residual passages | False characters | New false vs incumbent |
|
| 124 |
-
|---|---:|---:|---:|---:|
|
| 125 |
-
| 1 | 272 | 59 | 175 | 42 |
|
| 126 |
-
| 2 | 291 | 49 | 141 | 8 |
|
| 127 |
-
| 3 | 317 | 30 | 140 | 7 |
|
| 128 |
-
| 4 | 320 | 28 | 143 | 10 |
|
| 129 |
-
| **5** | **323** | **25** | **133** | **0** |
|
| 130 |
-
| 6 | 328 | 20 | 147 | 14 |
|
| 131 |
-
| 7 | 334 | 16 | 148 | 15 |
|
| 132 |
|
| 133 |
## Compared with other systems
|
| 134 |
|
| 135 |
-
Same
|
| 136 |
|
| 137 |
-
| System | Mode | Whole /
|
| 138 |
|---|---|---:|---:|---:|---:|---:|
|
| 139 |
-
| Regex (`scrub_pii`) | rules |
|
| 140 |
-
| GLiNER 2.5-multi zero-shot | naked |
|
| 141 |
-
| GLiNER 2.5-multi zero-shot | ∪ regex |
|
| 142 |
-
|
|
| 143 |
-
|
|
| 144 |
-
| XLM-R epoch 5 | naked |
|
| 145 |
-
| **NERGAL 1.
|
| 146 |
-
|
| 147 |
-
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL 1.1.2 is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. 1.2.0 is scored only on the restated gold (/315): 124/130 phone, 179/185 other PII, union FP 80, exact-span precision 91.05%, recall 93.65%, F1 92.33%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
|
| 148 |
-
|
| 149 |
-
Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
|
| 150 |
-
|
| 151 |
-
## Extra seeds
|
| 152 |
-
|
| 153 |
-
Historical seed-comparison results, before the 1.0.2 parser fix.
|
| 154 |
|
| 155 |
-
|
| 156 |
-
|---|---:|---:|---:|
|
| 157 |
-
| 202609160 (selected weights) | 323 | 123 | 0 |
|
| 158 |
-
| 202609161 | 322 | 134 | 1 |
|
| 159 |
-
| 202609162 | 316 | 151 | 18 |
|
| 160 |
|
| 161 |
-
|
| 162 |
|
| 163 |
## Load
|
| 164 |
|
| 165 |
-
|
| 166 |
|
| 167 |
```python
|
| 168 |
from pathlib import Path
|
|
@@ -179,6 +121,6 @@ masked, counts = nergal.scrub(text)
|
|
| 179 |
|
| 180 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 181 |
|
| 182 |
-
`hybrid.json` records
|
| 183 |
|
| 184 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
|
|
|
| 23 |
|
| 24 |
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked.
|
| 25 |
|
| 26 |
+
- **Ground:** `scrub_pii.py` rules
|
| 27 |
+
- **Additive labels:** XLM-RoBERTa-large token classifier (epoch 5 of 7), BIO tags `phone` / `pii`, threshold 0.95
|
| 28 |
+
- **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes
|
| 29 |
+
- **Changes:** `CHANGELOG.md`
|
|
|
|
| 30 |
|
| 31 |
## What NERGAL detects — and what it does not
|
| 32 |
|
|
|
|
| 63 |
|
| 64 |
### Known gaps in 1.2.0
|
| 65 |
|
| 66 |
+
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.2.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (`12 / 345 / 678`) is left as text. An identifier with a `[PII]` / `[Telefon]` placeholder inside it or right before it can be missed.
|
| 67 |
|
| 68 |
## Versions
|
| 69 |
|
| 70 |
+
841-dev (below), union at 0.95. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. 1.2.0 restated the gold to phone policy v3 and rescored 1.1.2 on it. The restatement and the mask-time cut apply the same policy, so the gain measures agreement with it, not an independent test.
|
| 71 |
|
| 72 |
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 73 |
|---|---:|---:|---:|---:|---:|---:|---|
|
| 74 |
+
| 1.1.2 | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Email boundary rules |
|
| 75 |
| **1.2.0** | **303** | **9** | **24** | **80** | **98.59%** | **97.83%** | Phone policy v3 at mask time, rules and model spans |
|
| 76 |
|
| 77 |
+
1.0.0–1.1.2 were scored on the earlier 354-span gold (whole 323 → 326, union FP 133 → 80); per-version numbers are in `CHANGELOG.md`.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
## Cue-less phone test
|
| 80 |
|
| 81 |
+
841-dev holds few phones without a contact label. This targeted set does: 150 web passages selected for them, labelled and frozen before any prediction. 1.1.1 added the grouped national forms and raised phones wholly masked from 37 to 66 of 111, losing none. Its benchmark half, restated to phone policy v3 (70 passages): 1.2.0 wholly masks 62 of 83 values, with 33 false characters. The set is enriched by selection and has a single reviewer, so it says nothing about how common such phones are.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
## 841-dev
|
| 84 |
|
| 85 |
+
841 development passages, 215 with gold PII, **315 spans** (130 phone, 185 other PII). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review. Since 1.2.0 the gold follows phone policy v3 (short and emergency numbers dropped, joined numbers split) and includes reviewed corrections; before, it held 354 spans. The files contain real identifiers, so they are not released with the weights.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 86 |
|
| 87 |
## Compared with other systems
|
| 88 |
|
| 89 |
+
Same split and gold. Every model runs through the 1.2.0 recipe: its spans at 0.95, phone spans cut by phone policy v3, unioned with the 1.2.0 rules. **Naked** is the model alone. **Residual** is passages with any gold character left unmasked. Character scores are gold vs masked characters.
|
| 90 |
|
| 91 |
+
| System | Mode | Whole /315 | Residual | False chars | Char P | Char R |
|
| 92 |
|---|---|---:|---:|---:|---:|---:|
|
| 93 |
+
| Regex (`scrub_pii`) | rules | 268 | 28 | 24 | 99.53% | 89.85% |
|
| 94 |
+
| GLiNER 2.5-multi zero-shot | naked | 84 | 151 | 451 | 72.96% | 21.33% |
|
| 95 |
+
| GLiNER 2.5-multi zero-shot | ∪ regex | 275 | 25 | 475 | 91.68% | 91.73% |
|
| 96 |
+
| Fine-tuned GLiNER (previous incumbent) | naked | 263 | 28 | 15 | 99.70% | 88.80% |
|
| 97 |
+
| Fine-tuned GLiNER (previous incumbent) | ∪ regex | 299 | 12 | 39 | 99.30% | 97.18% |
|
| 98 |
+
| XLM-R epoch 5 | naked | 275 | 28 | 67 | 98.72% | 90.25% |
|
| 99 |
+
| **NERGAL 1.2.0** | **∪ regex** | **303** | **9** | **80** | **98.59%** | **97.83%** |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 100 |
|
| 101 |
+
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive, especially on non-phone PII (7/185 whole, naked). The fine-tuned GLiNER that NERGAL replaced is the more precise union (39 false characters vs 80) and the stronger phone model (128/130 whole, naked, vs 115), but it masks 4 fewer values whole and leaves 12 residual passages vs 9; XLM-R is the stronger model on other PII (160/185 vs 135). NERGAL 1.2.0 wholly masks 124/130 phones and 179/185 other PII; exact-span precision 91.05%, recall 93.65%, F1 92.33%.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
|
| 103 |
+
XLM-R and epoch 5 were chosen in September 2026, on the earlier gold and before the phone-policy cut. Fine-tuned GLiNER, HerBERT-large and XLM-R-large were trained on the same split; only XLM-R passed the content-preservation gate, and epoch 5 at 0.95 (seed `202609160`) was the only one of 133 epoch/threshold points that covered more than the fine-tuned GLiNER union of that time while adding no false-mask characters it did not already make. Under the 1.2.0 rules, gold and phone cut, that no longer holds: the GLiNER union makes fewer false masks (table above).
|
| 104 |
|
| 105 |
## Load
|
| 106 |
|
| 107 |
+
Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[Telefon]` / `[PII]`.
|
| 108 |
|
| 109 |
```python
|
| 110 |
from pathlib import Path
|
|
|
|
| 121 |
|
| 122 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 123 |
|
| 124 |
+
`hybrid.json` records the version, threshold, 841-dev `eval` block and weight hashes. `test_nergal.py` is synthetic (no corpus text): `python -m unittest test_nergal`.
|
| 125 |
|
| 126 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
figures/primary-three-model-curves.png
DELETED
Git LFS Details
|
figures/xlmr-seven-epoch-curves.png
DELETED
Git LFS Details
|