Instructions to use SlayerLab/NERGAL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/NERGAL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="SlayerLab/NERGAL")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("SlayerLab/NERGAL") model = AutoModelForTokenClassification.from_pretrained("SlayerLab/NERGAL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
1.0.3: registry and label-note rules, placeholder and card fixes
#1
by ppuzio - opened
- CHANGELOG.md +19 -0
- README.md +8 -7
- hybrid.json +4 -3
- nergal.py +11 -13
- scrub_pii.py +23 -4
- test_nergal.py +9 -2
CHANGELOG.md
CHANGED
|
@@ -8,6 +8,25 @@ Semver for this island:
|
|
| 8 |
|
| 9 |
Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
|
| 10 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
## 1.0.2
|
| 12 |
|
| 13 |
Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path.
|
|
|
|
| 8 |
|
| 9 |
Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
|
| 10 |
|
| 11 |
+
## 1.0.3
|
| 12 |
+
|
| 13 |
+
Same weights, threshold, and API. Rules SHA `f32d5c54…`.
|
| 14 |
+
|
| 15 |
+
- **Rules:** `NIP`/`REGON` labels with footnote marks or a short gloss (`NIP*:`, `NIP (Wykonawcy):`); e-Delivery addresses broken across a line; `DUNS`, `BDO` and `RPWDL` numbers right after their own label, at a fixed length.
|
| 16 |
+
- **Wrapper:** existing `[PII]`/`[Telefon]` placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document.
|
| 17 |
+
- **Card:** `weights_sha256` was the hash of the source training checkpoint, not of `model.safetensors`. It is now `source_checkpoint_sha256`, and `model_safetensors_sha256` holds the published file's hash. Readers of `weights_sha256` must switch keys.
|
| 18 |
+
|
| 19 |
+
841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed.
|
| 20 |
+
|
| 21 |
+
Known limit: an identifier with a placeholder inside it or right before it (`NIP [PII] …`, `8503[PII]…`) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules.
|
| 22 |
+
|
| 23 |
+
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 24 |
+
|---|---:|---:|---:|---:|---:|---:|---|
|
| 25 |
+
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
|
| 26 |
+
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
|
| 27 |
+
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
|
| 28 |
+
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
|
| 29 |
+
|
| 30 |
## 1.0.2
|
| 31 |
|
| 32 |
Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path.
|
README.md
CHANGED
|
@@ -13,7 +13,7 @@ tags:
|
|
| 13 |
- hybrid
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# NERGAL 1.0.
|
| 17 |
|
| 18 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 19 |
|
|
@@ -23,8 +23,8 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
|
|
| 23 |
|
| 24 |
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
|
| 25 |
|
| 26 |
-
- **Version:** `1.0.
|
| 27 |
-
- **Ground:** `scrub_pii` regex (SHA256 `
|
| 28 |
- **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
|
| 29 |
- **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
|
| 30 |
|
|
@@ -36,7 +36,8 @@ Python rules do the identifiers they can prove. A transformer NER head adds phon
|
|
| 36 |
|---|---:|---:|---:|---:|---:|---:|---|
|
| 37 |
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
|
| 38 |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
|
| 39 |
-
|
|
|
|
|
| 40 |
|
| 41 |
## 841-dev
|
| 42 |
|
|
@@ -78,7 +79,7 @@ Same 841-dev split, threshold **0.95**. **Naked** is the transformer alone. **
|
|
| 78 |
| Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
|
| 79 |
| Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
|
| 80 |
| XLM-R epoch 5 | naked | 298 | 42 | 57 | 98.96% | 89.38% |
|
| 81 |
-
| **NERGAL 1.0.
|
| 82 |
|
| 83 |
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 145/169 phone, 179/185 other PII. Exact-span precision 86.34%, recall 89.27%, F1 87.78%.
|
| 84 |
|
|
@@ -98,7 +99,7 @@ Seed 160 is the published weights. 161 and 162 were confirmation runs of the sam
|
|
| 98 |
|
| 99 |
## Load
|
| 100 |
|
| 101 |
-
This repo is the PII island: `scrub_pii.py` plus `nergal.py`. `pipeline("token-classification")` will not match. Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[Telefon]` / `[PII]`.
|
| 102 |
|
| 103 |
```python
|
| 104 |
from pathlib import Path
|
|
@@ -113,6 +114,6 @@ nergal = Nergal.from_pretrained(root, local_files_only=True)
|
|
| 113 |
masked, counts = nergal.scrub(text)
|
| 114 |
```
|
| 115 |
|
| 116 |
-
`hybrid.json` records version `1.0.
|
| 117 |
|
| 118 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
|
|
|
| 13 |
- hybrid
|
| 14 |
---
|
| 15 |
|
| 16 |
+
# NERGAL 1.0.3
|
| 17 |
|
| 18 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 19 |
|
|
|
|
| 23 |
|
| 24 |
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
|
| 25 |
|
| 26 |
+
- **Version:** `1.0.3` (`hybrid.json`, `CHANGELOG.md`)
|
| 27 |
+
- **Ground:** `scrub_pii` regex (SHA256 `f32d5c54…`)
|
| 28 |
- **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
|
| 29 |
- **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
|
| 30 |
|
|
|
|
| 36 |
|---|---:|---:|---:|---:|---:|---:|---|
|
| 37 |
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
|
| 38 |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
|
| 39 |
+
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
|
| 40 |
+
| **1.0.3** | **324** | **24** | **98** | **123** | **97.93%** | **96.12%** | Label-note, e-Delivery and registry rules; placeholder and card fixes |
|
| 41 |
|
| 42 |
## 841-dev
|
| 43 |
|
|
|
|
| 79 |
| Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
|
| 80 |
| Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
|
| 81 |
| XLM-R epoch 5 | naked | 298 | 42 | 57 | 98.96% | 89.38% |
|
| 82 |
+
| **NERGAL 1.0.3** | **∪ regex** | 324 | 24 | 123 | 97.93% | 96.12% |
|
| 83 |
|
| 84 |
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 145/169 phone, 179/185 other PII. Exact-span precision 86.34%, recall 89.27%, F1 87.78%.
|
| 85 |
|
|
|
|
| 99 |
|
| 100 |
## Load
|
| 101 |
|
| 102 |
+
This repo is the PII island: `scrub_pii.py` plus `nergal.py`. `pipeline("token-classification")` will not match. Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[Telefon]` / `[PII]`. Text that already holds `[PII]` / `[Telefon]` is scrubbed as usual, but an identifier with a placeholder inside it or right before it can be missed.
|
| 103 |
|
| 104 |
```python
|
| 105 |
from pathlib import Path
|
|
|
|
| 114 |
masked, counts = nergal.scrub(text)
|
| 115 |
```
|
| 116 |
|
| 117 |
+
`hybrid.json` records version `1.0.3`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block, and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
|
| 118 |
|
| 119 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
hybrid.json
CHANGED
|
@@ -1,12 +1,12 @@
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
-
"version": "1.0.
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
| 8 |
"threshold": 0.95,
|
| 9 |
-
"rules_sha256": "
|
| 10 |
"eval": {
|
| 11 |
"split": "841-dev",
|
| 12 |
"gold_entities": 354,
|
|
@@ -24,7 +24,8 @@
|
|
| 24 |
"pii_whole": 179,
|
| 25 |
"pii_gold": 185
|
| 26 |
},
|
| 27 |
-
"
|
|
|
|
| 28 |
"gaps": [
|
| 29 |
"[PII_SPACE]",
|
| 30 |
"[PII_BREAK]"
|
|
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
+
"version": "1.0.3",
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
| 8 |
"threshold": 0.95,
|
| 9 |
+
"rules_sha256": "f32d5c5452fc47178e109d4bc248a0d8234ea6e59e8cf79407f4eb8451581d67",
|
| 10 |
"eval": {
|
| 11 |
"split": "841-dev",
|
| 12 |
"gold_entities": 354,
|
|
|
|
| 24 |
"pii_whole": 179,
|
| 25 |
"pii_gold": 185
|
| 26 |
},
|
| 27 |
+
"source_checkpoint_sha256": "063ee5f9782c1b4d99e838be5a836328718e1372e213d5b87098b371b5c162af",
|
| 28 |
+
"model_safetensors_sha256": "1d42c34459e90cd25db44f92bb31fb2be1ef3a1b1e34c599162e1895057c08f6",
|
| 29 |
"gaps": [
|
| 30 |
"[PII_SPACE]",
|
| 31 |
"[PII_BREAK]"
|
nergal.py
CHANGED
|
@@ -17,13 +17,13 @@ import scrub_pii
|
|
| 17 |
from scrub_pii import PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
-
VERSION = '1.0.
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
| 24 |
LABELS = ['phone', 'pii']
|
| 25 |
THRESHOLD = 0.95
|
| 26 |
-
RULES_SHA = '
|
| 27 |
|
| 28 |
|
| 29 |
def sha(path):
|
|
@@ -165,16 +165,19 @@ class Encoding:
|
|
| 165 |
return units, windows(units, self.count) if units else []
|
| 166 |
|
| 167 |
|
| 168 |
-
def
|
| 169 |
-
|
| 170 |
result = []
|
| 171 |
-
|
| 172 |
-
if '[PII]' in text or '[Telefon]' in text:
|
| 173 |
-
return []
|
| 174 |
return sorted(({k: s[k] for k in ('start', 'end', 'label')} | {'score': 1.0} for s in result),
|
| 175 |
key=lambda s: s['start'])
|
| 176 |
|
| 177 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 178 |
def apply_union(text, spans):
|
| 179 |
labels = [None] * len(text)
|
| 180 |
for span in spans:
|
|
@@ -294,12 +297,7 @@ class Nergal:
|
|
| 294 |
return decode_bio(units, (sums / counts).tolist())
|
| 295 |
|
| 296 |
def rule_spans(self, text):
|
| 297 |
-
|
| 298 |
-
self._scrub.scrub_pii(text, spans=result)
|
| 299 |
-
if '[PII]' in text or '[Telefon]' in text:
|
| 300 |
-
return []
|
| 301 |
-
return sorted(({k: s[k] for k in ('start', 'end', 'label')} | {'score': 1.0} for s in result),
|
| 302 |
-
key=lambda s: s['start'])
|
| 303 |
|
| 304 |
def scrub(self, text):
|
| 305 |
if not text:
|
|
|
|
| 17 |
from scrub_pii import PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
+
VERSION = '1.0.3'
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
| 24 |
LABELS = ['phone', 'pii']
|
| 25 |
THRESHOLD = 0.95
|
| 26 |
+
RULES_SHA = 'f32d5c5452fc47178e109d4bc248a0d8234ea6e59e8cf79407f4eb8451581d67'
|
| 27 |
|
| 28 |
|
| 29 |
def sha(path):
|
|
|
|
| 165 |
return units, windows(units, self.count) if units else []
|
| 166 |
|
| 167 |
|
| 168 |
+
def spans_from(module, text):
|
| 169 |
+
"""Rule spans on the original text. Existing [PII]/[Telefon] placeholders do not switch the rules off."""
|
| 170 |
result = []
|
| 171 |
+
module.scrub_pii(text, spans=result)
|
|
|
|
|
|
|
| 172 |
return sorted(({k: s[k] for k in ('start', 'end', 'label')} | {'score': 1.0} for s in result),
|
| 173 |
key=lambda s: s['start'])
|
| 174 |
|
| 175 |
|
| 176 |
+
def rules(text):
|
| 177 |
+
verify_rules()
|
| 178 |
+
return spans_from(scrub_pii, text)
|
| 179 |
+
|
| 180 |
+
|
| 181 |
def apply_union(text, spans):
|
| 182 |
labels = [None] * len(text)
|
| 183 |
for span in spans:
|
|
|
|
| 297 |
return decode_bio(units, (sums / counts).tolist())
|
| 298 |
|
| 299 |
def rule_spans(self, text):
|
| 300 |
+
return spans_from(self._scrub, text)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 301 |
|
| 302 |
def scrub(self, text):
|
| 303 |
if not text:
|
scrub_pii.py
CHANGED
|
@@ -37,7 +37,7 @@ import re
|
|
| 37 |
PHONE_TAG = "[Telefon]"
|
| 38 |
PII_TAG = "[PII]"
|
| 39 |
COUNTS = ("email", "phone", "pesel", "nip", "regon", "account", "document", "krs", "electronic_address", "pin",
|
| 40 |
-
"land_register")
|
| 41 |
|
| 42 |
# Mobile + geographic area codes (2-digit national prefix after trunk 0 / +48).
|
| 43 |
_PL_PREFIX = {
|
|
@@ -60,7 +60,10 @@ _EMAIL_WRAP_RE = re.compile(_EMAIL_START + r"[ \t]{0,3}(?:" + _EMAIL_DOMAIN
|
|
| 60 |
# Damaged contact fields can retain only the local part and @.
|
| 61 |
_EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
|
| 62 |
_EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
|
| 63 |
-
|
|
|
|
|
|
|
|
|
|
| 64 |
_EPUAP_RE = re.compile(r"(?<![\w/])/[\w-]+/[\w-]+(?![\w/-])")
|
| 65 |
_EPUAP_LABEL_RE = re.compile(
|
| 66 |
r"\b(?:e[ -]?puap|elektroniczna[ \t]+skrzynka[ \t]+podawcza)\b[^\n/\[\]]{0,50}\Z", re.I)
|
|
@@ -120,13 +123,27 @@ _LAND_REGISTER_LABEL_RE = re.compile(
|
|
| 120 |
_LAND_REGISTER_VALUES = {c: i for i, c in enumerate("0123456789XABCDEFGHIJKLMNOPRSTUWYZ")}
|
| 121 |
_REGON9_RE = re.compile(r"\b\d{9}\b")
|
| 122 |
_ID_LABEL_END = r"[ \t]{0,8}(?:(?:nr\.?|numer)[ \t]{0,8})?[:=.\-]?[ \t]{0,8}(?:\r?\n[ \t]{0,8})?\Z"
|
| 123 |
-
|
| 124 |
-
|
|
|
|
|
|
|
|
|
|
| 125 |
_KRS_RE = re.compile(r"\b\d{6,10}(?!\w|[ \t-]*\d)")
|
| 126 |
# Abbreviation or written-out register name, optionally "pod nr/numerem".
|
| 127 |
_KRS_LABEL_RE = re.compile(
|
| 128 |
r"(?:\bKRS\b|\bKrajow\w*[ \t]+Rejestr\w*[ \t]+S[aą]dow\w*)"
|
| 129 |
r"(?:[ \t]{0,8},?[ \t]{0,8}pod[ \t]{1,8}(?:nr\.?|numerem))?" + _ID_LABEL_END, re.I)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 130 |
# Separators are short and local: one newline *or* a few punct/spaces.
|
| 131 |
# Letters and blank lines break the match so body text is not swallowed.
|
| 132 |
_SEP = r"(?:[ \t./\-()\u2013\u2014]{0,3}|\n)"
|
|
@@ -771,6 +788,8 @@ def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, i
|
|
| 771 |
counts['electronic_address'] = apply(_replace_checked, _EDELIVERY_RE, PII_TAG, lambda _: True)
|
| 772 |
counts['electronic_address'] += apply(_replace_epuap)
|
| 773 |
counts['krs'] = apply(_replace_checked, _KRS_RE, PII_TAG, lambda _: True, _KRS_LABEL_RE)
|
|
|
|
|
|
|
| 774 |
counts["document"] = apply(_replace_documents)
|
| 775 |
counts["land_register"] = apply(_replace_land_registers)
|
| 776 |
def labelled_ids(value, *, mask):
|
|
|
|
| 37 |
PHONE_TAG = "[Telefon]"
|
| 38 |
PII_TAG = "[PII]"
|
| 39 |
COUNTS = ("email", "phone", "pesel", "nip", "regon", "account", "document", "krs", "electronic_address", "pin",
|
| 40 |
+
"land_register", "registry")
|
| 41 |
|
| 42 |
# Mobile + geographic area codes (2-digit national prefix after trunk 0 / +48).
|
| 43 |
_PL_PREFIX = {
|
|
|
|
| 60 |
# Damaged contact fields can retain only the local part and @.
|
| 61 |
_EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
|
| 62 |
_EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
|
| 63 |
+
# A layout wrap may follow any hyphen; a blank line still ends the address.
|
| 64 |
+
_EDELIVERY_WRAP = r"-(?:[ \t]*\r?\n[ \t]*)?"
|
| 65 |
+
_EDELIVERY_RE = re.compile(r"\bAE:PL" + _EDELIVERY_WRAP + r"\d{5}" + _EDELIVERY_WRAP + r"\d{5}"
|
| 66 |
+
+ _EDELIVERY_WRAP + r"[A-Z0-9]{5}" + _EDELIVERY_WRAP + r"\d{2}(?![\w-])", re.I)
|
| 67 |
_EPUAP_RE = re.compile(r"(?<![\w/])/[\w-]+/[\w-]+(?![\w/-])")
|
| 68 |
_EPUAP_LABEL_RE = re.compile(
|
| 69 |
r"\b(?:e[ -]?puap|elektroniczna[ \t]+skrzynka[ \t]+podawcza)\b[^\n/\[\]]{0,50}\Z", re.I)
|
|
|
|
| 123 |
_LAND_REGISTER_VALUES = {c: i for i, c in enumerate("0123456789XABCDEFGHIJKLMNOPRSTUWYZ")}
|
| 124 |
_REGON9_RE = re.compile(r"\b\d{9}\b")
|
| 125 |
_ID_LABEL_END = r"[ \t]{0,8}(?:(?:nr\.?|numer)[ \t]{0,8})?[:=.\-]?[ \t]{0,8}(?:\r?\n[ \t]{0,8})?\Z"
|
| 126 |
+
# Form labels may carry footnote asterisks and one short digit-free gloss: "NIP (Wykonawcy)**:".
|
| 127 |
+
# Only checksum-gated NIP/REGON take it; KRS has no checksum.
|
| 128 |
+
_ID_LABEL_NOTE = r"(?:[ \t]{0,8}\*{1,3})?(?:[ \t]{0,8}\([^()\d\n]{1,40}\))?(?:[ \t]{0,8}\*{1,3})?"
|
| 129 |
+
_NIP_LABEL_RE = re.compile(r"\bNIP\b" + _ID_LABEL_NOTE + _ID_LABEL_END, re.I)
|
| 130 |
+
_REGON_LABEL_RE = re.compile(r"\bREGON\b" + _ID_LABEL_NOTE + _ID_LABEL_END, re.I)
|
| 131 |
_KRS_RE = re.compile(r"\b\d{6,10}(?!\w|[ \t-]*\d)")
|
| 132 |
# Abbreviation or written-out register name, optionally "pod nr/numerem".
|
| 133 |
_KRS_LABEL_RE = re.compile(
|
| 134 |
r"(?:\bKRS\b|\bKrajow\w*[ \t]+Rejestr\w*[ \t]+S[aą]dow\w*)"
|
| 135 |
r"(?:[ \t]{0,8},?[ \t]{0,8}pod[ \t]{1,8}(?:nr\.?|numerem))?" + _ID_LABEL_END, re.I)
|
| 136 |
+
# Registry numbers without a checksum (user, 2026-09-25): masked only right after their own label, one fixed length
|
| 137 |
+
# each. DUNS 9 digits (also 2-3-4 dashed), BDO 9 digits, RPWDL book (księga rejestrowa) 12 digits.
|
| 138 |
+
_REGISTRY_BY = r"(?:[ \t]{0,8},?[ \t]{0,8}pod[ \t]{1,8}(?:nr\.?|numerem))?(?:[ \t]{0,8}[\u2013\u2014])?"
|
| 139 |
+
_REGISTRY = (
|
| 140 |
+
(re.compile(r"\b(?:\d{9}|\d{2}-\d{3}-\d{4})(?!\w|[ \t-]*\d)"),
|
| 141 |
+
re.compile(r"\bD-?U-?N-?S\b(?:[ \t]?®)?(?:[ \t]{1,8}number)?" + _REGISTRY_BY + _ID_LABEL_END, re.I)),
|
| 142 |
+
(re.compile(r"\b\d{9}(?!\w|[ \t-]*\d)"),
|
| 143 |
+
re.compile(r"\bBDO\b" + _REGISTRY_BY + _ID_LABEL_END, re.I)),
|
| 144 |
+
(re.compile(r"\b\d{12}(?!\w|[ \t-]*\d)"),
|
| 145 |
+
re.compile(r"(?:\bRPWDL\b|\bksi[eę]g\w*[ \t]+rejestrow\w*)" + _REGISTRY_BY + _ID_LABEL_END, re.I)),
|
| 146 |
+
)
|
| 147 |
# Separators are short and local: one newline *or* a few punct/spaces.
|
| 148 |
# Letters and blank lines break the match so body text is not swallowed.
|
| 149 |
_SEP = r"(?:[ \t./\-()\u2013\u2014]{0,3}|\n)"
|
|
|
|
| 788 |
counts['electronic_address'] = apply(_replace_checked, _EDELIVERY_RE, PII_TAG, lambda _: True)
|
| 789 |
counts['electronic_address'] += apply(_replace_epuap)
|
| 790 |
counts['krs'] = apply(_replace_checked, _KRS_RE, PII_TAG, lambda _: True, _KRS_LABEL_RE)
|
| 791 |
+
counts['registry'] = sum(apply(_replace_checked, number, PII_TAG, lambda _: True, label)
|
| 792 |
+
for number, label in _REGISTRY)
|
| 793 |
counts["document"] = apply(_replace_documents)
|
| 794 |
counts["land_register"] = apply(_replace_land_registers)
|
| 795 |
def labelled_ids(value, *, mask):
|
test_nergal.py
CHANGED
|
@@ -5,7 +5,7 @@ import unittest
|
|
| 5 |
from pathlib import Path
|
| 6 |
|
| 7 |
HERE = Path(__file__).resolve().parent
|
| 8 |
-
RULES_SHA = '
|
| 9 |
|
| 10 |
|
| 11 |
class NergalTests(unittest.TestCase):
|
|
@@ -13,7 +13,7 @@ class NergalTests(unittest.TestCase):
|
|
| 13 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 14 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 15 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 16 |
-
self.assertEqual(VERSION, '1.0.
|
| 17 |
self.assertEqual(card['version'], VERSION)
|
| 18 |
self.assertEqual(card['eval']['union_fp'], 123)
|
| 19 |
self.assertEqual(card['eval']['rules_fp'], 98)
|
|
@@ -35,6 +35,13 @@ class NergalTests(unittest.TestCase):
|
|
| 35 |
self.assertEqual(len(first), len(words))
|
| 36 |
self.assertEqual([encoded.word_ids(0)[i] for i in first], [0, 1, 2])
|
| 37 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
def test_union_keeps_regex_and_adds_model_spans(self):
|
| 39 |
from nergal import apply_union, scrub_spans
|
| 40 |
text = 'Ring 000000000 then extra.'
|
|
|
|
| 5 |
from pathlib import Path
|
| 6 |
|
| 7 |
HERE = Path(__file__).resolve().parent
|
| 8 |
+
RULES_SHA = 'f32d5c5452fc47178e109d4bc248a0d8234ea6e59e8cf79407f4eb8451581d67'
|
| 9 |
|
| 10 |
|
| 11 |
class NergalTests(unittest.TestCase):
|
|
|
|
| 13 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 14 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 15 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 16 |
+
self.assertEqual(VERSION, '1.0.3')
|
| 17 |
self.assertEqual(card['version'], VERSION)
|
| 18 |
self.assertEqual(card['eval']['union_fp'], 123)
|
| 19 |
self.assertEqual(card['eval']['rules_fp'], 98)
|
|
|
|
| 35 |
self.assertEqual(len(first), len(words))
|
| 36 |
self.assertEqual([encoded.word_ids(0)[i] for i in first], [0, 1, 2])
|
| 37 |
|
| 38 |
+
def test_existing_placeholders_do_not_switch_the_rules_off(self):
|
| 39 |
+
from nergal import rules
|
| 40 |
+
text = 'Kontakt [Telefon], NIP 1234567802.' # invented, checksum-valid
|
| 41 |
+
[span] = rules(text)
|
| 42 |
+
self.assertEqual(text[span['start']:span['end']], '1234567802')
|
| 43 |
+
self.assertEqual(rules('a [PII] b [Telefon] c'), [])
|
| 44 |
+
|
| 45 |
def test_union_keeps_regex_and_adds_model_spans(self):
|
| 46 |
from nergal import apply_union, scrub_spans
|
| 47 |
text = 'Ring 000000000 then extra.'
|