1.0.3: registry and label-note rules, placeholder and card fixes

#1
by ppuzio - opened
Files changed (6) hide show
  1. CHANGELOG.md +19 -0
  2. README.md +8 -7
  3. hybrid.json +4 -3
  4. nergal.py +11 -13
  5. scrub_pii.py +23 -4
  6. test_nergal.py +9 -2
CHANGELOG.md CHANGED
@@ -8,6 +8,25 @@ Semver for this island:
8
 
9
  Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11
  ## 1.0.2
12
 
13
  Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path.
 
8
 
9
  Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
10
 
11
+ ## 1.0.3
12
+
13
+ Same weights, threshold, and API. Rules SHA `f32d5c54…`.
14
+
15
+ - **Rules:** `NIP`/`REGON` labels with footnote marks or a short gloss (`NIP*:`, `NIP (Wykonawcy):`); e-Delivery addresses broken across a line; `DUNS`, `BDO` and `RPWDL` numbers right after their own label, at a fixed length.
16
+ - **Wrapper:** existing `[PII]`/`[Telefon]` placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document.
17
+ - **Card:** `weights_sha256` was the hash of the source training checkpoint, not of `model.safetensors`. It is now `source_checkpoint_sha256`, and `model_safetensors_sha256` holds the published file's hash. Readers of `weights_sha256` must switch keys.
18
+
19
+ 841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed.
20
+
21
+ Known limit: an identifier with a placeholder inside it or right before it (`NIP [PII] …`, `8503[PII]…`) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules.
22
+
23
+ | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
24
+ |---|---:|---:|---:|---:|---:|---:|---|
25
+ | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
26
+ | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
27
+ | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
28
+ | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
29
+
30
  ## 1.0.2
31
 
32
  Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path.
README.md CHANGED
@@ -13,7 +13,7 @@ tags:
13
  - hybrid
14
  ---
15
 
16
- # NERGAL 1.0.2
17
 
18
  **Named Entity Recognition with Grounded Additive Labels**
19
 
@@ -23,8 +23,8 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
23
 
24
  Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
25
 
26
- - **Version:** `1.0.2` (`hybrid.json`, `CHANGELOG.md`)
27
- - **Ground:** `scrub_pii` regex (SHA256 `3016ae5b…`)
28
  - **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
29
  - **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
30
 
@@ -36,7 +36,8 @@ Python rules do the identifiers they can prove. A transformer NER head adds phon
36
  |---|---:|---:|---:|---:|---:|---:|---|
37
  | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
38
  | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
39
- | **1.0.2** | **324** | **24** | **98** | **123** | **97.93%** | **96.12%** | Labelled country-area phone fix |
 
40
 
41
  ## 841-dev
42
 
@@ -78,7 +79,7 @@ Same 841-dev split, threshold **0.95**. **Naked** is the transformer alone. **
78
  | Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
79
  | Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
80
  | XLM-R epoch 5 | naked | 298 | 42 | 57 | 98.96% | 89.38% |
81
- | **NERGAL 1.0.2** | **∪ regex** | 324 | 24 | 123 | 97.93% | 96.12% |
82
 
83
  Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 145/169 phone, 179/185 other PII. Exact-span precision 86.34%, recall 89.27%, F1 87.78%.
84
 
@@ -98,7 +99,7 @@ Seed 160 is the published weights. 161 and 162 were confirmation runs of the sam
98
 
99
  ## Load
100
 
101
- This repo is the PII island: `scrub_pii.py` plus `nergal.py`. `pipeline("token-classification")` will not match. Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[Telefon]` / `[PII]`.
102
 
103
  ```python
104
  from pathlib import Path
@@ -113,6 +114,6 @@ nergal = Nergal.from_pretrained(root, local_files_only=True)
113
  masked, counts = nergal.scrub(text)
114
  ```
115
 
116
- `hybrid.json` records version `1.0.2`, threshold 0.95, gap ids `250002` / `250003`, and the 841-dev `eval` block. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
117
 
118
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
 
13
  - hybrid
14
  ---
15
 
16
+ # NERGAL 1.0.3
17
 
18
  **Named Entity Recognition with Grounded Additive Labels**
19
 
 
23
 
24
  Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
25
 
26
+ - **Version:** `1.0.3` (`hybrid.json`, `CHANGELOG.md`)
27
+ - **Ground:** `scrub_pii` regex (SHA256 `f32d5c54…`)
28
  - **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
29
  - **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
30
 
 
36
  |---|---:|---:|---:|---:|---:|---:|---|
37
  | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
38
  | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
39
+ | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
40
+ | **1.0.3** | **324** | **24** | **98** | **123** | **97.93%** | **96.12%** | Label-note, e-Delivery and registry rules; placeholder and card fixes |
41
 
42
  ## 841-dev
43
 
 
79
  | Historical GLiNER email12 | naked | 249 | 70 | 10 | 99.81% | 85.39% |
80
  | Historical GLiNER email12 | ∪ regex | 290 | 50 | 108 | 98.13% | 93.76% |
81
  | XLM-R epoch 5 | naked | 298 | 42 | 57 | 98.96% | 89.38% |
82
+ | **NERGAL 1.0.3** | **∪ regex** | 324 | 24 | 123 | 97.93% | 96.12% |
83
 
84
  Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 145/169 phone, 179/185 other PII. Exact-span precision 86.34%, recall 89.27%, F1 87.78%.
85
 
 
99
 
100
  ## Load
101
 
102
+ This repo is the PII island: `scrub_pii.py` plus `nergal.py`. `pipeline("token-classification")` will not match. Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[Telefon]` / `[PII]`. Text that already holds `[PII]` / `[Telefon]` is scrubbed as usual, but an identifier with a placeholder inside it or right before it can be missed.
103
 
104
  ```python
105
  from pathlib import Path
 
114
  masked, counts = nergal.scrub(text)
115
  ```
116
 
117
+ `hybrid.json` records version `1.0.3`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block, and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
118
 
119
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
hybrid.json CHANGED
@@ -1,12 +1,12 @@
1
  {
2
  "full_name": "Named Entity Recognition with Grounded Additive Labels",
3
  "hub_id": "SlayerLab/NERGAL",
4
- "version": "1.0.2",
5
  "mode": "rules_union",
6
  "epoch": 5,
7
  "seed": 202609160,
8
  "threshold": 0.95,
9
- "rules_sha256": "3016ae5bd403ff997458f9dd74bad8c6ed1388eb83dadc1b31cdb182f9ed607f",
10
  "eval": {
11
  "split": "841-dev",
12
  "gold_entities": 354,
@@ -24,7 +24,8 @@
24
  "pii_whole": 179,
25
  "pii_gold": 185
26
  },
27
- "weights_sha256": "063ee5f9782c1b4d99e838be5a836328718e1372e213d5b87098b371b5c162af",
 
28
  "gaps": [
29
  "[PII_SPACE]",
30
  "[PII_BREAK]"
 
1
  {
2
  "full_name": "Named Entity Recognition with Grounded Additive Labels",
3
  "hub_id": "SlayerLab/NERGAL",
4
+ "version": "1.0.3",
5
  "mode": "rules_union",
6
  "epoch": 5,
7
  "seed": 202609160,
8
  "threshold": 0.95,
9
+ "rules_sha256": "f32d5c5452fc47178e109d4bc248a0d8234ea6e59e8cf79407f4eb8451581d67",
10
  "eval": {
11
  "split": "841-dev",
12
  "gold_entities": 354,
 
24
  "pii_whole": 179,
25
  "pii_gold": 185
26
  },
27
+ "source_checkpoint_sha256": "063ee5f9782c1b4d99e838be5a836328718e1372e213d5b87098b371b5c162af",
28
+ "model_safetensors_sha256": "1d42c34459e90cd25db44f92bb31fb2be1ef3a1b1e34c599162e1895057c08f6",
29
  "gaps": [
30
  "[PII_SPACE]",
31
  "[PII_BREAK]"
nergal.py CHANGED
@@ -17,13 +17,13 @@ import scrub_pii
17
  from scrub_pii import PHONE_TAG, PII_TAG
18
 
19
  HUB_ID = 'SlayerLab/NERGAL'
20
- VERSION = '1.0.2'
21
  GAPS = ['[PII_SPACE]', '[PII_BREAK]']
22
  GAP_IDS = [250002, 250003]
23
  BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
24
  LABELS = ['phone', 'pii']
25
  THRESHOLD = 0.95
26
- RULES_SHA = '3016ae5bd403ff997458f9dd74bad8c6ed1388eb83dadc1b31cdb182f9ed607f'
27
 
28
 
29
  def sha(path):
@@ -165,16 +165,19 @@ class Encoding:
165
  return units, windows(units, self.count) if units else []
166
 
167
 
168
- def rules(text):
169
- verify_rules()
170
  result = []
171
- scrub_pii.scrub_pii(text, spans=result)
172
- if '[PII]' in text or '[Telefon]' in text:
173
- return []
174
  return sorted(({k: s[k] for k in ('start', 'end', 'label')} | {'score': 1.0} for s in result),
175
  key=lambda s: s['start'])
176
 
177
 
 
 
 
 
 
178
  def apply_union(text, spans):
179
  labels = [None] * len(text)
180
  for span in spans:
@@ -294,12 +297,7 @@ class Nergal:
294
  return decode_bio(units, (sums / counts).tolist())
295
 
296
  def rule_spans(self, text):
297
- result = []
298
- self._scrub.scrub_pii(text, spans=result)
299
- if '[PII]' in text or '[Telefon]' in text:
300
- return []
301
- return sorted(({k: s[k] for k in ('start', 'end', 'label')} | {'score': 1.0} for s in result),
302
- key=lambda s: s['start'])
303
 
304
  def scrub(self, text):
305
  if not text:
 
17
  from scrub_pii import PHONE_TAG, PII_TAG
18
 
19
  HUB_ID = 'SlayerLab/NERGAL'
20
+ VERSION = '1.0.3'
21
  GAPS = ['[PII_SPACE]', '[PII_BREAK]']
22
  GAP_IDS = [250002, 250003]
23
  BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
24
  LABELS = ['phone', 'pii']
25
  THRESHOLD = 0.95
26
+ RULES_SHA = 'f32d5c5452fc47178e109d4bc248a0d8234ea6e59e8cf79407f4eb8451581d67'
27
 
28
 
29
  def sha(path):
 
165
  return units, windows(units, self.count) if units else []
166
 
167
 
168
+ def spans_from(module, text):
169
+ """Rule spans on the original text. Existing [PII]/[Telefon] placeholders do not switch the rules off."""
170
  result = []
171
+ module.scrub_pii(text, spans=result)
 
 
172
  return sorted(({k: s[k] for k in ('start', 'end', 'label')} | {'score': 1.0} for s in result),
173
  key=lambda s: s['start'])
174
 
175
 
176
+ def rules(text):
177
+ verify_rules()
178
+ return spans_from(scrub_pii, text)
179
+
180
+
181
  def apply_union(text, spans):
182
  labels = [None] * len(text)
183
  for span in spans:
 
297
  return decode_bio(units, (sums / counts).tolist())
298
 
299
  def rule_spans(self, text):
300
+ return spans_from(self._scrub, text)
 
 
 
 
 
301
 
302
  def scrub(self, text):
303
  if not text:
scrub_pii.py CHANGED
@@ -37,7 +37,7 @@ import re
37
  PHONE_TAG = "[Telefon]"
38
  PII_TAG = "[PII]"
39
  COUNTS = ("email", "phone", "pesel", "nip", "regon", "account", "document", "krs", "electronic_address", "pin",
40
- "land_register")
41
 
42
  # Mobile + geographic area codes (2-digit national prefix after trunk 0 / +48).
43
  _PL_PREFIX = {
@@ -60,7 +60,10 @@ _EMAIL_WRAP_RE = re.compile(_EMAIL_START + r"[ \t]{0,3}(?:" + _EMAIL_DOMAIN
60
  # Damaged contact fields can retain only the local part and @.
61
  _EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
62
  _EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
63
- _EDELIVERY_RE = re.compile(r"\bAE:PL-\d{5}-\d{5}-[A-Z0-9]{5}-\d{2}(?![\w-])", re.I)
 
 
 
64
  _EPUAP_RE = re.compile(r"(?<![\w/])/[\w-]+/[\w-]+(?![\w/-])")
65
  _EPUAP_LABEL_RE = re.compile(
66
  r"\b(?:e[ -]?puap|elektroniczna[ \t]+skrzynka[ \t]+podawcza)\b[^\n/\[\]]{0,50}\Z", re.I)
@@ -120,13 +123,27 @@ _LAND_REGISTER_LABEL_RE = re.compile(
120
  _LAND_REGISTER_VALUES = {c: i for i, c in enumerate("0123456789XABCDEFGHIJKLMNOPRSTUWYZ")}
121
  _REGON9_RE = re.compile(r"\b\d{9}\b")
122
  _ID_LABEL_END = r"[ \t]{0,8}(?:(?:nr\.?|numer)[ \t]{0,8})?[:=.\-]?[ \t]{0,8}(?:\r?\n[ \t]{0,8})?\Z"
123
- _NIP_LABEL_RE = re.compile(r"\bNIP\b" + _ID_LABEL_END, re.I)
124
- _REGON_LABEL_RE = re.compile(r"\bREGON\b" + _ID_LABEL_END, re.I)
 
 
 
125
  _KRS_RE = re.compile(r"\b\d{6,10}(?!\w|[ \t-]*\d)")
126
  # Abbreviation or written-out register name, optionally "pod nr/numerem".
127
  _KRS_LABEL_RE = re.compile(
128
  r"(?:\bKRS\b|\bKrajow\w*[ \t]+Rejestr\w*[ \t]+S[aą]dow\w*)"
129
  r"(?:[ \t]{0,8},?[ \t]{0,8}pod[ \t]{1,8}(?:nr\.?|numerem))?" + _ID_LABEL_END, re.I)
 
 
 
 
 
 
 
 
 
 
 
130
  # Separators are short and local: one newline *or* a few punct/spaces.
131
  # Letters and blank lines break the match so body text is not swallowed.
132
  _SEP = r"(?:[ \t./\-()\u2013\u2014]{0,3}|\n)"
@@ -771,6 +788,8 @@ def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, i
771
  counts['electronic_address'] = apply(_replace_checked, _EDELIVERY_RE, PII_TAG, lambda _: True)
772
  counts['electronic_address'] += apply(_replace_epuap)
773
  counts['krs'] = apply(_replace_checked, _KRS_RE, PII_TAG, lambda _: True, _KRS_LABEL_RE)
 
 
774
  counts["document"] = apply(_replace_documents)
775
  counts["land_register"] = apply(_replace_land_registers)
776
  def labelled_ids(value, *, mask):
 
37
  PHONE_TAG = "[Telefon]"
38
  PII_TAG = "[PII]"
39
  COUNTS = ("email", "phone", "pesel", "nip", "regon", "account", "document", "krs", "electronic_address", "pin",
40
+ "land_register", "registry")
41
 
42
  # Mobile + geographic area codes (2-digit national prefix after trunk 0 / +48).
43
  _PL_PREFIX = {
 
60
  # Damaged contact fields can retain only the local part and @.
61
  _EMAIL_FRAGMENT_RE = re.compile(_EMAIL_START + r"(?=[ \t]*\r?$)", re.M)
62
  _EMAIL_LABEL_RE = re.compile(r"\be[ -]?mail[ \t]*:[ \t]*\Z", re.I)
63
+ # A layout wrap may follow any hyphen; a blank line still ends the address.
64
+ _EDELIVERY_WRAP = r"-(?:[ \t]*\r?\n[ \t]*)?"
65
+ _EDELIVERY_RE = re.compile(r"\bAE:PL" + _EDELIVERY_WRAP + r"\d{5}" + _EDELIVERY_WRAP + r"\d{5}"
66
+ + _EDELIVERY_WRAP + r"[A-Z0-9]{5}" + _EDELIVERY_WRAP + r"\d{2}(?![\w-])", re.I)
67
  _EPUAP_RE = re.compile(r"(?<![\w/])/[\w-]+/[\w-]+(?![\w/-])")
68
  _EPUAP_LABEL_RE = re.compile(
69
  r"\b(?:e[ -]?puap|elektroniczna[ \t]+skrzynka[ \t]+podawcza)\b[^\n/\[\]]{0,50}\Z", re.I)
 
123
  _LAND_REGISTER_VALUES = {c: i for i, c in enumerate("0123456789XABCDEFGHIJKLMNOPRSTUWYZ")}
124
  _REGON9_RE = re.compile(r"\b\d{9}\b")
125
  _ID_LABEL_END = r"[ \t]{0,8}(?:(?:nr\.?|numer)[ \t]{0,8})?[:=.\-]?[ \t]{0,8}(?:\r?\n[ \t]{0,8})?\Z"
126
+ # Form labels may carry footnote asterisks and one short digit-free gloss: "NIP (Wykonawcy)**:".
127
+ # Only checksum-gated NIP/REGON take it; KRS has no checksum.
128
+ _ID_LABEL_NOTE = r"(?:[ \t]{0,8}\*{1,3})?(?:[ \t]{0,8}\([^()\d\n]{1,40}\))?(?:[ \t]{0,8}\*{1,3})?"
129
+ _NIP_LABEL_RE = re.compile(r"\bNIP\b" + _ID_LABEL_NOTE + _ID_LABEL_END, re.I)
130
+ _REGON_LABEL_RE = re.compile(r"\bREGON\b" + _ID_LABEL_NOTE + _ID_LABEL_END, re.I)
131
  _KRS_RE = re.compile(r"\b\d{6,10}(?!\w|[ \t-]*\d)")
132
  # Abbreviation or written-out register name, optionally "pod nr/numerem".
133
  _KRS_LABEL_RE = re.compile(
134
  r"(?:\bKRS\b|\bKrajow\w*[ \t]+Rejestr\w*[ \t]+S[aą]dow\w*)"
135
  r"(?:[ \t]{0,8},?[ \t]{0,8}pod[ \t]{1,8}(?:nr\.?|numerem))?" + _ID_LABEL_END, re.I)
136
+ # Registry numbers without a checksum (user, 2026-09-25): masked only right after their own label, one fixed length
137
+ # each. DUNS 9 digits (also 2-3-4 dashed), BDO 9 digits, RPWDL book (księga rejestrowa) 12 digits.
138
+ _REGISTRY_BY = r"(?:[ \t]{0,8},?[ \t]{0,8}pod[ \t]{1,8}(?:nr\.?|numerem))?(?:[ \t]{0,8}[\u2013\u2014])?"
139
+ _REGISTRY = (
140
+ (re.compile(r"\b(?:\d{9}|\d{2}-\d{3}-\d{4})(?!\w|[ \t-]*\d)"),
141
+ re.compile(r"\bD-?U-?N-?S\b(?:[ \t]?®)?(?:[ \t]{1,8}number)?" + _REGISTRY_BY + _ID_LABEL_END, re.I)),
142
+ (re.compile(r"\b\d{9}(?!\w|[ \t-]*\d)"),
143
+ re.compile(r"\bBDO\b" + _REGISTRY_BY + _ID_LABEL_END, re.I)),
144
+ (re.compile(r"\b\d{12}(?!\w|[ \t-]*\d)"),
145
+ re.compile(r"(?:\bRPWDL\b|\bksi[eę]g\w*[ \t]+rejestrow\w*)" + _REGISTRY_BY + _ID_LABEL_END, re.I)),
146
+ )
147
  # Separators are short and local: one newline *or* a few punct/spaces.
148
  # Letters and blank lines break the match so body text is not swallowed.
149
  _SEP = r"(?:[ \t./\-()\u2013\u2014]{0,3}|\n)"
 
788
  counts['electronic_address'] = apply(_replace_checked, _EDELIVERY_RE, PII_TAG, lambda _: True)
789
  counts['electronic_address'] += apply(_replace_epuap)
790
  counts['krs'] = apply(_replace_checked, _KRS_RE, PII_TAG, lambda _: True, _KRS_LABEL_RE)
791
+ counts['registry'] = sum(apply(_replace_checked, number, PII_TAG, lambda _: True, label)
792
+ for number, label in _REGISTRY)
793
  counts["document"] = apply(_replace_documents)
794
  counts["land_register"] = apply(_replace_land_registers)
795
  def labelled_ids(value, *, mask):
test_nergal.py CHANGED
@@ -5,7 +5,7 @@ import unittest
5
  from pathlib import Path
6
 
7
  HERE = Path(__file__).resolve().parent
8
- RULES_SHA = '3016ae5bd403ff997458f9dd74bad8c6ed1388eb83dadc1b31cdb182f9ed607f'
9
 
10
 
11
  class NergalTests(unittest.TestCase):
@@ -13,7 +13,7 @@ class NergalTests(unittest.TestCase):
13
  from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
14
  card = json.loads((HERE / 'hybrid.json').read_text())
15
  self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
16
- self.assertEqual(VERSION, '1.0.2')
17
  self.assertEqual(card['version'], VERSION)
18
  self.assertEqual(card['eval']['union_fp'], 123)
19
  self.assertEqual(card['eval']['rules_fp'], 98)
@@ -35,6 +35,13 @@ class NergalTests(unittest.TestCase):
35
  self.assertEqual(len(first), len(words))
36
  self.assertEqual([encoded.word_ids(0)[i] for i in first], [0, 1, 2])
37
 
 
 
 
 
 
 
 
38
  def test_union_keeps_regex_and_adds_model_spans(self):
39
  from nergal import apply_union, scrub_spans
40
  text = 'Ring 000000000 then extra.'
 
5
  from pathlib import Path
6
 
7
  HERE = Path(__file__).resolve().parent
8
+ RULES_SHA = 'f32d5c5452fc47178e109d4bc248a0d8234ea6e59e8cf79407f4eb8451581d67'
9
 
10
 
11
  class NergalTests(unittest.TestCase):
 
13
  from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
14
  card = json.loads((HERE / 'hybrid.json').read_text())
15
  self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
16
+ self.assertEqual(VERSION, '1.0.3')
17
  self.assertEqual(card['version'], VERSION)
18
  self.assertEqual(card['eval']['union_fp'], 123)
19
  self.assertEqual(card['eval']['rules_fp'], 98)
 
35
  self.assertEqual(len(first), len(words))
36
  self.assertEqual([encoded.word_ids(0)[i] for i in first], [0, 1, 2])
37
 
38
+ def test_existing_placeholders_do_not_switch_the_rules_off(self):
39
+ from nergal import rules
40
+ text = 'Kontakt [Telefon], NIP 1234567802.' # invented, checksum-valid
41
+ [span] = rules(text)
42
+ self.assertEqual(text[span['start']:span['end']], '1234567802')
43
+ self.assertEqual(rules('a [PII] b [Telefon] c'), [])
44
+
45
  def test_union_keeps_regex_and_adds_model_spans(self):
46
  from nergal import apply_union, scrub_spans
47
  text = 'Ring 000000000 then extra.'