NERGAL / CHANGELOG.md
ppuzio's picture
Claude Opus 5.5
2.0.1: a [PERSON] placeholder ends an other-number label
960c515
|
Raw History Blame Contribute Delete
23.7 kB

NERGAL versions

Semver for this island:

  • MAJOR β€” API, threshold, or weight recipe changes
  • MINOR β€” new capability, same API
  • PATCH β€” rules or card fix, same weights and API

Accuracy is 841-dev, union at 0.95. Through 1.1.2 the gold has 354 spans; from 1.2.0 it is restated to phone policy v3 (315 spans), and the two are not comparable. A version that changes those numbers must update hybrid.json eval and the tables below.

2.0.1

Same NERGAL weights, threshold and API. Rules SHA c1b924a8…: one fix for input that already holds [PERSON].

  • A masked name ends an other-number label. A Kod, NIP, REGON, PESEL, KRS or ISBN label up to 20 characters before a number keeps it from being a phone. In 2.0.0 that held across a [PERSON] placeholder, so scrubbing names-on output again could leave a phone that follows a masked name. In 2.0.1 a [PERSON] between the label and the number ends the label's reach.
  • Phone cues still cross [PERSON]. A contact or phone label before a masked name still counts for the number after it. Unlike [PHONE] and [PII], [PERSON] is not a rules boundary.
  • A first scrub is unchanged. The rules read the raw text and name spans are joined after, so scrubbing raw text gives 2.0.0's rule spans, names on or off. The fix shows only when the input already holds [PERSON], such as a second scrub.

841-dev, names off: unchanged from 2.0.0. Release gates, 2.0.1 rules against 2.0.0's: 841-dev plus the invented placeholder cases, 0 of 845 passages changed, scrubbed once or twice; 841-dev scrubbed names-on and then again, 0 of 841 changed, so the fix is exercised only by the invented cases in test_nergal.py; the benchmark replay gives 2.0.0's numbers on all 12 cohorts; every on-disk Dynaword source (37 sources, 4,144,908 documents, full text, placeholder documents included), 0 changed.

Version Whole /315 Residual Rules FP Union FP Char P Char R What changed
2.0.0, names off 303 9 24 80 98.59% 97.83% [PHONE] replaces [Telefon]; opt-in [PERSON] names
2.0.1, names off 303 9 24 80 98.59% 97.83% A [PERSON] placeholder ends an other-number label

2.0.0

Same NERGAL weights and threshold. Rules SHA d1866243…: the placeholder rename only; every rule span is unchanged. Major because the output placeholder changes.

  • Breaking β€” [PHONE] replaces [Telefon]. Every phone mask now reads [PHONE]; [PII] is unchanged. Code that matches [Telefon] in the output must switch. Text scrubbed by 1.x still works as input: [Telefon] stays a rules boundary (a previous redaction, not a fresh Telefon label), like [PHONE] and [PII].
  • Names, opt-in: Nergal.from_pretrained(..., names=True) (CLI --names) loads names/, FastPDN NER β€” Polish PII by ArkadiuszPawlak (CC-BY-4.0, names/NOTICE.md), converted from ONNX to safetensors with no value changed. Its PERSON, PERSON_F and PERSON_L tags become [PERSON]: whole words, joined across spaces, tabs and no-break spaces (not punctuation or line breaks); a word touching a placeholder is skipped. It always runs in float32: dtype sets NERGAL only. Without names=True, from_pretrained skips names/ (β‰ˆ490 MB) and the output equals 1.2.0's except for the rename.
  • Union: each character takes its highest-ranked label, phone > pii > person, and each run becomes one placeholder. A name inside a masked email stays [PII].
  • Counts: scrub() counts gain person (0 with names off). union_placeholder_chars includes [PERSON] placeholders. predict() / predict_many() still return NERGAL spans only; nergal.names.spans_many(texts) returns the person spans.

841-dev, names off: unchanged from 1.2.0 (whole 303/315, residual 9, rules FP 24, union FP 80, char P 98.59%, char R 97.83%). Its gold has no person labels, so names-on is not scored on it. The benchmark replay at 2.0.0 with names off gives 1.2.0's numbers on every one of its 12 cohorts.

Name test (nergal_names_test_v1). Dynaword and Polish mC4 web passages (70/30 in each panel), component-disjoint from every earlier set, labelled under docs/pii-annotation-names.md, frozen before any prediction and scored once. The random panel estimates the corpus; the targeted panel is enriched by selection (correspondence, signatures, honorifics, hard negatives) and says nothing about prevalence. Panels are never pooled. A name is whole when every letter and digit of it is masked ([PII] counts); a component is the independent unit, covered when every name in it is whole. Names-on, float32:

Panel Passages Names Whole Exact Components covered (95% CI) Person false chars Passages with one (95% CI)
Random 300 614 586 (95.44%) 565 98/116 (76.6–90.5%) 1,051 34 (8.0–15.5%)
Targeted 150 292 286 (97.95%) 285 94/99 (88.6–98.3%) 976 42 (21.0–35.9%)

Names off hides no name; its false characters are 1 (random) and 29 (targeted). The set's phones (2 and 7) and other PII (2 and 5) are whole in both modes. For the record, neither shipped: float16 on Apple MPS equals float32 on both panels; FastPDN's INT8 ONNX on CPU hides 589 and 287 names whole with 1,009 and 935 person false characters. One reviewer labelled the set; the second pass is a same-day blind self-review of 100 passages (11 disagreed, adjudication changed 5), not independent agreement. hybrid.json eval_names holds these numbers and the score report's SHA.

Release gates. The name test replaces the planned 200-document spot check: a spot check reviews only what the model masked, so it measures neither recall nor names missed on clean-looking text. Throughput, local: 200 seeded Dynaword documents (1,103,229 characters), Apple M4 Max, MPS, float32, scrub_many, one process, off and on alternating twice. Names off 7,914 and 7,922 chars/s, peak RSS 1,393 and 1,388 MiB; names on 7,444 and 7,597 chars/s (βˆ’5%), 1,679 and 1,682 MiB (+290 MiB). MPS driver memory at the end of a run is 70,261 MiB off and 87,788 MiB on; that is the allocator cache, not a working set. Repeat runs mask identically; names on gives 1,309 [PERSON] placeholders, names off none, and both give 7 [PHONE] and 7 [PII]. CUDA, the shipped release: the v2.0.0 Hub snapshot loaded on an RTX 4090 pod (driver 570.172.08, torch 2.8.0+cu128), the runtime study's 1,000 Dynaword documents (4,215,149 characters), NERGAL float16 and names float32, off and on interleaved, twice. One process 38,654 β†’ 35,108 chars/s (βˆ’9%); 3 processes 74,520 β†’ 66,492 chars/s (βˆ’11%), peak device memory 9,006 β†’ 15,416 MiB; per process, torch peak 1,806 β†’ 2,668 MiB. Repeats and 1 vs 3 processes mask identically; names on changes 666 of the 1,000 documents (6,209 [PERSON]). This host was slower than the runtime study's (names off, 3 processes: 74,520 vs 83,479 chars/s), so compare within one host. A float16 names model changed spans in 11 of the 1,000 documents in the runtime study, so it is not offered.

Known limits: names-on masks every person the policy covers, public officials and historical figures included, which is why it is opt-in. On the random panel 18 of 116 components keep part of a name, and 34 of 300 passages have a false person mask. The names model was fine-tuned partly on LLM-synthetic data that SlayerLab has not audited.

Version Whole /315 Residual Rules FP Union FP Char P Char R What changed
1.2.0 303 9 24 80 98.59% 97.83% Phone policy v3 at mask time, rules and model spans
2.0.0, names off 303 9 24 80 98.59% 97.83% [PHONE] replaces [Telefon]; opt-in [PERSON] names

1.2.0

Same weights, threshold and API. Rules SHA b238d5b8…. The rules file is the frozen and scored candidate 73c185f7… with two comments reworded; its syntax tree is identical. Phone policy v3 (labelling policy amended 2026-10-02) now applies at mask time, to the rules and to the model spans.

  • Phone policy v3: each number of 7+ digits is its own span, and connectors between numbers (/, ,, lub, a spaced dash) stay as text. A short part (an extension wew. 101, an alternative ending /90, a wrapped line) stays inside its number's span. A standalone number under 7 digits is not masked: emergency numbers (112, 997), helplines and service numbers (116 111), short codes. Country and area codes count as digits, and so do keypad letters. A code before a slash or in parentheses belongs to the number after it (032/2345678 β€” 032/2345679 is two numbers), and a + slash chain (+420/55/123456) is one.
  • Rules: the service-number and lone-extension passes are removed. Detected phones are merged into runs and cut by scrub_pii._phone_spans. A written connector joins two detections, but a bare line break joins only two short detections (one wrapped number). A lone slash followed by at least 5 digits joins its two sides into one number (022/123-456).
  • Model spans: at 0.95, each model phone span is cut by the same _phone_spans before the union: one span per full number, and short numbers are dropped. pii spans are unchanged. nergal.model_keep(text, model_spans) returns the spans that are kept. The weights did not change, so predict() still returns short numbers; only scrub() / scrub_many() drop them.
  • Gold: 841-dev is restated to policy v3 mechanically by the policy's own rule (policy_restatement_v2: 55 passages change; checked on a seeded 41-passage human spot-check, which it matches except for one passage a later ruling overrode), and 5 passages are amended by a reviewed gold check (gold_check_v4). Gold spans: 354 β†’ 315 (phone 169 β†’ 130). Both are listed in hybrid.json eval.

841-dev (restated gold), 1.1.2 β†’ 1.2.0: whole 303/315 and residual 9 are unchanged. Rules FP 83 β†’ 24, union FP 333 β†’ 80, passages with false characters 42 β†’ 1. Char P 94.37% β†’ 98.59%, char R 97.83%, exact-span F1 82.28% β†’ 92.33%. Without the model-span cut, rules 1.1.3 alone give union FP 327: the published model masks short numbers. The restatement and the mask-time cut apply the same policy, so these gains measure agreement with it and are not an independent test.

On the benchmark halves of the human test sets (restated the same way), no whole value is lost in any cohort. Union FP: human_test_v2 random 51 β†’ 39, email 136 β†’ 127, registry 4 β†’ 0, bare-ID 3 β†’ 0, e-Delivery 1 β†’ 0; mC4 negatives email 224 β†’ 222; the other cohorts are unchanged. The human_test_v2 half is not independent: one of its rows set the line-break rule and another the 5-digit slash cut.

Dynaword rules delta, human-reviewed (40 passages around changed masks; 38 scored, 2 left uncertain and unscored): phone values whole 57 β†’ 65 of 70, false phone characters 373 β†’ 113, no value lost. Two of these passages set the + chain and spaced-dash rules, so this is development evidence. Full Dynaword rules delta against 1.1.2 (36 sources, 4,094,255 documents, counts only, unreviewed): rule spans 294,819 β†’ 291,651, masked characters βˆ’10,766 (0.21%). In steps: the 1.1.3 v2 candidate changes 2,092 documents (4,661 spans removed, 1,153 added), and v3 then changes 293 (eurlex 287, wikivoyage 6; 335 removed, 675 added). Both steps read the same per-unit input hashes. The 1.1.2 entry's run covered 21 sources on an older checkout, and samorzad_gov_pl has changed upstream since, so document counts across entries are not comparable.

Known limits: two slashes between short parts (12 / 345 / 678) are left as text. The 5 reviewed Dynaword phone values that 1.2.0 does not wholly mask are also missed by 1.1.2. The model-span cut was not reviewed on new text: no new inference was run for this release.

Version Whole /315 Residual Rules FP Union FP Char P Char R What changed
1.1.2, restated gold 303 9 83 333 94.37% 97.83% Restated on phone policy v3 and gold_check_v4
1.2.0 303 9 24 80 98.59% 97.83% Phone policy v3 at mask time, rules and model spans

1.1.2

Same weights, threshold and API. Rules SHA 08faef84…. The rules file is the frozen and scored candidate 08faef84….

  • Rules (email only): an address match ends where glued text begins, before a www. host glued onto the domain (jan@firma.plwww.…) or at a capital glued onto a lowercase domain ending (jan@firma.plKontakt). A cut is kept only when what remains is a complete address. A match right after another @ is dropped when a space comes before the match's own @ (a list of @-mentions). Without the space it keeps its mask: a handle (@jan@firma.social), a label glued on with @ or glued addresses can hold a real one.
  • Gold: one 841-dev email span ran on into a glued URL host, unlike the other seven addresses in its passage and the labelling policy. It was reviewed and trimmed by 10 characters (gold_check_v2, hybrid.json eval.gold_amendments). Numbers from 1.1.2 on use the amended gold; earlier rows use the original. On the amended gold, 1.1.1 has rules FP 108, union FP 133, char P 97.77% and char R 96.49%.

841-dev (amended gold): whole 326/354 and residual 23 unchanged; union FP 133 β†’ 80, rules FP 108 β†’ 24; exact-span and per-label numbers unchanged. All changes are in one passage. Its 24 remaining rules FP are capitalised words glued before a local part, which the 1.0.1 prefix trim cuts only partly. Web development set (41,204 passages): 13 passages change, 0 characters added, and a review ruled all dropped runs out of scope. Human test v2 email benchmark half: rules FP 89 β†’ 65, union FP 162 β†’ 138; not independent, since that half surfaced the cases. Cue-less phone test (150 passages): one passage changes, false characters 95 β†’ 78, phones unchanged. Full Dynaword (2,797,400 documents): 165 documents change, almost all EUR-Lex. There are 553 cuts: 521 at a capital and 32 at a glued www.. No cut frees an @, and 26 cuts free a glued Tel/Fax label, whose phone number the rules now mask. No mask is dropped. A first candidate also dropped a match after an @ inside a word (a@b@firma.pl). On full Dynaword that lost real addresses in glued lists, so it was removed.

Known limits: a lowercase word glued onto the domain (jan@firma.plkontakt) stays masked with the address, because the TLD list does not separate the two; text glued before the local part (a www. host or a word) is not cut.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix
1.0.3 324 24 98 123 97.93% 96.12% Label-note, e-Delivery and registry rules; placeholder and card fixes
1.1.0 324 24 98 123 97.93% 96.12% Batch API (predict_many, scrub_many), opt-in float16
1.1.1 326 23 98 123 97.94% 96.50% Grouped national phones without a label
1.1.1, amended gold 326 23 108 133 97.77% 96.49% Restated on gold_check_v2
1.1.2 326 23 24 80 98.65% 96.49% Email boundary rules

1.1.1

Same weights, threshold and API. Rules SHA ad5c51f7…. The rules file is the frozen and scored candidate 7412261c… with comment tags removed; its syntax tree is identical.

  • Rules: Polish phones in the grouped national forms are masked without a contact label: mobile 601 234 567 and landline 22 123 45 67 (area code optionally in parentheses), optional +48, one separator kind (space or hyphen) throughout. Not taken: plain 9-digit strings (still label-gated), amounts, round counts (500 000 000), decimal figures (601 234 567,89), and numbers after another identifier's label (NIP, REGON, codes). After a bare grouped phone, a list goes on only with complete Polish numbers.

841-dev: 326/354 whole (1.1.0: 324), residual 23 (24), 0 new false characters; phone 147/169 (145). A frozen blind set of 150 web passages selected for unlabelled phones: phones wholly masked 66/111 (1.1.0: 37), all values 127/178 (98), 0 lost, 1 passage with new false characters. Dynaword (60,000 documents): 79 documents change, 247 phone spans added, 0 removed; in a reviewed sample 1 of 23 sampled documents' added spans was not a phone.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix
1.0.3 324 24 98 123 97.93% 96.12% Label-note, e-Delivery and registry rules; placeholder and card fixes
1.1.0 324 24 98 123 97.93% 96.12% Batch API (predict_many, scrub_many), opt-in float16
1.1.1 326 23 98 123 97.94% 96.50% Grouped national phones without a label

1.1.0

Same weights, rules, threshold and default outputs; hybrid.json eval is unchanged. New batch API and opt-in fp16.

  • Nergal.predict_many(texts) / scrub_many(texts): windows from up to 64 texts are sorted by token length and packed into batches of at most 32,768 padded tokens and 128 rows. predict / scrub are now the one-text case of these.
  • dtype='float16' on Nergal(...) / from_pretrained(...) (CLI --dtype): casts the float32 weights at load time. Needs CUDA or MPS. The default stays float32, and model.safetensors is unchanged.
  • Faster window sizing: Encoding.count adds up cached unit pieces instead of re-tokenizing inside the window search. It yields the same windows, because encode() still rejects any unit whose pieces change with context.
  • Default device: from_pretrained now tries CUDA, then MPS, then CPU. 1.0.x used CPU even on CUDA machines.

Throughput on one RTX 4090 (13.88M chars of FineWeb-2, kchar/s): 1.0.3-style per-document batches 11.0 (23.0 with 3 processes); predict_many float32 21.3 (28.1 with 2); predict_many float16 39.4 (79.6 with 3). The float16 path is limited by CPU-side tokenization, so run 2–3 processes per GPU.

Equivalence: on the 1,685 labelled dev rows, float32 predict_many gives 0 span changes at 0.95 against the cached model spans of the published weights (841-dev max score change 2.6e-5). Float16 adds 2 spans on gold (1 on 841-dev, already covered by the rules) and removes none; 841-dev union numbers are identical. On 6,182 FineWeb-2 documents (1,236 with MPS and 4,946 with CUDA float32 references), CUDA float16 gave 0 span changes (max score change 0.012).

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix
1.0.3 324 24 98 123 97.93% 96.12% Label-note, e-Delivery and registry rules; placeholder and card fixes
1.1.0 324 24 98 123 97.93% 96.12% Batch API (predict_many, scrub_many), opt-in float16

1.0.3

Same weights, threshold, and API. Rules SHA f32d5c54….

  • Rules: NIP/REGON labels with footnote marks or a short gloss (NIP*:, NIP (Wykonawcy):); e-Delivery addresses broken across a line; DUNS, BDO and RPWDL numbers right after their own label, at a fixed length.
  • Wrapper: existing [PII]/[Telefon] placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document.
  • Card: weights_sha256 was the hash of the source training checkpoint, not of model.safetensors. It is now source_checkpoint_sha256, and model_safetensors_sha256 holds the published file's hash. Readers of weights_sha256 must switch keys.

841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed.

Known limit: an identifier with a placeholder inside it or right before it (NIP [PII] …, 8503[PII]…) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix
1.0.3 324 24 98 123 97.93% 96.12% Label-note, e-Delivery and registry rules; placeholder and card fixes

1.0.2

Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path.

Fix immediately labelled hyphenated country-area phones with one-digit country codes or wider area codes. Numeric continuations, slash lists, weak contact cues, and nearby title prose still abstain. Same weights, threshold, and API.

841-dev: rules recover 14 more whole spans; the union recovers one (323 β†’ 324), reducing residual passages 25 β†’ 24. Rules FP remain 98; union FP remain 123. Lost gold, new false characters, and new clean-passage damage are all zero. Rules SHA 3016ae5b….

Exact-span scores are recomputed from deduplicated raw rule/model span triples: precision 86.34%, recall 89.27%, F1 87.78%. The 1.0.1 card's precision/F1 were stale after the prefix trim changed exact rule/model duplicate counts; its character metrics were correct.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix

1.0.1

Prefix-only glued-email trim. Cluster-gated title-case 5–11 letter prefixes are dropped from the redaction when the remainder is already a lowercase-local email. Same epoch-5 weights. Lost gold 0. New clean-passage damage 0.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim

Rules SHA 4dcc441c…. Seven emails trimmed, 35 prefix characters dropped; the model still covers 25 of those positions. Suffix glue is unchanged.

1.0.0

First Hub snapshot. Seed 202609160, epoch 5, rules 547c0428…. Union FP 133 was the glued-email regex floor. This seed added no new false characters versus the historical incumbent.