Instructions to use SlayerLab/NERGAL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/NERGAL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="SlayerLab/NERGAL")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("SlayerLab/NERGAL") model = AutoModelForTokenClassification.from_pretrained("SlayerLab/NERGAL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
2.0.0
Browse filesCo-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- CHANGELOG.md +29 -0
- README.md +41 -14
- hybrid.json +224 -1
- names/NOTICE.md +5 -1
- names/convert.py +88 -0
- nergal.py +1 -1
- test_nergal.py +1 -1
CHANGELOG.md
CHANGED
|
@@ -8,6 +8,35 @@ Semver for this island:
|
|
| 8 |
|
| 9 |
Accuracy is 841-dev, union at 0.95. Through 1.1.2 the gold has 354 spans; from 1.2.0 it is restated to phone policy v3 (315 spans), and the two are not comparable. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
|
| 10 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
## 1.2.0
|
| 12 |
|
| 13 |
Same weights, threshold and API. Rules SHA `b238d5b8…`. The rules file is the frozen and scored candidate `73c185f7…` with two comments reworded; its syntax tree is identical. Phone policy v3 (labelling policy amended 2026-10-02) now applies at mask time, to the rules and to the model spans.
|
|
|
|
| 8 |
|
| 9 |
Accuracy is 841-dev, union at 0.95. Through 1.1.2 the gold has 354 spans; from 1.2.0 it is restated to phone policy v3 (315 spans), and the two are not comparable. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
|
| 10 |
|
| 11 |
+
## 2.0.0
|
| 12 |
+
|
| 13 |
+
Same NERGAL weights and threshold. Rules SHA `d1866243…`: the placeholder rename only; every rule span is unchanged. Major because the output placeholder changes.
|
| 14 |
+
|
| 15 |
+
- **Breaking — `[PHONE]` replaces `[Telefon]`.** Every phone mask now reads `[PHONE]`; `[PII]` is unchanged. Code that matches `[Telefon]` in the output must switch. Text scrubbed by 1.x still works as input: `[Telefon]` stays a rules boundary (a previous redaction, not a fresh `Telefon` label), like `[PHONE]` and `[PII]`.
|
| 16 |
+
- **Names, opt-in:** `Nergal.from_pretrained(..., names=True)` (CLI `--names`) loads `names/`, FastPDN NER — Polish PII by ArkadiuszPawlak (CC-BY-4.0, `names/NOTICE.md`), converted from ONNX to safetensors with no value changed. Its `PERSON`, `PERSON_F` and `PERSON_L` tags become `[PERSON]`: whole words, joined across spaces, tabs and no-break spaces (not punctuation or line breaks); a word touching a placeholder is skipped. It always runs in float32: `dtype` sets NERGAL only. Without `names=True`, `from_pretrained` skips `names/` (≈490 MB) and the output equals 1.2.0's except for the rename.
|
| 17 |
+
- **Union:** each character takes its highest-ranked label, phone > pii > person, and each run becomes one placeholder. A name inside a masked email stays `[PII]`.
|
| 18 |
+
- **Counts:** `scrub()` counts gain `person` (0 with names off). `union_placeholder_chars` includes `[PERSON]` placeholders. `predict()` / `predict_many()` still return NERGAL spans only; `nergal.names.spans_many(texts)` returns the person spans.
|
| 19 |
+
|
| 20 |
+
841-dev, names off: unchanged from 1.2.0 (whole 303/315, residual 9, rules FP 24, union FP 80, char P 98.59%, char R 97.83%). Its gold has no person labels, so names-on is not scored on it. The benchmark replay at 2.0.0 with names off gives 1.2.0's numbers on every one of its 12 cohorts.
|
| 21 |
+
|
| 22 |
+
**Name test (`nergal_names_test_v1`).** Dynaword and Polish mC4 web passages (70/30 in each panel), component-disjoint from every earlier set, labelled under `docs/pii-annotation-names.md`, frozen before any prediction and scored once. The random panel estimates the corpus; the targeted panel is enriched by selection (correspondence, signatures, honorifics, hard negatives) and says nothing about prevalence. Panels are never pooled. A name is whole when every letter and digit of it is masked (`[PII]` counts); a component is the independent unit, covered when every name in it is whole. Names-on, float32:
|
| 23 |
+
|
| 24 |
+
| Panel | Passages | Names | Whole | Exact | Components covered (95% CI) | Person false chars | Passages with one (95% CI) |
|
| 25 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 26 |
+
| Random | 300 | 614 | 586 (95.44%) | 565 | 98/116 (76.6–90.5%) | 1,051 | 34 (8.0–15.5%) |
|
| 27 |
+
| Targeted | 150 | 292 | 286 (97.95%) | 285 | 94/99 (88.6–98.3%) | 976 | 42 (21.0–35.9%) |
|
| 28 |
+
|
| 29 |
+
Names off hides no name; its false characters are 1 (random) and 29 (targeted). The set's phones (2 and 7) and other PII (2 and 5) are whole in both modes. For the record, neither shipped: float16 on Apple MPS equals float32 on both panels; FastPDN's INT8 ONNX on CPU hides 589 and 287 names whole with 1,009 and 935 person false characters. One reviewer labelled the set; the second pass is a same-day blind self-review of 100 passages (11 disagreed, adjudication changed 5), not independent agreement. `hybrid.json` `eval_names` holds these numbers and the score report's SHA.
|
| 30 |
+
|
| 31 |
+
Release gates. The name test replaces the planned 200-document spot check: a spot check reviews only what the model masked, so it measures neither recall nor names missed on clean-looking text. Throughput, local: 200 seeded Dynaword documents (1,103,229 characters), Apple M4 Max, MPS, float32, `scrub_many`, one process, off and on alternating twice. Names off 7,914 and 7,922 chars/s, peak RSS 1,393 and 1,388 MiB; names on 7,444 and 7,597 chars/s (−5%), 1,679 and 1,682 MiB (+290 MiB). MPS driver memory at the end of a run is 70,261 MiB off and 87,788 MiB on; that is the allocator cache, not a working set. Repeat runs mask identically; names on gives 1,309 `[PERSON]` placeholders, names off none, and both give 7 `[PHONE]` and 7 `[PII]`. CUDA reference from the runtime study (RTX 4090, 1,000 Dynaword documents, 4,215,149 characters, a prototype of this runtime with NERGAL float16 and names float32): 3 processes 83,479 → 72,662 chars/s (−13%), peak device memory 9,012 → 12,602 MiB; per process, torch peak 1,806 → 2,668 MiB. A float16 names model changed spans in 11 of those 1,000 documents, so it is not offered.
|
| 32 |
+
|
| 33 |
+
Known limits: names-on masks every person the policy covers, public officials and historical figures included, which is why it is opt-in. On the random panel 18 of 116 components keep part of a name, and 34 of 300 passages have a false person mask. The names model was fine-tuned partly on LLM-synthetic data that SlayerLab has not audited.
|
| 34 |
+
|
| 35 |
+
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 36 |
+
|---|---:|---:|---:|---:|---:|---:|---|
|
| 37 |
+
| 1.2.0 | 303 | 9 | 24 | 80 | 98.59% | 97.83% | Phone policy v3 at mask time, rules and model spans |
|
| 38 |
+
| 2.0.0, names off | 303 | 9 | 24 | 80 | 98.59% | 97.83% | `[PHONE]` replaces `[Telefon]`; opt-in `[PERSON]` names |
|
| 39 |
+
|
| 40 |
## 1.2.0
|
| 41 |
|
| 42 |
Same weights, threshold and API. Rules SHA `b238d5b8…`. The rules file is the frozen and scored candidate `73c185f7…` with two comments reworded; its syntax tree is identical. Phone policy v3 (labelling policy amended 2026-10-02) now applies at mask time, to the rules and to the model spans.
|
README.md
CHANGED
|
@@ -1,5 +1,8 @@
|
|
| 1 |
---
|
| 2 |
-
|
|
|
|
|
|
|
|
|
|
| 3 |
language:
|
| 4 |
- pl
|
| 5 |
base_model: FacebookAI/xlm-roberta-large
|
|
@@ -13,7 +16,7 @@ tags:
|
|
| 13 |
- hybrid
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# NERGAL
|
| 17 |
|
| 18 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 19 |
|
|
@@ -21,16 +24,17 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
|
|
| 21 |
|
| 22 |
## TL;DR
|
| 23 |
|
| 24 |
-
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[
|
| 25 |
|
| 26 |
- **Ground:** `scrub_pii.py` rules
|
| 27 |
- **Additive labels:** XLM-RoBERTa-large token classifier (epoch 5 of 7), BIO tags `phone` / `pii`, threshold 0.95
|
| 28 |
-
- **
|
|
|
|
| 29 |
- **Changes:** `CHANGELOG.md`
|
| 30 |
|
| 31 |
## What NERGAL detects — and what it does not
|
| 32 |
|
| 33 |
-
NERGAL masks **contact details and selected identifiers in Polish text**. It is not a general-purpose anonymizer: names, postal addresses and other personal information can remain in the output. Its corpus-masking policy also includes public, institutional and company contacts and identifiers.
|
| 34 |
|
| 35 |
### Detection scope
|
| 36 |
|
|
@@ -38,13 +42,14 @@ These are target categories, not a guarantee that every occurrence or format is
|
|
| 38 |
|
| 39 |
| Category | Values in scope | Replacement |
|
| 40 |
|---|---|---|
|
| 41 |
-
| Phone contacts | Phone, fax and SMS contact numbers of 7+ digits (country and area codes and keypad letters count), including foreign and vanity numbers, one span per number; an extension or alternative ending stays inside its number. Emergency, helpline, service and other numbers under 7 digits are not masked on their own | `[
|
| 42 |
| Email | Email addresses, including recoverable broken or incomplete addresses | `[PII]` |
|
| 43 |
| Personal identifiers | PESEL, passport and identity-document numbers | `[PII]` |
|
| 44 |
| Organization identifiers | NIP/VAT, REGON, KRS, GEMI, LEI and equivalent foreign company registration numbers; labelled DUNS, BDO and RPWDL register-book numbers | `[PII]` |
|
| 45 |
| Financial identifiers | Bank/account numbers, including Polish accounts and foreign IBANs | `[PII]` |
|
| 46 |
| Property identifiers | Land-register (księga wieczysta, KW) numbers | `[PII]` |
|
| 47 |
| Electronic contacts and access | e-Doręczenia and ePUAP addresses, explicitly labelled numeric access PINs (including `pin=` URL values), GG account IDs | `[PII]` |
|
|
|
|
| 48 |
|
| 49 |
Context matters: a number that resembles a phone or identifier is not automatically in scope. Rules use labels, format checks and, for some unlabelled identifiers, checksums; the model adds contextual detections. Coverage varies by category, and the aggregate benchmark below does not establish recall for every category.
|
| 50 |
|
|
@@ -52,7 +57,7 @@ Context matters: a number that resembles a phone or identifier is not automatica
|
|
| 52 |
|
| 53 |
NERGAL is not designed to remove:
|
| 54 |
|
| 55 |
-
- **Personal names**, including private individuals and public officials; organization names.
|
| 56 |
- **Postal/street addresses, dates of birth and ages.**
|
| 57 |
- **Social-media handles, ordinary URLs and filenames.** An in-scope value inside a URL, such as a labelled numeric PIN, can still be masked.
|
| 58 |
- **Vehicle registration plates and generic serial, model or version codes.**
|
|
@@ -61,9 +66,9 @@ NERGAL is not designed to remove:
|
|
| 61 |
|
| 62 |
These are intended exclusions; false positives can still mask some of this content.
|
| 63 |
|
| 64 |
-
### Known gaps in
|
| 65 |
|
| 66 |
-
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released
|
| 67 |
|
| 68 |
## Versions
|
| 69 |
|
|
@@ -72,10 +77,28 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
|
|
| 72 |
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 73 |
|---|---:|---:|---:|---:|---:|---:|---|
|
| 74 |
| 1.1.2 | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Email boundary rules |
|
| 75 |
-
|
|
|
|
|
| 76 |
|
| 77 |
1.0.0–1.1.2 were scored on the earlier 354-span gold (whole 323 → 326, union FP 133 → 80); per-version numbers are in `CHANGELOG.md`.
|
| 78 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
## Cue-less phone test
|
| 80 |
|
| 81 |
841-dev holds few phones without a contact label. This targeted set does: 150 web passages selected for them, labelled and frozen before any prediction. 1.1.1 added the grouped national forms and raised phones wholly masked from 37 to 66 of 111, losing none. Its benchmark half, restated to phone policy v3 (70 passages): 1.2.0 wholly masks 62 of 83 values, with 33 false characters. The set is enriched by selection and has a single reviewer, so it says nothing about how common such phones are.
|
|
@@ -104,7 +127,7 @@ XLM-R and epoch 5 were chosen in September 2026, on the earlier gold and before
|
|
| 104 |
|
| 105 |
## Load
|
| 106 |
|
| 107 |
-
Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[
|
| 108 |
|
| 109 |
```python
|
| 110 |
from pathlib import Path
|
|
@@ -116,11 +139,15 @@ sys.path.insert(0, str(root))
|
|
| 116 |
from nergal import Nergal
|
| 117 |
|
| 118 |
nergal = Nergal.from_pretrained(root, local_files_only=True)
|
| 119 |
-
masked, counts = nergal.scrub(text)
|
| 120 |
```
|
| 121 |
|
|
|
|
|
|
|
| 122 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 123 |
|
| 124 |
-
`hybrid.json` records the version, threshold, 841-dev `eval` block and
|
|
|
|
|
|
|
| 125 |
|
| 126 |
-
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
|
|
|
| 1 |
---
|
| 2 |
+
# MIT for NERGAL (code, rules, weights); names/ (FastPDN NER — Polish PII) is CC-BY-4.0, see names/NOTICE.md
|
| 3 |
+
license: other
|
| 4 |
+
license_name: mit-and-cc-by-4.0
|
| 5 |
+
license_link: https://huggingface.co/SlayerLab/NERGAL/blob/main/README.md#license
|
| 6 |
language:
|
| 7 |
- pl
|
| 8 |
base_model: FacebookAI/xlm-roberta-large
|
|
|
|
| 16 |
- hybrid
|
| 17 |
---
|
| 18 |
|
| 19 |
+
# NERGAL 2.0.0
|
| 20 |
|
| 21 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 22 |
|
|
|
|
| 24 |
|
| 25 |
## TL;DR
|
| 26 |
|
| 27 |
+
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[PHONE]` or `[PII]`. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked. With `names=True`, a second model adds person names as `[PERSON]` ([Names](#names)).
|
| 28 |
|
| 29 |
- **Ground:** `scrub_pii.py` rules
|
| 30 |
- **Additive labels:** XLM-RoBERTa-large token classifier (epoch 5 of 7), BIO tags `phone` / `pii`, threshold 0.95
|
| 31 |
+
- **Names (opt-in):** [FastPDN NER — Polish PII](https://huggingface.co/ArkadiuszPawlak/fastpdn-ner-polish-pii) (CC-BY-4.0) in `names/`, float32
|
| 32 |
+
- **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes; names cost about 13%
|
| 33 |
- **Changes:** `CHANGELOG.md`
|
| 34 |
|
| 35 |
## What NERGAL detects — and what it does not
|
| 36 |
|
| 37 |
+
NERGAL masks **contact details and selected identifiers in Polish text**. It is not a general-purpose anonymizer: names (masked only with `names=True`), postal addresses and other personal information can remain in the output. Its corpus-masking policy also includes public, institutional and company contacts and identifiers.
|
| 38 |
|
| 39 |
### Detection scope
|
| 40 |
|
|
|
|
| 42 |
|
| 43 |
| Category | Values in scope | Replacement |
|
| 44 |
|---|---|---|
|
| 45 |
+
| Phone contacts | Phone, fax and SMS contact numbers of 7+ digits (country and area codes and keypad letters count), including foreign and vanity numbers, one span per number; an extension or alternative ending stays inside its number. Emergency, helpline, service and other numbers under 7 digits are not masked on their own | `[PHONE]` |
|
| 46 |
| Email | Email addresses, including recoverable broken or incomplete addresses | `[PII]` |
|
| 47 |
| Personal identifiers | PESEL, passport and identity-document numbers | `[PII]` |
|
| 48 |
| Organization identifiers | NIP/VAT, REGON, KRS, GEMI, LEI and equivalent foreign company registration numbers; labelled DUNS, BDO and RPWDL register-book numbers | `[PII]` |
|
| 49 |
| Financial identifiers | Bank/account numbers, including Polish accounts and foreign IBANs | `[PII]` |
|
| 50 |
| Property identifiers | Land-register (księga wieczysta, KW) numbers | `[PII]` |
|
| 51 |
| Electronic contacts and access | e-Doręczenia and ePUAP addresses, explicitly labelled numeric access PINs (including `pin=` URL values), GG account IDs | `[PII]` |
|
| 52 |
+
| Person names, only with `names=True` | Names of people, public officials and historical figures included ([Names](#names)) | `[PERSON]` |
|
| 53 |
|
| 54 |
Context matters: a number that resembles a phone or identifier is not automatically in scope. Rules use labels, format checks and, for some unlabelled identifiers, checksums; the model adds contextual detections. Coverage varies by category, and the aggregate benchmark below does not establish recall for every category.
|
| 55 |
|
|
|
|
| 57 |
|
| 58 |
NERGAL is not designed to remove:
|
| 59 |
|
| 60 |
+
- **Personal names** with names off, including private individuals and public officials; organization names in either mode.
|
| 61 |
- **Postal/street addresses, dates of birth and ages.**
|
| 62 |
- **Social-media handles, ordinary URLs and filenames.** An in-scope value inside a URL, such as a labelled numeric PIN, can still be masked.
|
| 63 |
- **Vehicle registration plates and generic serial, model or version codes.**
|
|
|
|
| 66 |
|
| 67 |
These are intended exclusions; false positives can still mask some of this content.
|
| 68 |
|
| 69 |
+
### Known gaps in 2.0.0
|
| 70 |
|
| 71 |
+
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 2.0.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (`12 / 345 / 678`) is left as text. An identifier with a `[PII]` / `[PHONE]` / `[Telefon]` placeholder inside it or right before it can be missed. Names-on gaps are under [Names](#names).
|
| 72 |
|
| 73 |
## Versions
|
| 74 |
|
|
|
|
| 77 |
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 78 |
|---|---:|---:|---:|---:|---:|---:|---|
|
| 79 |
| 1.1.2 | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Email boundary rules |
|
| 80 |
+
| 1.2.0 | 303 | 9 | 24 | 80 | 98.59% | 97.83% | Phone policy v3 at mask time, rules and model spans |
|
| 81 |
+
| **2.0.0, names off** | **303** | **9** | **24** | **80** | **98.59%** | **97.83%** | `[PHONE]` replaces `[Telefon]`; opt-in `[PERSON]` names |
|
| 82 |
|
| 83 |
1.0.0–1.1.2 were scored on the earlier 354-span gold (whole 323 → 326, union FP 133 → 80); per-version numbers are in `CHANGELOG.md`.
|
| 84 |
|
| 85 |
+
## Names
|
| 86 |
+
|
| 87 |
+
`names=True` (CLI `--names`) adds a second model, [FastPDN NER — Polish PII](https://huggingface.co/ArkadiuszPawlak/fastpdn-ner-polish-pii) by ArkadiuszPawlak, revision `636c57fe…`, fine-tuned from [clarin-pl/FastPDN](https://huggingface.co/clarin-pl/FastPDN) on filtered KPWr plus LLM-synthetic data. `names/` holds it converted from ONNX to safetensors with no value changed. **`names/` is licensed [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/), not MIT**; `names/NOTICE.md` has the attribution, the changes and the file hashes. NERGAL uses only its person tags (`PERSON`, `PERSON_F`, `PERSON_L`), masks whole words, and always runs it in float32.
|
| 88 |
+
|
| 89 |
+
**Scope with names on.** The name test labels every person: private individuals, public officials, politicians and historical figures alike, plus fictional characters, nicknames, partial names and initials. Titles and roles stay outside the span. Names inside a street, patronage, organisation or work title, eponyms, usernames, and names inside emails or URLs are not person spans. Masking public figures removes who said, signed or wrote what, which is why names are opt-in. This is the labelling target; the model does not meet it everywhere.
|
| 90 |
+
|
| 91 |
+
**Name test** (`nergal_names_test_v1`): Dynaword and Polish mC4 web passages, component-disjoint from every earlier set, labelled by one reviewer, frozen before any prediction and scored once. The random panel estimates the corpus; the targeted panel is enriched by selection and says nothing about prevalence. A name is whole when every letter and digit of it is masked.
|
| 92 |
+
|
| 93 |
+
| Panel | Passages | Names | Whole | Components covered (95% CI) | Person false chars | Passages with one |
|
| 94 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 95 |
+
| Random | 300 | 614 | 586 (95.44%) | 98/116 (76.6–90.5%) | 1,051 | 34 (11.3%) |
|
| 96 |
+
| Targeted | 150 | 292 | 286 (97.95%) | 94/99 (88.6–98.3%) | 976 | 42 (28.0%) |
|
| 97 |
+
|
| 98 |
+
Names off hides no name. A component is the independent unit, covered when every name in it is whole: on the random panel 18 of 116 keep part of a name. The second pass is a same-day blind self-review of 100 passages by the same reviewer (11 disagreed), not independent agreement. `hybrid.json` `eval_names` has the full numbers, including float16 and INT8 scored for the record; neither ships.
|
| 99 |
+
|
| 100 |
+
**Cost:** on one RTX 4090 with 3 processes, names took a prototype of this runtime from 83,479 to 72,662 chars/s (−13%) and peak device memory from 9,012 to 12,602 MiB. Local measurements are in `CHANGELOG.md`.
|
| 101 |
+
|
| 102 |
## Cue-less phone test
|
| 103 |
|
| 104 |
841-dev holds few phones without a contact label. This targeted set does: 150 web passages selected for them, labelled and frozen before any prediction. 1.1.1 added the grouped national forms and raised phones wholly masked from 37 to 66 of 111, losing none. Its benchmark half, restated to phone policy v3 (70 passages): 1.2.0 wholly masks 62 of 83 values, with 33 false characters. The set is enriched by selection and has a single reviewer, so it says nothing about how common such phones are.
|
|
|
|
| 127 |
|
| 128 |
## Load
|
| 129 |
|
| 130 |
+
Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with `[PHONE]` / `[PII]` (and `[PERSON]` with `names=True`). Each character takes its highest-ranked label, phone > pii > person.
|
| 131 |
|
| 132 |
```python
|
| 133 |
from pathlib import Path
|
|
|
|
| 139 |
from nergal import Nergal
|
| 140 |
|
| 141 |
nergal = Nergal.from_pretrained(root, local_files_only=True)
|
| 142 |
+
masked, counts = nergal.scrub(text) # counts: phone, pii, person
|
| 143 |
```
|
| 144 |
|
| 145 |
+
Names: `snapshot_download("SlayerLab/NERGAL")` fetches `names/` too; `Nergal.from_pretrained("SlayerLab/NERGAL", names=True)` downloads it only when asked. From the command line: `python nergal.py --names < in.txt > out.txt`.
|
| 146 |
+
|
| 147 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 148 |
|
| 149 |
+
`hybrid.json` records the version, threshold, 841-dev `eval` block, weight hashes, the `names` source and file hashes, and the name test `eval_names`. `test_nergal.py` is synthetic (no corpus text): `python -m unittest test_nergal`.
|
| 150 |
+
|
| 151 |
+
## License
|
| 152 |
|
| 153 |
+
NERGAL code, rules and weights: MIT. `names/`: CC-BY-4.0, attribution in `names/NOTICE.md`. Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
hybrid.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
-
"version": "
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
|
@@ -64,5 +64,228 @@
|
|
| 64 |
"model.safetensors": "2fa84ec6abd0c1b12ba0ba1548e89c9fca36fc58f03e66074db423b0e98ec8e6",
|
| 65 |
"tokenizer.json": "3a690c2d605076ad3901d946d4c0145fbfb19fef2f01370c7388430aa9f31edf"
|
| 66 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 67 |
}
|
| 68 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
+
"version": "2.0.0",
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
|
|
|
| 64 |
"model.safetensors": "2fa84ec6abd0c1b12ba0ba1548e89c9fca36fc58f03e66074db423b0e98ec8e6",
|
| 65 |
"tokenizer.json": "3a690c2d605076ad3901d946d4c0145fbfb19fef2f01370c7388430aa9f31edf"
|
| 66 |
}
|
| 67 |
+
},
|
| 68 |
+
"eval_names": {
|
| 69 |
+
"split": "nergal_names_test_v1",
|
| 70 |
+
"policy": "docs/pii-annotation-names.md",
|
| 71 |
+
"reviewers": 1,
|
| 72 |
+
"second_pass": {
|
| 73 |
+
"passages": 100,
|
| 74 |
+
"disagreeing": 11,
|
| 75 |
+
"adjudication_changed": 5,
|
| 76 |
+
"gap_days": 0
|
| 77 |
+
},
|
| 78 |
+
"note": "single reviewer; the second pass is a same-day blind self-review of a sample, not independent agreement",
|
| 79 |
+
"runtime": "fp32",
|
| 80 |
+
"devices": {
|
| 81 |
+
"fp32": "cpu",
|
| 82 |
+
"fp16": "mps",
|
| 83 |
+
"int8": "cpu"
|
| 84 |
+
},
|
| 85 |
+
"score_report_sha256": "9dec1517dae8fefb4857acab975d83d82b217373b1e7efdee56a685eaf9dc6b2",
|
| 86 |
+
"panels": {
|
| 87 |
+
"random": {
|
| 88 |
+
"passages": 300,
|
| 89 |
+
"systems": {
|
| 90 |
+
"fp16": {
|
| 91 |
+
"person_gold": 614,
|
| 92 |
+
"person_whole": 586,
|
| 93 |
+
"person_whole_exact": 565,
|
| 94 |
+
"name_recall": 0.9544,
|
| 95 |
+
"person_components": 116,
|
| 96 |
+
"covered_components": 98,
|
| 97 |
+
"coverage_ci95": [
|
| 98 |
+
0.7659,
|
| 99 |
+
0.9054
|
| 100 |
+
],
|
| 101 |
+
"person_false_chars": 1051,
|
| 102 |
+
"person_false_passages": 34,
|
| 103 |
+
"person_false_ci95": [
|
| 104 |
+
0.0798,
|
| 105 |
+
0.1548
|
| 106 |
+
],
|
| 107 |
+
"false_chars": 1052,
|
| 108 |
+
"false_mask_passages": 35,
|
| 109 |
+
"phone_gold": 2,
|
| 110 |
+
"phone_whole": 2,
|
| 111 |
+
"pii_gold": 2,
|
| 112 |
+
"pii_whole": 2
|
| 113 |
+
},
|
| 114 |
+
"fp32": {
|
| 115 |
+
"person_gold": 614,
|
| 116 |
+
"person_whole": 586,
|
| 117 |
+
"person_whole_exact": 565,
|
| 118 |
+
"name_recall": 0.9544,
|
| 119 |
+
"person_components": 116,
|
| 120 |
+
"covered_components": 98,
|
| 121 |
+
"coverage_ci95": [
|
| 122 |
+
0.7659,
|
| 123 |
+
0.9054
|
| 124 |
+
],
|
| 125 |
+
"person_false_chars": 1051,
|
| 126 |
+
"person_false_passages": 34,
|
| 127 |
+
"person_false_ci95": [
|
| 128 |
+
0.0798,
|
| 129 |
+
0.1548
|
| 130 |
+
],
|
| 131 |
+
"false_chars": 1052,
|
| 132 |
+
"false_mask_passages": 35,
|
| 133 |
+
"phone_gold": 2,
|
| 134 |
+
"phone_whole": 2,
|
| 135 |
+
"pii_gold": 2,
|
| 136 |
+
"pii_whole": 2
|
| 137 |
+
},
|
| 138 |
+
"int8": {
|
| 139 |
+
"person_gold": 614,
|
| 140 |
+
"person_whole": 589,
|
| 141 |
+
"person_whole_exact": 568,
|
| 142 |
+
"name_recall": 0.9593,
|
| 143 |
+
"person_components": 116,
|
| 144 |
+
"covered_components": 101,
|
| 145 |
+
"coverage_ci95": [
|
| 146 |
+
0.7957,
|
| 147 |
+
0.9258
|
| 148 |
+
],
|
| 149 |
+
"person_false_chars": 1009,
|
| 150 |
+
"person_false_passages": 37,
|
| 151 |
+
"person_false_ci95": [
|
| 152 |
+
0.0883,
|
| 153 |
+
0.166
|
| 154 |
+
],
|
| 155 |
+
"false_chars": 1010,
|
| 156 |
+
"false_mask_passages": 38,
|
| 157 |
+
"phone_gold": 2,
|
| 158 |
+
"phone_whole": 2,
|
| 159 |
+
"pii_gold": 2,
|
| 160 |
+
"pii_whole": 2
|
| 161 |
+
},
|
| 162 |
+
"off": {
|
| 163 |
+
"person_gold": 614,
|
| 164 |
+
"person_whole": 0,
|
| 165 |
+
"person_whole_exact": 0,
|
| 166 |
+
"name_recall": 0.0,
|
| 167 |
+
"person_components": 116,
|
| 168 |
+
"covered_components": 0,
|
| 169 |
+
"coverage_ci95": [
|
| 170 |
+
0.0,
|
| 171 |
+
0.0313
|
| 172 |
+
],
|
| 173 |
+
"person_false_chars": 0,
|
| 174 |
+
"person_false_passages": 0,
|
| 175 |
+
"person_false_ci95": [
|
| 176 |
+
0.0,
|
| 177 |
+
0.0122
|
| 178 |
+
],
|
| 179 |
+
"false_chars": 1,
|
| 180 |
+
"false_mask_passages": 1,
|
| 181 |
+
"phone_gold": 2,
|
| 182 |
+
"phone_whole": 2,
|
| 183 |
+
"pii_gold": 2,
|
| 184 |
+
"pii_whole": 2
|
| 185 |
+
}
|
| 186 |
+
}
|
| 187 |
+
},
|
| 188 |
+
"targeted": {
|
| 189 |
+
"passages": 150,
|
| 190 |
+
"systems": {
|
| 191 |
+
"fp16": {
|
| 192 |
+
"person_gold": 292,
|
| 193 |
+
"person_whole": 286,
|
| 194 |
+
"person_whole_exact": 285,
|
| 195 |
+
"name_recall": 0.9795,
|
| 196 |
+
"person_components": 99,
|
| 197 |
+
"covered_components": 94,
|
| 198 |
+
"coverage_ci95": [
|
| 199 |
+
0.8861,
|
| 200 |
+
0.9834
|
| 201 |
+
],
|
| 202 |
+
"person_false_chars": 976,
|
| 203 |
+
"person_false_passages": 42,
|
| 204 |
+
"person_false_ci95": [
|
| 205 |
+
0.2098,
|
| 206 |
+
0.3591
|
| 207 |
+
],
|
| 208 |
+
"false_chars": 1005,
|
| 209 |
+
"false_mask_passages": 43,
|
| 210 |
+
"phone_gold": 7,
|
| 211 |
+
"phone_whole": 7,
|
| 212 |
+
"pii_gold": 5,
|
| 213 |
+
"pii_whole": 5
|
| 214 |
+
},
|
| 215 |
+
"fp32": {
|
| 216 |
+
"person_gold": 292,
|
| 217 |
+
"person_whole": 286,
|
| 218 |
+
"person_whole_exact": 285,
|
| 219 |
+
"name_recall": 0.9795,
|
| 220 |
+
"person_components": 99,
|
| 221 |
+
"covered_components": 94,
|
| 222 |
+
"coverage_ci95": [
|
| 223 |
+
0.8861,
|
| 224 |
+
0.9834
|
| 225 |
+
],
|
| 226 |
+
"person_false_chars": 976,
|
| 227 |
+
"person_false_passages": 42,
|
| 228 |
+
"person_false_ci95": [
|
| 229 |
+
0.2098,
|
| 230 |
+
0.3591
|
| 231 |
+
],
|
| 232 |
+
"false_chars": 1005,
|
| 233 |
+
"false_mask_passages": 43,
|
| 234 |
+
"phone_gold": 7,
|
| 235 |
+
"phone_whole": 7,
|
| 236 |
+
"pii_gold": 5,
|
| 237 |
+
"pii_whole": 5
|
| 238 |
+
},
|
| 239 |
+
"int8": {
|
| 240 |
+
"person_gold": 292,
|
| 241 |
+
"person_whole": 287,
|
| 242 |
+
"person_whole_exact": 285,
|
| 243 |
+
"name_recall": 0.9829,
|
| 244 |
+
"person_components": 99,
|
| 245 |
+
"covered_components": 95,
|
| 246 |
+
"coverage_ci95": [
|
| 247 |
+
0.8998,
|
| 248 |
+
0.9889
|
| 249 |
+
],
|
| 250 |
+
"person_false_chars": 935,
|
| 251 |
+
"person_false_passages": 41,
|
| 252 |
+
"person_false_ci95": [
|
| 253 |
+
0.2038,
|
| 254 |
+
0.352
|
| 255 |
+
],
|
| 256 |
+
"false_chars": 964,
|
| 257 |
+
"false_mask_passages": 42,
|
| 258 |
+
"phone_gold": 7,
|
| 259 |
+
"phone_whole": 7,
|
| 260 |
+
"pii_gold": 5,
|
| 261 |
+
"pii_whole": 5
|
| 262 |
+
},
|
| 263 |
+
"off": {
|
| 264 |
+
"person_gold": 292,
|
| 265 |
+
"person_whole": 0,
|
| 266 |
+
"person_whole_exact": 0,
|
| 267 |
+
"name_recall": 0.0,
|
| 268 |
+
"person_components": 99,
|
| 269 |
+
"covered_components": 0,
|
| 270 |
+
"coverage_ci95": [
|
| 271 |
+
0.0,
|
| 272 |
+
0.0366
|
| 273 |
+
],
|
| 274 |
+
"person_false_chars": 0,
|
| 275 |
+
"person_false_passages": 0,
|
| 276 |
+
"person_false_ci95": [
|
| 277 |
+
0.0,
|
| 278 |
+
0.0243
|
| 279 |
+
],
|
| 280 |
+
"false_chars": 29,
|
| 281 |
+
"false_mask_passages": 2,
|
| 282 |
+
"phone_gold": 7,
|
| 283 |
+
"phone_whole": 7,
|
| 284 |
+
"pii_gold": 5,
|
| 285 |
+
"pii_whole": 5
|
| 286 |
+
}
|
| 287 |
+
}
|
| 288 |
+
}
|
| 289 |
+
}
|
| 290 |
}
|
| 291 |
}
|
names/NOTICE.md
CHANGED
|
@@ -14,7 +14,10 @@
|
|
| 14 |
- `model.safetensors` is upstream `model.onnx` (FP32, sha256
|
| 15 |
`11924bfd20eb478ee614e5e1184536825db4747753d341af7c04b940efa84645`) converted by `convert.py` (sha256
|
| 16 |
`a698bc993b253fc70d960aa14af37a5088afb79c77dae70557c7056ce394333e`): 199 tensors copied, 73 linear weights
|
| 17 |
-
transposed from ONNX [in, out] to torch [out, in], no value changed.
|
|
|
|
|
|
|
|
|
|
| 18 |
- Review Gate 14, on 1,000 documents (1,714 windows, 614,466 tokens): torch FP32 against ONNX Runtime FP32, max
|
| 19 |
absolute logit difference 3.4e-4 (within 1e-3), every token's argmax identical, every person range identical.
|
| 20 |
- NERGAL uses only the PERSON, PERSON_F and PERSON_L tags, masked as `[PERSON]`.
|
|
@@ -26,4 +29,5 @@
|
|
| 26 |
| `config.json` | `3cef51cd7d184478adbdec9f68730db279386fe145cb32473a0622edea923578` |
|
| 27 |
| `model.safetensors` | `2fa84ec6abd0c1b12ba0ba1548e89c9fca36fc58f03e66074db423b0e98ec8e6` |
|
| 28 |
| `tokenizer.json` | `3a690c2d605076ad3901d946d4c0145fbfb19fef2f01370c7388430aa9f31edf` |
|
|
|
|
| 29 |
| `NOTICE.md` | this file |
|
|
|
|
| 14 |
- `model.safetensors` is upstream `model.onnx` (FP32, sha256
|
| 15 |
`11924bfd20eb478ee614e5e1184536825db4747753d341af7c04b940efa84645`) converted by `convert.py` (sha256
|
| 16 |
`a698bc993b253fc70d960aa14af37a5088afb79c77dae70557c7056ce394333e`): 199 tensors copied, 73 linear weights
|
| 17 |
+
transposed from ONNX [in, out] to torch [out, in], no value changed. `convert.py` is included: it reads the upstream
|
| 18 |
+
snapshot at that revision from the default Hugging Face cache and is create-only, so run a copy outside `names/`. A
|
| 19 |
+
rerun under another `safetensors` version can give a different file sha256 with every tensor identical: only the
|
| 20 |
+
`__metadata__` key order differs.
|
| 21 |
- Review Gate 14, on 1,000 documents (1,714 windows, 614,466 tokens): torch FP32 against ONNX Runtime FP32, max
|
| 22 |
absolute logit difference 3.4e-4 (within 1e-3), every token's argmax identical, every person range identical.
|
| 23 |
- NERGAL uses only the PERSON, PERSON_F and PERSON_L tags, masked as `[PERSON]`.
|
|
|
|
| 29 |
| `config.json` | `3cef51cd7d184478adbdec9f68730db279386fe145cb32473a0622edea923578` |
|
| 30 |
| `model.safetensors` | `2fa84ec6abd0c1b12ba0ba1548e89c9fca36fc58f03e66074db423b0e98ec8e6` |
|
| 31 |
| `tokenizer.json` | `3a690c2d605076ad3901d946d4c0145fbfb19fef2f01370c7388430aa9f31edf` |
|
| 32 |
+
| `convert.py` | `a698bc993b253fc70d960aa14af37a5088afb79c77dae70557c7056ce394333e` |
|
| 33 |
| `NOTICE.md` | this file |
|
names/convert.py
ADDED
|
@@ -0,0 +1,88 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""FastPDN-PII FP32 ONNX -> model.safetensors for BertForTokenClassification (plan Task 4a Step 2, Review Gate 14 A).
|
| 2 |
+
|
| 3 |
+
Named initializers keep their torch names. Each linear weight is the MatMul input that feeds the Add or BiasGelu taking
|
| 4 |
+
its named bias, transposed from ONNX's [in, out] to torch's [out, in]. Every weight is used once, the state dict loads
|
| 5 |
+
strictly, and no unused float tensor is left in the graph; anything else stops. Weights are copied, never changed.
|
| 6 |
+
Usage: python convert.py # writes model.safetensors and convert.json here; create-only
|
| 7 |
+
"""
|
| 8 |
+
from pathlib import Path
|
| 9 |
+
import hashlib, json
|
| 10 |
+
|
| 11 |
+
import numpy as np
|
| 12 |
+
from onnx import numpy_helper
|
| 13 |
+
|
| 14 |
+
HERE = Path(__file__).resolve().parent
|
| 15 |
+
REPO, REVISION = 'ArkadiuszPawlak/fastpdn-ner-polish-pii', '636c57fe77b0afad8f065e7c76243dc051437035'
|
| 16 |
+
SNAPSHOT = Path.home()/'.cache/huggingface/hub/models--ArkadiuszPawlak--fastpdn-ner-polish-pii/snapshots'/REVISION
|
| 17 |
+
ONNX_SHA256 = '11924bfd20eb478ee614e5e1184536825db4747753d341af7c04b940efa84645' # model.onnx, LFS pin at REVISION
|
| 18 |
+
CONFIG_SHA256 = '3cef51cd7d184478adbdec9f68730db279386fe145cb32473a0622edea923578'
|
| 19 |
+
NAMED = ('bert.', 'classifier.')
|
| 20 |
+
BUFFERS = ('bert.embeddings.position_ids',) # non-persistent in transformers: checked, not saved
|
| 21 |
+
|
| 22 |
+
|
| 23 |
+
def sha(path):
|
| 24 |
+
h = hashlib.sha256()
|
| 25 |
+
with open(path, 'rb') as f:
|
| 26 |
+
for block in iter(lambda: f.read(1 << 20), b''):
|
| 27 |
+
h.update(block)
|
| 28 |
+
return h.hexdigest()
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
def weights(graph):
|
| 32 |
+
"""({torch name: array}, used initializer names)."""
|
| 33 |
+
init = {i.name: numpy_helper.to_array(i) for i in graph.initializer}
|
| 34 |
+
producer = {o: n for n in graph.node for o in n.output}
|
| 35 |
+
state = {k: v for k, v in init.items() if k.startswith(NAMED)}
|
| 36 |
+
used = set(state)
|
| 37 |
+
for n in graph.node:
|
| 38 |
+
bias = [x for x in n.input if x.endswith('.bias') and 'LayerNorm' not in x and x in init]
|
| 39 |
+
if n.op_type not in ('Add', 'BiasGelu') or not bias:
|
| 40 |
+
continue
|
| 41 |
+
(b,), (other,) = bias, [x for x in n.input if x not in bias]
|
| 42 |
+
matmul = producer.get(other)
|
| 43 |
+
assert matmul is not None and matmul.op_type == 'MatMul' and matmul.input[1] in init, b
|
| 44 |
+
w, key = matmul.input[1], b[:-len('bias')] + 'weight'
|
| 45 |
+
assert w not in used and key not in state, (w, key)
|
| 46 |
+
state[key] = init[w].T.copy()
|
| 47 |
+
used.add(w)
|
| 48 |
+
return state, used
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
def leftovers(init, used):
|
| 52 |
+
"""Unused float tensors with more than one value: a weight the mapping missed. Scalars and int shapes pass."""
|
| 53 |
+
return sorted(k for k, v in init.items() if k not in used and np.asarray(v).dtype.kind == 'f' and np.size(v) > 1)
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
def main():
|
| 57 |
+
import onnx, torch
|
| 58 |
+
from safetensors.torch import save_file
|
| 59 |
+
from transformers import BertConfig, BertForTokenClassification
|
| 60 |
+
out = HERE/'model.safetensors'
|
| 61 |
+
assert not out.exists(), 'create-only'
|
| 62 |
+
assert sha(SNAPSHOT/'model.onnx') == ONNX_SHA256 and sha(SNAPSHOT/'config.json') == CONFIG_SHA256
|
| 63 |
+
graph = onnx.load(SNAPSHOT/'model.onnx').graph
|
| 64 |
+
init = {i.name: numpy_helper.to_array(i) for i in graph.initializer}
|
| 65 |
+
state, used = weights(graph)
|
| 66 |
+
left = leftovers(init, used)
|
| 67 |
+
assert not left, f'{len(left)} unused float tensors'
|
| 68 |
+
for k in BUFFERS:
|
| 69 |
+
v = state.pop(k)
|
| 70 |
+
assert np.array_equal(v, np.arange(v.shape[-1]).reshape(v.shape)), k
|
| 71 |
+
tensors = {k: torch.from_numpy(np.ascontiguousarray(v)) for k, v in sorted(state.items())}
|
| 72 |
+
model = BertForTokenClassification(BertConfig.from_pretrained(SNAPSHOT))
|
| 73 |
+
expected = model.state_dict()
|
| 74 |
+
assert {k: tuple(v.shape) for k, v in tensors.items()} == {k: tuple(v.shape) for k, v in expected.items()}, 'keys/shapes'
|
| 75 |
+
assert all(t.dtype == torch.float32 for t in tensors.values())
|
| 76 |
+
model.load_state_dict(tensors, strict=True)
|
| 77 |
+
save_file(tensors, out, metadata=dict(source=f'{REPO}@{REVISION}/model.onnx', onnx_sha256=ONNX_SHA256,
|
| 78 |
+
note='converted from ONNX, weights unchanged'))
|
| 79 |
+
report = dict(repo=REPO, revision=REVISION, onnx_sha256=ONNX_SHA256, config_sha256=CONFIG_SHA256,
|
| 80 |
+
initializers=len(init), used=len(used), tensors=len(tensors),
|
| 81 |
+
transposed=len(used) - sum(k.startswith(NAMED) for k in used),
|
| 82 |
+
unused_constants=len(init) - len(used), safetensors_sha256=sha(out), convert_sha256=sha(__file__))
|
| 83 |
+
(HERE/'convert.json').write_text(json.dumps(report, indent=1) + '\n')
|
| 84 |
+
return report
|
| 85 |
+
|
| 86 |
+
|
| 87 |
+
if __name__ == '__main__':
|
| 88 |
+
print(json.dumps(main()))
|
nergal.py
CHANGED
|
@@ -17,7 +17,7 @@ import scrub_pii
|
|
| 17 |
from scrub_pii import LEGACY_PHONE_TAG, PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
-
VERSION = '
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
|
|
|
| 17 |
from scrub_pii import LEGACY_PHONE_TAG, PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
+
VERSION = '2.0.0'
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
test_nergal.py
CHANGED
|
@@ -15,7 +15,7 @@ class NergalTests(unittest.TestCase):
|
|
| 15 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 16 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 17 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 18 |
-
self.assertEqual(VERSION, '
|
| 19 |
self.assertEqual(card['version'], VERSION)
|
| 20 |
self.assertEqual(card['eval']['union_fp'], 80)
|
| 21 |
self.assertEqual(card['eval']['rules_fp'], 24)
|
|
|
|
| 15 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 16 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 17 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 18 |
+
self.assertEqual(VERSION, '2.0.0')
|
| 19 |
self.assertEqual(card['version'], VERSION)
|
| 20 |
self.assertEqual(card['eval']['union_fp'], 80)
|
| 21 |
self.assertEqual(card['eval']['rules_fp'], 24)
|