ppuzio Claude Opus 5.5 commited on
Commit
8329d5b
Β·
1 Parent(s): be6377b

2.0.0: CUDA release gate numbers

Browse files

Shipped runtime on an RTX 4090 (Hub v2.0.0 snapshot): names cost 9% at
1 process and 11% at 3 processes; CHANGELOG release gates and the card's
Cost paragraph now cite these instead of the prototype.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Files changed (2) hide show
  1. CHANGELOG.md +1 -1
  2. README.md +2 -2
CHANGELOG.md CHANGED
@@ -28,7 +28,7 @@ Same NERGAL weights and threshold. Rules SHA `d1866243…`: the placeholder rena
28
 
29
  Names off hides no name; its false characters are 1 (random) and 29 (targeted). The set's phones (2 and 7) and other PII (2 and 5) are whole in both modes. For the record, neither shipped: float16 on Apple MPS equals float32 on both panels; FastPDN's INT8 ONNX on CPU hides 589 and 287 names whole with 1,009 and 935 person false characters. One reviewer labelled the set; the second pass is a same-day blind self-review of 100 passages (11 disagreed, adjudication changed 5), not independent agreement. `hybrid.json` `eval_names` holds these numbers and the score report's SHA.
30
 
31
- Release gates. The name test replaces the planned 200-document spot check: a spot check reviews only what the model masked, so it measures neither recall nor names missed on clean-looking text. Throughput, local: 200 seeded Dynaword documents (1,103,229 characters), Apple M4 Max, MPS, float32, `scrub_many`, one process, off and on alternating twice. Names off 7,914 and 7,922 chars/s, peak RSS 1,393 and 1,388 MiB; names on 7,444 and 7,597 chars/s (βˆ’5%), 1,679 and 1,682 MiB (+290 MiB). MPS driver memory at the end of a run is 70,261 MiB off and 87,788 MiB on; that is the allocator cache, not a working set. Repeat runs mask identically; names on gives 1,309 `[PERSON]` placeholders, names off none, and both give 7 `[PHONE]` and 7 `[PII]`. CUDA reference from the runtime study (RTX 4090, 1,000 Dynaword documents, 4,215,149 characters, a prototype of this runtime with NERGAL float16 and names float32): 3 processes 83,479 β†’ 72,662 chars/s (βˆ’13%), peak device memory 9,012 β†’ 12,602 MiB; per process, torch peak 1,806 β†’ 2,668 MiB. A float16 names model changed spans in 11 of those 1,000 documents, so it is not offered.
32
 
33
  Known limits: names-on masks every person the policy covers, public officials and historical figures included, which is why it is opt-in. On the random panel 18 of 116 components keep part of a name, and 34 of 300 passages have a false person mask. The names model was fine-tuned partly on LLM-synthetic data that SlayerLab has not audited.
34
 
 
28
 
29
  Names off hides no name; its false characters are 1 (random) and 29 (targeted). The set's phones (2 and 7) and other PII (2 and 5) are whole in both modes. For the record, neither shipped: float16 on Apple MPS equals float32 on both panels; FastPDN's INT8 ONNX on CPU hides 589 and 287 names whole with 1,009 and 935 person false characters. One reviewer labelled the set; the second pass is a same-day blind self-review of 100 passages (11 disagreed, adjudication changed 5), not independent agreement. `hybrid.json` `eval_names` holds these numbers and the score report's SHA.
30
 
31
+ Release gates. The name test replaces the planned 200-document spot check: a spot check reviews only what the model masked, so it measures neither recall nor names missed on clean-looking text. Throughput, local: 200 seeded Dynaword documents (1,103,229 characters), Apple M4 Max, MPS, float32, `scrub_many`, one process, off and on alternating twice. Names off 7,914 and 7,922 chars/s, peak RSS 1,393 and 1,388 MiB; names on 7,444 and 7,597 chars/s (βˆ’5%), 1,679 and 1,682 MiB (+290 MiB). MPS driver memory at the end of a run is 70,261 MiB off and 87,788 MiB on; that is the allocator cache, not a working set. Repeat runs mask identically; names on gives 1,309 `[PERSON]` placeholders, names off none, and both give 7 `[PHONE]` and 7 `[PII]`. CUDA, the shipped release: the `v2.0.0` Hub snapshot loaded on an RTX 4090 pod (driver 570.172.08, torch 2.8.0+cu128), the runtime study's 1,000 Dynaword documents (4,215,149 characters), NERGAL float16 and names float32, off and on interleaved, twice. One process 38,654 β†’ 35,108 chars/s (βˆ’9%); 3 processes 74,520 β†’ 66,492 chars/s (βˆ’11%), peak device memory 9,006 β†’ 15,416 MiB; per process, torch peak 1,806 β†’ 2,668 MiB. Repeats and 1 vs 3 processes mask identically; names on changes 666 of the 1,000 documents (6,209 `[PERSON]`). This host was slower than the runtime study's (names off, 3 processes: 74,520 vs 83,479 chars/s), so compare within one host. A float16 names model changed spans in 11 of the 1,000 documents in the runtime study, so it is not offered.
32
 
33
  Known limits: names-on masks every person the policy covers, public officials and historical figures included, which is why it is opt-in. On the random panel 18 of 116 components keep part of a name, and 34 of 300 passages have a false person mask. The names model was fine-tuned partly on LLM-synthetic data that SlayerLab has not audited.
34
 
README.md CHANGED
@@ -29,7 +29,7 @@ Python rules do the identifiers they can prove. A transformer NER head adds phon
29
  - **Ground:** `scrub_pii.py` rules
30
  - **Additive labels:** XLM-RoBERTa-large token classifier (epoch 5 of 7), BIO tags `phone` / `pii`, threshold 0.95
31
  - **Names (opt-in):** [FastPDN NER β€” Polish PII](https://huggingface.co/ArkadiuszPawlak/fastpdn-ner-polish-pii) (CC-BY-4.0) in `names/`, float32
32
- - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes; names cost about 13%
33
  - **Changes:** `CHANGELOG.md`
34
 
35
  ## What NERGAL detects β€” and what it does not
@@ -97,7 +97,7 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
97
 
98
  Names off hides no name. A component is the independent unit, covered when every name in it is whole: on the random panel 18 of 116 keep part of a name. The second pass is a same-day blind self-review of 100 passages by the same reviewer (11 disagreed), not independent agreement. `hybrid.json` `eval_names` has the full numbers, including float16 and INT8 scored for the record; neither ships.
99
 
100
- **Cost:** on one RTX 4090 with 3 processes, names took a prototype of this runtime from 83,479 to 72,662 chars/s (βˆ’13%) and peak device memory from 9,012 to 12,602 MiB. Local measurements are in `CHANGELOG.md`.
101
 
102
  ## Cue-less phone test
103
 
 
29
  - **Ground:** `scrub_pii.py` rules
30
  - **Additive labels:** XLM-RoBERTa-large token classifier (epoch 5 of 7), BIO tags `phone` / `pii`, threshold 0.95
31
  - **Names (opt-in):** [FastPDN NER β€” Polish PII](https://huggingface.co/ArkadiuszPawlak/fastpdn-ner-polish-pii) (CC-BY-4.0) in `names/`, float32
32
+ - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes; names cost about 11%
33
  - **Changes:** `CHANGELOG.md`
34
 
35
  ## What NERGAL detects β€” and what it does not
 
97
 
98
  Names off hides no name. A component is the independent unit, covered when every name in it is whole: on the random panel 18 of 116 keep part of a name. The second pass is a same-day blind self-review of 100 passages by the same reviewer (11 disagreed), not independent agreement. `hybrid.json` `eval_names` has the full numbers, including float16 and INT8 scored for the record; neither ships.
99
 
100
+ **Cost:** on one RTX 4090 with 3 processes, names took 2.0.0 from 74,520 to 66,492 chars/s (βˆ’11%) and peak device memory from 9,006 to 15,416 MiB; with one process, βˆ’9%. Local measurements are in `CHANGELOG.md`.
101
 
102
  ## Cue-less phone test
103