File size: 14,624 Bytes
bc39933
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
---
language:
- ko
- en
license: apache-2.0
library_name: transformers
pipeline_tag: token-classification
base_model: openai/privacy-filter
tags:
- token-classification
- ner
- pii
- privacy
- pii-masking
- korean
- finance
- bioes
- viterbi
- mixture-of-experts
datasets:
- BCCard/pii-masking-openpii-finance
metrics:
- f1
- precision
- recall
---

# 1. Overview
A Korean/English **PII detection model for the finance domain**, built by full fine-tuning
[`openai/privacy-filter`](https://huggingface.co/openai/privacy-filter) (1.4B MoE, 50M active) on synthetic finance-domain PII data. It tags **18 PII entity types** (73 BIOES classes) at token level and is intended as the **NER layer of a multi-layer PII-masking gateway** in front of LLM services —
behind a regex backstop for fully structured identifiers, never as a standalone compliance guarantee.

On held-out validation it reaches **strict span-F1 0.956 (ko) / 0.969 (en)**. On an independent, adversarially-hardened Golden Set it holds **0.944 (ko) / 0.907 (en)** with **masking coverage 0.996 (ko) / 0.998 (en)** — i.e. ≥99.5% of gold PII characters are covered by predicted spans.

## 1.1. TL;DR
* **Base model**: [`openai/privacy-filter`](https://huggingface.co/openai/privacy-filter) — 1.4B-parameter MoE (128 experts, 50M active), 8 layers, hidden 640, bidirectional banded attention (±128), o200k tokenizer
* **Domain / Language**: Finance (BC Card — cards, accounts, national IDs, customer service text) / Korean + English
* **Task**: Token classification (BIOES) → character-offset PII spans → masking
* **Labels (18)**: `PERSON, RRN, FRN, CARD_NUMBER, ACCOUNT_NUMBER, SECRET, USER_ID, EMAIL, PHONE, PASSPORT, DRIVER_LICENSE, GENERIC_ID, ADDRESS, ZIPCODE, DATE, CARD_EXPIRY, CVC, IPIN`
* **Method**: Full fine-tuning (all parameters incl. experts & router) with a re-initialized 73-class head (rows copied from the base head by taxonomy mapping)
* **Decoding**: **constrained BIOES Viterbi** (not per-token argmax) + whitespace span refinement — the bundled `viterbi_calibration.json` exposes precision↔recall operating-point biases without retraining
* **Format**: BF16 (attention `sinks` kept FP32), single safetensors + tokenizer + label taxonomy + Viterbi calibration sidecar
* **Sequence length**: trained on sequences ≤768 tokens — chunk longer inputs
* **Intended use**
  - In-house **PII masking gateway** (detect → mask before text reaches an LLM)
  - Korean-centric finance text with mixed English (IDs, e-mails, card numbers)

## 1.2. Label Taxonomy (N=18)
The 18 labels re-map the upstream ai4privacy source labels to the granularity a Korean financial masking policy needs - merging fragments into single spans (`GIVENNAME`/`SURNAME``PERSON`, `CITY`/`STREET`/`BUILDINGNUM``ADDRESS`) and adding Korea-specific classes absent upstream (`RRN`, `FRN`, `IPIN`, `CARD_EXPIRY`, `CVC`, `SECRET`). `data source` records the row-source buckets in which each label occurs: `ko` means `openpii-1.5m-ko`, `en` means `openpii-1.5m-en`, and `domain` means locally synthesized rows.

| label | description | data source |
|-------|-------------|-------------|
| `PERSON` | full name (surname + given, single span) | ko, en, domain |
| `RRN` | resident registration number (Korea) | ko, domain |
| `FRN` | foreign registration number | domain |
| `CARD_NUMBER` | credit/debit card PAN | ko, en, domain |
| `ACCOUNT_NUMBER` | bank account number | ko, domain |
| `SECRET` | auth secret (password / API key / token) | ko, domain |
| `USER_ID` | online member ID | ko, en, domain |
| `EMAIL` | email address | ko, en, domain |
| `PHONE` | phone number (mobile / landline) | ko, en, domain |
| `PASSPORT` | passport number | ko, en, domain |
| `DRIVER_LICENSE` | driver's license number | ko, en, domain |
| `GENERIC_ID` | generic identifier (no KO counterpart) | ko, en, domain |
| `ADDRESS` | address (city / street / building, single span) | ko, en, domain |
| `ZIPCODE` | postal code | ko, en, domain |
| `DATE` | date / time | ko, en, domain |
| `CARD_EXPIRY` | card expiry date | domain |
| `CVC` | card verification code | domain |
| `IPIN` | I-PIN number | domain |

Each entity type has `B-`, `I-`, `E-` and `S-` boundary classes, plus the background class `O`. This yields 73 output classes. The bundled `label-taxonomy.yaml` and `config.json` must remain in the same label order.

## 1.3. Usage

```python
import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer

model_id = "BCCard/MoAI-Privacy-Filter"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)
model.eval()

text = "고객 모아이님(000000-0000000)께서 010-0000-0000로 연락 요청하셨습니다."
enc = tokenizer(text, return_offsets_mapping=True, add_special_tokens=False, return_tensors="pt")
offsets = enc.pop("offset_mapping")[0].tolist()

with torch.no_grad():
    logits = model(**enc).logits[0]          # [T, 73] — raw output

# 1) Decode the logits with constrained BIOES Viterbi (recommended; see note below)
#    and map token paths to character spans via `offsets` — the detector's output:
# -> [{"start": 3, "end": 6, "label": "PERSON"},   '모아이'
#     {"start": 8, "end": 22, "label": "RRN"},     '000000-0000000'
#     {"start": 26, "end": 39, "label": "PHONE"}]  '010-0000-0000'
# 2) Masking is downstream application logic — replace each span according to
#    your masking policy, e.g.:
# -> "고객 [PERSON]님([RRN])께서 [PHONE]로 연락 요청하셨습니다."
```

> **Decoding note** — this model (like its base) is post-trained for **constrained Viterbi decoding**
> over the BIOES transition grammar, *not* independent per-token argmax. Argmax can emit invalid tag
> sequences (span splits / orphan tags) and measurably lowers span-F1. The bundled
> `viterbi_calibration.json` follows the upstream operating-point schema: its six transition biases
> shift the precision↔recall trade-off at deploy time without retraining (all `0.0` = neutral).

## 1.4. Training Data
| Dataset | Role | Size |
|---------|------|------|
| (Public) [BCCard/pii-masking-openpii-finance](https://huggingface.co/datasets/BCCard/pii-masking-openpii-finance) (v2) | Training / validation | ~58.5k train rows · ~14.5k validation rows |
| (Private) BCCard/pii-masking-openpii-finance-test (v2) | Golden Set (release gate; never used for training/tuning) | 2,000 rows (ko 1,460 / en 540) |

* Sources: curated Korean subset of `ai4privacy/pii-masking-openpii-1.5m` (label taxonomy remapped, name spans merged & naturalized) + finance-domain synthetic templates + **~30% English replay** (forgetting guard)
* Hard-example design baked into v2: surface-similar non-PII decoys (FP suppression), label-confusion pairs in one sentence (RRN↔FRN, DRIVER_LICENSE↔GENERIC_ID), weak-context true PII (FN suppression), long-span address boundary variants
* All values are synthetic; validity-pattern collisions with real identifiers are removed at generation time (e.g. card numbers are forced to fail Luhn)

## 1.5. Training Procedure
| Item | Value |
|------|-------|
| Method | Full fine-tuning (1.4B params — experts and router included) |
| Head | 33-class base head → 73-class head, initialized by copying base rows via taxonomy mapping |
| Loss | Token-level cross-entropy |
| Batch | effective 16 (per-device × world × accum), fixed across hardware layouts |
| LR / scheduler | 1e-4 / linear decay, warmup 3% |
| Optimizer | AdamW (fused), weight decay 0.0, max_grad_norm 1.0 |
| Epochs | 5 — best checkpoint by validation span micro-F1, **decoded with the same constrained Viterbi as deployment** |
| Precision | FP32 master weights + BF16 autocast; MoE router/experts explicitly kept FP32 during compute |
| Hardware | 1× NVIDIA H100 (~5h) |

<div align="center">
  <img src="figures/evaluation-train-1-1.png" alt="Training loss, learning-rate and gradient-norm curves for the v1 and v2 models" >
</div>

<div align="center">
  <img src="figures/evaluation-test-1-1.png" alt="Training-time validation metric curves for the v1 and v2 models" >
</div>

<br>

# 2. Evaluation
## 2.1. Setup
* **Golden Set**: independently generated 2,000-row test set (ko 1,460 / en 540), **adversarially hardened** — weak-context PII, decoys, confusion pairs and boundary variants are deliberately over-represented, so scores here read *lower* than typical in-distribution synthetic benchmarks
* **Protocol**: strict exact-match span P/R/F1 (CoNLL-style; boundary and label must both match) + **masking coverage** (share of gold PII *characters* covered by predicted spans, label-agnostic — the leakage-oriented metric)
* **Decoding**: constrained Viterbi + whitespace refinement — identical to the deployment chain

## 2.2. Results
### `validation` dataset

| Metric | v2 model / v2 validation |
|---|---:|
| micro F1 | 0.9599 |
| macro F1 | 0.9603 |
| ko strict micro F1 | 0.9562 |
| ko macro F1 | 0.9568 |
| en strict micro F1 | 0.9688 |
| en macro F1 | 0.9631 |
| **masking coverage** | **0.9979** |
| **ko masking coverage** | **0.9987** |
| **en masking coverage** | **0.9965** |

* These values were measured post-hoc by running the exported `final-bf16` artifact over all 14,543 v2 validation rows (ko 10,460 / en 4,083) through the deployment-equivalent chain: constrained Viterbi, actual tokenizer character offsets and whitespace refinement.
* Overall micro F1 and masking coverage pool all ko/en spans or characters before scoring. Overall macro F1 pools per-label TP/FP/FN across both languages and then averages the 18 label F1 values.
* Character coverage counts are ko **482,861 / 483,488** and en **264,979 / 265,898** gold PII characters.
* The training-time checkpoint-selection metrics remain ko micro F1 **0.9820**, en micro F1 **0.9739** and global macro F1 **0.9764** at epoch 5. They compare entity spans on token indices, so they are not interchangeable with the character-span values above and do not include masking coverage.

### `test` dataset
Independently generated Golden Set — deliberately harder than validation: weak-context PII, surface-similar decoys, label-confusion pairs and long-span boundary variants are over-represented.
**Δ = vs. the post-hoc validation baseline above** using the same `final-bf16` artifact and character-span evaluation chain. The difference measures test hardening and distribution shift, not model regression.

| Metric | v2 model / v2 test | Δ |
|---|---:|---:|
| micro F1 | 0.9336 | -2.63%p |
| macro F1 | 0.9308 | -2.94%p |
| ko strict micro F1 | 0.9441 | -1.21%p |
| ko macro F1 | 0.9416 | -1.53%p |
| en strict micro F1 | 0.9065 | -6.24%p |
| en macro F1 | 0.9017 | -6.14%p |
| **masking coverage** | **0.9964** | -0.16%p |
| **ko masking coverage** | **0.9956** | -0.31%p |
| **en masking coverage** | **0.9984** | +0.18%p |

* **Masking coverage stays ≥0.9956 on the adversarial set** — only 0.44% (ko) / 0.16% (en) of gold PII characters are uncovered; most strict-F1 losses are boundary or label-name errors, not leaks
* **English ADDRESS holds on hard boundary variants**: strict recall **0.983** on long-span address forms (state suffixes, unit/floor tails) that are heavily represented in this set
* **Weak-context person names are the main remaining leak channel**: ko `PERSON` strict recall 0.875 with 82 full-span misses (see Limitations)
* Label-swap errors (e.g. en `ACCOUNT_NUMBER` predicted as `GENERIC_ID`/`CARD_NUMBER`) keep **coverage 1.0** — the value is still masked; only the label name is wrong

## 2.3. Reading the numbers
Strict exact-match span-F1 on an adversarial test is a deliberately harsh score: a one-character boundary miss or a swapped label counts as a full error. For the product question — *"how much PII text leaks through?"* — masking coverage is the operative metric: **0.44% (ko) / 0.16% (en) of gold PII characters uncovered**, concentrated in weak-context person names.

<br>

## 2.4. Limitations
* **One layer of defense** — inherits the base model's positioning: not an anonymization or compliance guarantee. Deploy behind a regex backstop for fully structured identifiers (RRN patterns, card numbers, phones) and combine with policy-level controls.
* **Weak-context person names** — Korean names without honorifics/particles or list-form values are the main miss channel (ko `PERSON` recall 0.875 on the adversarial set). Consider a recall-leaning Viterbi operating point in high-sensitivity deployments.
* **Alphanumeric ID confusion**`USER_ID`/`SECRET`/`GENERIC_ID`/`ACCOUNT_NUMBER` share surface forms; without cue words the label may swap (masking still applies — coverage stays ~1.0).
* **Synthetic-only training & evaluation** — no real customer text was used or evaluated. Real-world robustness (typos, slang, OCR noise) is unvalidated; shadow-mode rollout is recommended before enforcement.
* **Fixed label policy** — the 18-label taxonomy is baked in at fine-tuning time; changing masking policy granularity requires re-fine-tuning (runtime keep/mask toggles must operate on these labels).
* **Context window** — banded attention limits each token's context to ±128 tokens; trained sequence regime is ≤768 tokens (chunk longer documents).

<br>

# 3. Future Work
* **v3 data** — weak-context person-name hard positives, cue-word diversification for the alphanumeric ID group, privacy-safe failure-collection loop from shadow operation
* **Operating point** - recall-leaning Viterbi transition biases tuned on validation or a dedicated calibration set without Golden Set feedback
* **Serving** - target-hardware latency, throughput and memory benchmarks for the separately published INT8 weight-only ONNX artifact

<br>

# 4. Meta Info
## 4.1. Citation
```bibtex
@misc{bccard2026moaiprivacyfilter,
  title        = {MoAI-Privacy-Filter: A Korean Finance-Domain PII Detection Model},
  author       = {BC Card AX Team},
  year         = {2026},
  howpublished = {https://huggingface.co/BCCard/MoAI-Privacy-Filter},
  note         = {Full fine-tune of openai/privacy-filter for Korean/English PII masking in the BC Card domain}
}
```

## 4.2. See Also
* **Base model**: [`openai/privacy-filter`](https://huggingface.co/openai/privacy-filter)
* **INT8 ONNX artifact**: [`BCCard/MoAI-Privacy-Filter-INT8`](https://huggingface.co/BCCard/MoAI-Privacy-Filter-INT8)
* **Training dataset**: [`BCCard/pii-masking-openpii-finance`](https://huggingface.co/datasets/BCCard/pii-masking-openpii-finance)
* **Source data attribution**: `ai4privacy/pii-masking-openpii-1.5m` (CC-BY-4.0)

<br>