AstrLink Guard
AstrLink Guard is a lightweight model for detecting personal information and credentials in Chinese, English and mixed Chinese–English text, including source code, configuration files and logs. It runs fully locally on CPU with ONNX Runtime and returns entity spans with labels and scores, so the calling application decides what to mask.
It is the privacy detection model used by AstrLink.
Highlights
- 10 entity types, from names, phones and addresses to API keys, tokens, cookies and keys.
- Bilingual: Chinese, English and code-mixed text, including code and logs.
- Small and fast: 37M parameters. The INT8 variant processes a 510-token window in about 67 ms (p95) on 2 CPU threads, with about 171 MiB peak memory.
- Long inputs: sliding windows cover requests up to 131,072 tokens.
- Local only: no network access at inference time.
Model details
| Version | 0.1.0 |
| Architecture | BertForTokenClassification, 12 layers, hidden size 384 |
| Parameters | 36.9M |
| Base model | microsoft/Multilingual-MiniLM-L12-H384 |
| Tokenizer | XLM-R SentencePiece Unigram, 40,003 tokens |
| Languages | Chinese, English |
| Window | 512 tokens (128-token overlap for long inputs) |
| Formats | ONNX FP32, ONNX INT8 |
| License | Apache-2.0 |
Variants
| Folder | Format | Size | general F1 | code F1 | p95 latency, 510 tokens | Peak RSS |
|---|---|---|---|---|---|---|
fp32/ |
FP32 | 141 MiB | 96.49 | 88.97 | 130–138 ms | 246 MiB |
int8/ |
Dynamic INT8 | 35 MiB | 96.48 | 87.91 | 67–68 ms | 171 MiB |
Use fp32/ for the best accuracy, especially on code. Use int8/ when latency or memory matters more. Latency and memory are measured on an Apple M5 Pro CPU with 2 intra-op threads, including tokenization and decoding.
Each folder contains the ONNX model with its external data file, config.json, tokenizer.json, tokenizer_config.json, labels.json, label_mapping.json, CHECKSUMS.json and astrlink-model.json (the AstrLink import descriptor).
Entity types
BIO tagging, 21 tags.
| Label | Description |
|---|---|
private_person |
Person name |
phone |
Phone number (mobile, or landline with area code) |
email |
Email address |
private_address |
Physical address (home, shipping, office) |
private_date |
Date, time of day or timestamp |
account |
Personal account or service number: user name, QQ / WeChat ID, member or employee number, non-bank stored-value or campus card number |
payment_card |
Bank or credit card number |
ip_address |
IP address (without port) |
url |
URL |
common_secret |
Password, PIN, API key, access token, cookie key=value, private key, public key |
Not labeled as entities: order, ticket, tracking and transaction numbers; amounts, counts, durations and version numbers; hashes, checksums and request IDs; key fingerprints and signatures; OAuth client IDs; short internal numeric IDs.
Quick start
pip install onnxruntime tokenizers numpy huggingface_hub
import json
import numpy as np
import onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer
d = snapshot_download("QuantumNous/astrlink-guard", allow_patterns=["fp32/*"]) + "/fp32"
onnx_file = "model.onnx" # for the INT8 variant: allow_patterns=["int8/*"], "/int8", "model_int8.onnx"
tok = Tokenizer.from_file(f"{d}/tokenizer.json")
id2label = json.load(open(f"{d}/labels.json"))["id2label"]
sess = ort.InferenceSession(f"{d}/{onnx_file}", providers=["CPUExecutionProvider"])
def detect(text):
enc = tok.encode(text) # one window: at most 512 tokens including <s> and </s>
ids = np.array([enc.ids], dtype=np.int64)
feed = {"input_ids": ids, "attention_mask": np.ones_like(ids), "token_type_ids": np.zeros_like(ids)}
logits = sess.run(["logits"], feed)[0][0]
probs = np.exp(logits - logits.max(-1, keepdims=True))
probs /= probs.sum(-1, keepdims=True)
spans, cur = [], None
for label_id, p, (start, end) in zip(probs.argmax(-1), probs.max(-1), enc.offsets):
if start == end: # special tokens
continue
tag, _, kind = id2label[str(label_id)].partition("-")
if tag == "O":
cur = None
elif tag == "B" or cur is None or cur["label"] != kind:
cur = {"label": kind, "start": start, "end": end, "score": float(p)}
spans.append(cur)
else:
cur["end"] = end
cur["score"] = min(cur["score"], float(p))
for s in spans:
s["text"] = text[s["start"]:s["end"]]
return spans
print(detect("请联系张伟,电话 13812345678,邮箱 zhangwei@example.com,收货地址:北京市海淀区中关村大街 27 号。"))
Example results (FP32; all values are made up):
请联系张伟,电话 13812345678,邮箱 zhangwei@example.com,收货地址:北京市海淀区中关村大街 27 号。
private_person 张伟 0.9960
phone 13812345678 0.9953
email zhangwei@example.com 0.9966
private_address 北京市海淀区中关村大街 27 号 0.9954
Contact Jane Doe at jane.doe@example.org or +1 415-555-0132. Server 10.0.3.17, api_key = "q7Vt2LmX9pRk4sWz8bNc3yHd6fJg1aE5"
private_person Jane Doe 0.9956
email jane.doe@example.org 0.9968
phone +1 415-555-0132 0.9947
ip_address 10.0.3.17 0.9747
common_secret q7Vt2LmX9pRk4sWz8bNc3yHd6fJg1aE5 0.9943
订单号 202409301234567890,金额 128.00 元,已于 2026-09-30 14:05 发货。
private_date 2026-09-30 14:05 0.9917
Offsets are Unicode character positions in the input string.
Long inputs
The example above handles one 512-token window. For longer text, run 512-token windows with a 128-token overlap and merge the spans. AstrLink accepts up to 131,072 tokens per request and rejects longer input explicitly instead of truncating it.
The metrics below use the decoding rules recorded in config.json: BIO decoding constrained by character offsets (bio-offset-consistency-v1), plus boundary cleanup for IP addresses and URLs (existing-ip-and-dynamic-url-v1). Plain argmax decoding, as in the example, may differ slightly. No score threshold is applied; filter by score as your application needs.
Evaluation
Evaluated on in-house benchmarks. P / R / F1 are entity-level exact-match percentages; 95% confidence intervals come from document-level bootstrap.
General text (2,200 documents: 840 English, 840 Chinese, 520 mixed)
| Entities | FP32 P | FP32 R | FP32 F1 | INT8 F1 | |
|---|---|---|---|---|---|
| All | 4,336 | 95.98 | 97.00 | 96.49 [95.79, 97.13] | 96.48 |
| English | 1,730 | 93.23 | 94.80 | 94.01 | 94.43 |
| Chinese | 1,694 | 97.19 | 97.87 | 97.53 | 97.20 |
| Mixed | 912 | 99.02 | 99.56 | 99.29 | 99.07 |
| Entity type | Entities | P | R | F1 |
|---|---|---|---|---|
| url | 360 | 100 | 100 | 100 |
| 407 | 99.26 | 99.51 | 99.39 | |
| ip_address | 315 | 98.75 | 100 | 99.37 |
| payment_card | 319 | 99.68 | 98.75 | 99.21 |
| phone | 446 | 98.43 | 98.43 | 98.43 |
| private_person | 507 | 95.02 | 97.83 | 96.40 |
| common_secret | 352 | 94.51 | 97.73 | 96.09 |
| private_address | 365 | 91.79 | 98.08 | 94.83 |
| private_date | 751 | 93.34 | 93.34 | 93.34 |
| account | 514 | 93.48 | 92.02 | 92.75 |
FP32. Macro F1 over the 10 types: 96.98. False positives on documents with no entities: 2 of 662 (0.3%).
Code, configs and logs (709 documents)
| Entities | FP32 P | FP32 R | FP32 F1 | INT8 F1 | |
|---|---|---|---|---|---|
| All | 1,499 | 88.62 | 89.33 | 88.97 [86.77, 90.73] | 87.91 |
| Code only | 712 | 89.43 | 90.31 | 89.87 | 88.61 |
| English | 763 | 88.12 | 88.47 | 88.29 | 87.46 |
FP32 by type: url 95.47, common_secret 87.90, account 85.90, ip_address 100, private_date 96.30, private_person 94.74. False positives on documents with no entities: 13 of 261 (5.0%).
Additional benchmarks
| Benchmark | Documents | FP32 F1 | INT8 F1 |
|---|---|---|---|
| PII-600 (mixed Chinese / English PII) | 600 | 98.91 | 98.34 |
| Code-303 (secrets in code) | 303 | 79.62 | 77.55 |
Performance
Apple M5 Pro, CPU, 2 intra-op threads, one request at a time; includes tokenization and decoding.
| Peak RSS | 510 tokens, p50 / p95 | 8K tokens | 32K tokens | 128K tokens | |
|---|---|---|---|---|---|
| FP32 | 246 MiB | 128–136 / 130–138 ms | 2.8–2.9 s | 11.2–11.7 s | 46.3–46.9 s |
| INT8 | 171 MiB | 64–65 / 67–68 ms | 1.4 s | 5.5–5.6 s | 22.1–22.6 s |
Limitations
- Like any detector, the model can miss entities or flag harmless values. Use it as one layer of privacy protection, not as a guarantee.
- Code is harder than prose: about 5% of entity-free code documents get at least one false positive, and
accountis the weakest entity type in code. - Card numbers from issuers not seen in training, such as transit, membership or badge numbers, may be labeled
payment_cardinstead ofaccount. - Configuration values whose key name ends in
keymay be labeledcommon_secreteven when they are not secrets. - Government ID numbers (such as national ID cards) have no dedicated entity type.
- Quality is measured on Chinese and English only.
Training data
Public datasets and repositories used in training:
- NVIDIA Nemotron-PII (CC BY 4.0)
- Wismut/nym-pii-multilingual-data (MIT)
- Samsung CredData (Apache-2.0) and the open-source repositories it references
- Other open-source code and documentation, including django/django (BSD-3-Clause), dromara/hutool (MulanPSL-2.0) and google/eng-practices (CC BY 3.0)
The full list with revisions and license texts is in licenses/. No training data is included in this repository.
License
The model is released under the Apache License 2.0. The base model, microsoft/Multilingual-MiniLM-L12-H384, is MIT-licensed.
Attribution:
- Contains data derived from NVIDIA Nemotron-PII, licensed under CC BY 4.0. The data was modified.
- Contains data derived from google/eng-practices, licensed under CC BY 3.0. The data was modified.
- Contains data derived from Django (BSD-3-Clause). This model is not endorsed by the Django project or its contributors.
Full license texts and notices for the base model, datasets and source repositories are in licenses/. SHA256SUMS lists the sha256 of every file in this repository.
中文说明
AstrLink Guard 是一个轻量级的隐私与凭证检测模型,支持中文、英文和中英混排文本,也覆盖源代码、配置文件和日志。模型用 ONNX Runtime 在本地 CPU 上运行,输出实体片段的位置、类别和分数,是否遮蔽由调用方决定。它是 AstrLink 使用的隐私检测模型。
特点
- 10 类实体:从人名、电话、地址,到 API key、token、Cookie、密钥。
- 中英双语:中文、英文、中英混排,以及代码、日志。
- 小而快:3,700 万参数;INT8 版在 2 个 CPU 线程上处理 510 token 约 67 ms(p95),峰值内存约 171 MiB。
- 长文本:用滑动窗口处理,单次最长 131,072 token。
- 完全本地:推理时不联网。
两种格式
fp32/:精度最好,代码场景推荐用它。int8/:体积 35 MiB,速度约为 FP32 的 2 倍,内存更低;通用文本 F1 与 FP32 基本持平,代码场景约低 1 个点。
实体类别
人名、电话、邮箱、物理地址、日期时间、个人账号、银行卡、IP、URL、凭证(密码、PIN、API key、token、Cookie、私钥、公钥),BIO 共 21 个标签。
订单号、快递号、交易号、金额、版本号、哈希、请求 ID、OAuth client_id 等不属于实体。
用法
见上面的 Quick start。示例只处理一个 512 token 的窗口;长文本按 512 token 窗口、128 token 重叠切分,再合并结果。
效果
| 评测集 | FP32 F1 | INT8 F1 |
|---|---|---|
| 通用文本(2,200 篇,中 / 英 / 混排) | 96.49 | 96.48 |
| 代码、配置、日志(709 篇) | 88.97 | 87.91 |
| PII-600 | 98.91 | 98.34 |
| Code-303 | 79.62 | 77.55 |
通用文本上:中文 F1 97.53,英文 94.01,混排 99.29;不含实体的文档误报率为 0.3%。
局限
- 检测模型都会有漏检和误报,请把它当作隐私保护的一层,不要当作保证。
- 代码比普通文本难:不含实体的代码文档里约 5% 会出现误报;代码中的账号识别最弱。
- 训练中没见过的发卡方(交通卡、会员卡、工牌等)的卡号,可能被判成银行卡。
- 键名以 key 结尾的配置值,即使不是凭证,也可能被判成凭证。
- 身份证等政府证件号没有专门的类别。
许可
模型采用 Apache-2.0;底座模型 MiniLM 为 MIT。训练用到的公开数据包括 NVIDIA Nemotron-PII(CC BY 4.0)、google/eng-practices(CC BY 3.0)、Django(BSD-3-Clause)等,署名见上面 License 一节,许可原文见 licenses/。
Model tree for QuantumNous/astrlink-guard
Base model
microsoft/Multilingual-MiniLM-L12-H384