AstrLink Guard

AstrLink Guard is a lightweight model for detecting personal information and credentials in Chinese, English and mixed Chinese–English text, including source code, configuration files and logs. It runs fully locally on CPU with ONNX Runtime and returns entity spans with labels and scores, so the calling application decides what to mask.

It is the privacy detection model used by AstrLink.

中文说明

Highlights

  • 10 entity types, from names, phones and addresses to API keys, tokens, cookies and keys.
  • Bilingual: Chinese, English and code-mixed text, including code and logs.
  • Small and fast: 37M parameters. The INT8 variant processes a 510-token window in about 67 ms (p95) on 2 CPU threads, with about 171 MiB peak memory.
  • Long inputs: sliding windows cover requests up to 131,072 tokens.
  • Local only: no network access at inference time.

Model details

Version 0.1.0
Architecture BertForTokenClassification, 12 layers, hidden size 384
Parameters 36.9M
Base model microsoft/Multilingual-MiniLM-L12-H384
Tokenizer XLM-R SentencePiece Unigram, 40,003 tokens
Languages Chinese, English
Window 512 tokens (128-token overlap for long inputs)
Formats ONNX FP32, ONNX INT8
License Apache-2.0

Variants

Folder Format Size general F1 code F1 p95 latency, 510 tokens Peak RSS
fp32/ FP32 141 MiB 96.49 88.97 130–138 ms 246 MiB
int8/ Dynamic INT8 35 MiB 96.48 87.91 67–68 ms 171 MiB

Use fp32/ for the best accuracy, especially on code. Use int8/ when latency or memory matters more. Latency and memory are measured on an Apple M5 Pro CPU with 2 intra-op threads, including tokenization and decoding.

Each folder contains the ONNX model with its external data file, config.json, tokenizer.json, tokenizer_config.json, labels.json, label_mapping.json, CHECKSUMS.json and astrlink-model.json (the AstrLink import descriptor).

Entity types

BIO tagging, 21 tags.

Label Description
private_person Person name
phone Phone number (mobile, or landline with area code)
email Email address
private_address Physical address (home, shipping, office)
private_date Date, time of day or timestamp
account Personal account or service number: user name, QQ / WeChat ID, member or employee number, non-bank stored-value or campus card number
payment_card Bank or credit card number
ip_address IP address (without port)
url URL
common_secret Password, PIN, API key, access token, cookie key=value, private key, public key

Not labeled as entities: order, ticket, tracking and transaction numbers; amounts, counts, durations and version numbers; hashes, checksums and request IDs; key fingerprints and signatures; OAuth client IDs; short internal numeric IDs.

Quick start

pip install onnxruntime tokenizers numpy huggingface_hub
import json

import numpy as np
import onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer

d = snapshot_download("QuantumNous/astrlink-guard", allow_patterns=["fp32/*"]) + "/fp32"
onnx_file = "model.onnx"  # for the INT8 variant: allow_patterns=["int8/*"], "/int8", "model_int8.onnx"

tok = Tokenizer.from_file(f"{d}/tokenizer.json")
id2label = json.load(open(f"{d}/labels.json"))["id2label"]
sess = ort.InferenceSession(f"{d}/{onnx_file}", providers=["CPUExecutionProvider"])


def detect(text):
    enc = tok.encode(text)  # one window: at most 512 tokens including <s> and </s>
    ids = np.array([enc.ids], dtype=np.int64)
    feed = {"input_ids": ids, "attention_mask": np.ones_like(ids), "token_type_ids": np.zeros_like(ids)}
    logits = sess.run(["logits"], feed)[0][0]
    probs = np.exp(logits - logits.max(-1, keepdims=True))
    probs /= probs.sum(-1, keepdims=True)
    spans, cur = [], None
    for label_id, p, (start, end) in zip(probs.argmax(-1), probs.max(-1), enc.offsets):
        if start == end:  # special tokens
            continue
        tag, _, kind = id2label[str(label_id)].partition("-")
        if tag == "O":
            cur = None
        elif tag == "B" or cur is None or cur["label"] != kind:
            cur = {"label": kind, "start": start, "end": end, "score": float(p)}
            spans.append(cur)
        else:
            cur["end"] = end
            cur["score"] = min(cur["score"], float(p))
    for s in spans:
        s["text"] = text[s["start"]:s["end"]]
    return spans


print(detect("请联系张伟,电话 13812345678,邮箱 zhangwei@example.com,收货地址:北京市海淀区中关村大街 27 号。"))

Example results (FP32; all values are made up):

请联系张伟,电话 13812345678,邮箱 zhangwei@example.com,收货地址:北京市海淀区中关村大街 27 号。
  private_person   张伟                               0.9960
  phone            13812345678                        0.9953
  email            zhangwei@example.com               0.9966
  private_address  北京市海淀区中关村大街 27 号        0.9954

Contact Jane Doe at jane.doe@example.org or +1 415-555-0132. Server 10.0.3.17, api_key = "q7Vt2LmX9pRk4sWz8bNc3yHd6fJg1aE5"
  private_person   Jane Doe                           0.9956
  email            jane.doe@example.org               0.9968
  phone            +1 415-555-0132                    0.9947
  ip_address       10.0.3.17                          0.9747
  common_secret    q7Vt2LmX9pRk4sWz8bNc3yHd6fJg1aE5   0.9943

订单号 202409301234567890,金额 128.00 元,已于 2026-09-30 14:05 发货。
  private_date     2026-09-30 14:05                   0.9917

Offsets are Unicode character positions in the input string.

Long inputs

The example above handles one 512-token window. For longer text, run 512-token windows with a 128-token overlap and merge the spans. AstrLink accepts up to 131,072 tokens per request and rejects longer input explicitly instead of truncating it.

The metrics below use the decoding rules recorded in config.json: BIO decoding constrained by character offsets (bio-offset-consistency-v1), plus boundary cleanup for IP addresses and URLs (existing-ip-and-dynamic-url-v1). Plain argmax decoding, as in the example, may differ slightly. No score threshold is applied; filter by score as your application needs.

Evaluation

Evaluated on in-house benchmarks. P / R / F1 are entity-level exact-match percentages; 95% confidence intervals come from document-level bootstrap.

General text (2,200 documents: 840 English, 840 Chinese, 520 mixed)

Entities FP32 P FP32 R FP32 F1 INT8 F1
All 4,336 95.98 97.00 96.49 [95.79, 97.13] 96.48
English 1,730 93.23 94.80 94.01 94.43
Chinese 1,694 97.19 97.87 97.53 97.20
Mixed 912 99.02 99.56 99.29 99.07
Entity type Entities P R F1
url 360 100 100 100
email 407 99.26 99.51 99.39
ip_address 315 98.75 100 99.37
payment_card 319 99.68 98.75 99.21
phone 446 98.43 98.43 98.43
private_person 507 95.02 97.83 96.40
common_secret 352 94.51 97.73 96.09
private_address 365 91.79 98.08 94.83
private_date 751 93.34 93.34 93.34
account 514 93.48 92.02 92.75

FP32. Macro F1 over the 10 types: 96.98. False positives on documents with no entities: 2 of 662 (0.3%).

Code, configs and logs (709 documents)

Entities FP32 P FP32 R FP32 F1 INT8 F1
All 1,499 88.62 89.33 88.97 [86.77, 90.73] 87.91
Code only 712 89.43 90.31 89.87 88.61
English 763 88.12 88.47 88.29 87.46

FP32 by type: url 95.47, common_secret 87.90, account 85.90, ip_address 100, private_date 96.30, private_person 94.74. False positives on documents with no entities: 13 of 261 (5.0%).

Additional benchmarks

Benchmark Documents FP32 F1 INT8 F1
PII-600 (mixed Chinese / English PII) 600 98.91 98.34
Code-303 (secrets in code) 303 79.62 77.55

Performance

Apple M5 Pro, CPU, 2 intra-op threads, one request at a time; includes tokenization and decoding.

Peak RSS 510 tokens, p50 / p95 8K tokens 32K tokens 128K tokens
FP32 246 MiB 128–136 / 130–138 ms 2.8–2.9 s 11.2–11.7 s 46.3–46.9 s
INT8 171 MiB 64–65 / 67–68 ms 1.4 s 5.5–5.6 s 22.1–22.6 s

Limitations

  • Like any detector, the model can miss entities or flag harmless values. Use it as one layer of privacy protection, not as a guarantee.
  • Code is harder than prose: about 5% of entity-free code documents get at least one false positive, and account is the weakest entity type in code.
  • Card numbers from issuers not seen in training, such as transit, membership or badge numbers, may be labeled payment_card instead of account.
  • Configuration values whose key name ends in key may be labeled common_secret even when they are not secrets.
  • Government ID numbers (such as national ID cards) have no dedicated entity type.
  • Quality is measured on Chinese and English only.

Training data

Public datasets and repositories used in training:

The full list with revisions and license texts is in licenses/. No training data is included in this repository.

License

The model is released under the Apache License 2.0. The base model, microsoft/Multilingual-MiniLM-L12-H384, is MIT-licensed.

Attribution:

  • Contains data derived from NVIDIA Nemotron-PII, licensed under CC BY 4.0. The data was modified.
  • Contains data derived from google/eng-practices, licensed under CC BY 3.0. The data was modified.
  • Contains data derived from Django (BSD-3-Clause). This model is not endorsed by the Django project or its contributors.

Full license texts and notices for the base model, datasets and source repositories are in licenses/. SHA256SUMS lists the sha256 of every file in this repository.


中文说明

AstrLink Guard 是一个轻量级的隐私与凭证检测模型,支持中文、英文和中英混排文本,也覆盖源代码、配置文件和日志。模型用 ONNX Runtime 在本地 CPU 上运行,输出实体片段的位置、类别和分数,是否遮蔽由调用方决定。它是 AstrLink 使用的隐私检测模型。

特点

  • 10 类实体:从人名、电话、地址,到 API key、token、Cookie、密钥。
  • 中英双语:中文、英文、中英混排,以及代码、日志。
  • 小而快:3,700 万参数;INT8 版在 2 个 CPU 线程上处理 510 token 约 67 ms(p95),峰值内存约 171 MiB。
  • 长文本:用滑动窗口处理,单次最长 131,072 token。
  • 完全本地:推理时不联网。

两种格式

  • fp32/:精度最好,代码场景推荐用它。
  • int8/:体积 35 MiB,速度约为 FP32 的 2 倍,内存更低;通用文本 F1 与 FP32 基本持平,代码场景约低 1 个点。

实体类别

人名、电话、邮箱、物理地址、日期时间、个人账号、银行卡、IP、URL、凭证(密码、PIN、API key、token、Cookie、私钥、公钥),BIO 共 21 个标签。

订单号、快递号、交易号、金额、版本号、哈希、请求 ID、OAuth client_id 等不属于实体。

用法

见上面的 Quick start。示例只处理一个 512 token 的窗口;长文本按 512 token 窗口、128 token 重叠切分,再合并结果。

效果

评测集 FP32 F1 INT8 F1
通用文本(2,200 篇,中 / 英 / 混排) 96.49 96.48
代码、配置、日志(709 篇) 88.97 87.91
PII-600 98.91 98.34
Code-303 79.62 77.55

通用文本上:中文 F1 97.53,英文 94.01,混排 99.29;不含实体的文档误报率为 0.3%。

局限

  • 检测模型都会有漏检和误报,请把它当作隐私保护的一层,不要当作保证。
  • 代码比普通文本难:不含实体的代码文档里约 5% 会出现误报;代码中的账号识别最弱。
  • 训练中没见过的发卡方(交通卡、会员卡、工牌等)的卡号,可能被判成银行卡。
  • 键名以 key 结尾的配置值,即使不是凭证,也可能被判成凭证。
  • 身份证等政府证件号没有专门的类别。

许可

模型采用 Apache-2.0;底座模型 MiniLM 为 MIT。训练用到的公开数据包括 NVIDIA Nemotron-PII(CC BY 4.0)、google/eng-practices(CC BY 3.0)、Django(BSD-3-Clause)等,署名见上面 License 一节,许可原文见 licenses/。

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for QuantumNous/astrlink-guard

Quantized
(9)
this model