Swisscoding Technologies

Swisscoding Technologies Report

Swisscoding Name Filter

Swisscoding Name Filter is a family of ModernBERT-base token classifiers for detecting personal names. Each model is deliberately focused on this single important use case to maximize performance. Choose the multilingual model or a monolingual model for English, German, French, or Italian.

Benchmarks

MultiGraSCCo contains 63 clinical texts per language, derived from a German corpus whose documents were manually anonymized and extensively altered. It was prepared by researchers from Technische Universität Berlin and the German Research Center for Artificial Intelligence (DFKI). Its German source annotations were manually labeled by human annotators; translated annotations were automatically preserved, and the translations were reviewed by medical professionals.

Nemotron PII is a synthetic PII dataset released by NVIDIA. Our local evaluation used the complete Nemotron PII test split: 100,000 records (50,000 US and 50,000 international). We evaluated only the English and multilingual Swisscoding Name Filter models on these records. NVIDIA describes these records as English-language; international refers to PII locale conventions, not additional document languages.

Model Precision Recall F1
MultiGraSCCo - Multilingual (German, French, Italian, English)
pii-name-filter-149M97.9298.7698.34
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v193.8584.3688.85
openai/privacy-filter79.4771.1975.10
MultiGraSCCo - English
pii-EN-name-filter-149M100.0099.7299.86
pii-name-filter-149M100.0099.1699.58
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v198.3994.8296.57
openai/privacy-filter97.7390.1593.79
MultiGraSCCo - German
pii-DE-name-filter-149M97.2099.1098.14
pii-name-filter-149M96.9299.4098.14
OpenMed/OpenMed-PII-German-SuperClinical-Large-434M-v198.0756.3671.58
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v193.1375.3583.31
openai/privacy-filter68.3566.0367.17
MultiGraSCCo - French
pii-FR-name-filter-149M99.3597.6898.51
pii-name-filter-149M97.4397.0497.24
OpenMed/OpenMed-PII-French-SuperClinical-Large-434M-v197.2144.5961.13
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v193.3085.9189.45
openai/privacy-filter78.8767.3772.67
MultiGraSCCo - Italian
pii-IT-name-filter-149M96.9099.5998.23
pii-name-filter-149M98.3199.5998.95
OpenMed/OpenMed-PII-Italian-SuperClinical-Large-434M-v195.2235.1451.34
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v191.6586.4388.97
openai/privacy-filter81.1668.5574.32
Nemotron PII
pii-EN-name-filter-149M94.5699.7097.06
pii-name-filter-149M93.9499.5896.68
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v1 first_name99.4899.5199.50
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v1 last_name99.4299.2999.35

For MultiGraSCCo, we combined OpenMed token predictions for first_name and last_name into one name category before calculating Precision, Recall, and F1.

We evaluated the MultiGraSCCo scores for OpenMed and Privacy Filter locally; they are not official results published by the respective organizations.

The OpenMed Nemotron PII first_name and last_name scores come from its published evaluation results, not our 100,000-record evaluation. OpenMed reports the two labels separately, while our models use a single PERSON label. The evaluation settings and label definitions differ, so these scores should not be compared directly.

Size

Total parameter count comparison for Swisscoding Name Filter, OpenMed SuperClinical, and OpenAI Privacy Filter

Total parameter count of the evaluated models.

* OpenAI Privacy Filter uses a mixture-of-experts (MoE) architecture and reports 50M active parameters out of 1.5B total parameters. Only total parameters are plotted.

For operational context, our separate throughput comparison used the Hugging Face Transformers implementations on an NVIDIA A100 in BF16 with 1,000 randomly selected Nemotron examples processed one at a time without batching. The English Swisscoding Name Filter achieved 39.00 examples/second, OpenMed SuperClinical Large achieved 22.13 examples/second, and OpenAI Privacy Filter achieved 0.29 examples/second (3.42 seconds/example). Throughput is implementation- and hardware-dependent, so we do not use it as the primary cross-model comparison.

How to use

Each model supports input windows of up to 8,192 tokens. For longer documents, split the text into overlapping chunks.

We evaluated the Swisscoding Name Filter models using a 0.5 decision threshold. Treat this as a practical starting point and tune it on representative validation data to balance missed names against false positives.

from transformers import pipeline

model_id = "Swisscoding-Technologies/pii-IT-name-filter-149M"
threshold = 0.5

name_detector = pipeline(
    "token-classification",
    model=model_id,
    aggregation_strategy="simple",
)

text = "La paziente Alice Rossi è stata inviata dal Dr Marco Weber per un controllo."
entities = name_detector(text)
names = [entity for entity in entities if entity["score"] >= threshold]
print(names)

Intended use

This model is intended to detect personal names in Italian text, especially medical and medical-adjacent documents. It is designed as one component of a broader de-identification workflow and should be validated on representative local data.

Training data and privacy

Training started from de-identified real-world medical documents in Italian, French, and German. We added structured placeholders and synthetic personas, then used Qwen3.5-122B-A10B to translate the documents across English, German, French, and Italian, producing roughly 30,000 examples per language. This provides exact supervision without reintroducing real PII.

Limitations

This model detects names only and is not a complete anonymization system. It may miss uncommon, ambiguous, or unusually formatted names and may over-redact name-like words. Performance may be lower in other languages or outside medical and medical-adjacent text; high-sensitivity deployments should tune the threshold and validate on local data.

Citation

Authors: Paul Roeseler, Aurélien Ferlay, and Matteo Caliandro

@misc{roeseler2026swisscodingnamefilter,
  title        = {Introducing Swisscoding Name Filter},
  author       = {Roeseler, Paul and Ferlay, Aurélien and Caliandro, Matteo},
  year         = {2026},
  organization = {Swisscoding Technologies},
  url          = {https://swisscoding-labs.github.io/swisscoding-name-filter/}
}
Downloads last month
12
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Swisscoding-Technologies/pii-IT-name-filter-149M

Finetuned
(1532)
this model