gliner2-PII-basque

A Basque-adapted version of fastino/gliner2-privacy-filter-PII-multi, a 205M-parameter multilingual PII (personally identifiable information) detection model built on GLiNER2 with an mDeBERTa-v3-base encoder.

The base model recognizes 42 PII entity types, but underperforms on Basque text: person names, declension suffixes, and agglutinative forms are frequently missed. This model fine-tunes the base on Basque named-entity data while preserving the original multilingual PII detection capabilities.

Training

Two data sources were combined (experience replay):

  • Basque NER: the nerc_id split of orai-nlp/basqueGLUE (EIEC + naiz.eus, manually annotated PER/LOC/ORG/MISC; 2,842 sentences), mapped to the base model's existing types: person name, location, organization, miscellaneous.
  • Synthetic PII: the train split of tknika/pii-synthetic-basque — 6,000 synthetic sentences in Basque, Spanish and English covering 14 common PII types, with checksum-valid values (Spanish national IDs mod-23, IBANs mod-97, Luhn-valid cards, phone formats). This follows the synthetic-data methodology of the GLiNER2-PII paper, since the original training corpus is not public.

Training used LoRA (r=16, alpha=32) on the boundary/task heads, merged into the base model after training; the resulting adapter is only ~13 MB. The model card repository contains the merged standalone model.

Results

Basque NER

Evaluation on BasqueGLUE validation sets (500 sentences per split, exact match on type + mention string):

Metric Base (zero-shot) This model
nerc_id/val person name F1 0.664 0.850
nerc_id/val micro F1 0.464 0.787
nerc_od/val (Wikipedia) person name F1 0.721 0.813

PII detection

On a PII probe sentence combining several entity types:

PII type Base (zero-shot) This model
person name ✓ ✓
email ✓ ✓
phone number ✓ ✓
bank account number ✓ ✓
date of birth ✓ ✓
zip code ✓ ✓
national id ✗ ✓
address ✗ ✓

PII detection on the v2 synthetic benchmark

This model was later re-evaluated on the eval split of pii-synthetic-basque-v2 (1,000 sentences, 17 types — a benchmark this model was not trained on, with three educational types added after its release). Exact (type, mention) matching; gold mentions use the exact surface form (declension suffixes included):

Model Precision Recall Micro F1
Base (zero-shot) 0.680 0.794 0.733
This model (v1) 0.774 0.882 0.824
v2 (gliner2-PII-basque-v2, trained on this dataset) 0.960 0.990 0.975

Per-type F1 of this model on the benchmark (types sorted by frequency):

Type F1 Type F1
person name 0.899 national id 0.897
email 0.774 location 0.695
phone number 0.990 personal url 0.670
user name 0.615 (base 0.846) city 0.511
date of birth 0.916 zip code 0.720
student id 0.937 bank account number 0.844
address 0.866 age / date / credit card 1.000
organization 0.744

Two observations worth noting:

  • The three educational types (user name, personal url, student id) were introduced with the v2 dataset, after this model was trained. student id transfers well (0.937) and personal url improves over the base (0.670 vs. 0.149), but user name degrades below the base model (0.615 vs. 0.846): fine-tuning on replay data that does not contain a type partially erases it — the motivation for covering all target types in the v2 replay data.
  • city (0.511) is measured against gold mentions in their exact surface form (e.g. «Amorebietan»), while this model tends to output the undeclined form (e.g. «Amorebieta»), so part of that gap is an artifact of the stricter benchmark rather than pure model weakness.

Fine-tuning a PII model on new-language NER data alone typically causes catastrophic forgetting of the original entity types. The LoRA + experience replay combination used here preserves the base model's PII detection behavior (and even improves on national id and address) while achieving strong Basque NER results.

Usage

from gliner2 import AutoExtractor

model = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque")

text = ("Kaixo, Ane Azkune naiz, Bilboko La Casilla kalean bizi naiz. "
        "Posta: ane.azkune@euskaltel.eus, telefonoa 688 123 456.")
result = model.extract_entities(
    text,
    ["person name", "location", "address", "email", "phone number",
     "national id", "bank account number", "credit card number",
     "date of birth", "city", "zip code", "age", "date"],
)

Limitations

  • Research-quality evaluation (500 sentences per split); not production-tested.
  • The synthetic replay data covers 14 of the 42 PII types of the base model; types not covered may still degrade after fine-tuning.
  • The license of the BasqueGLUE/EIEC training corpus should be verified before redistribution of derivatives beyond this model.
  • The training corpus text is reconstructed by joining tokens with whitespace, so punctuation spacing is slightly non-standard.

Citation

If you use this model, please cite:

@misc{tknika2026gliner2piibasque,
  title  = {gliner2-PII-basque: A Basque-Adapted Multilingual PII Detection Model},
  author = {{TKNIKA} and Ezpeleta Mendikute, Xabier},
  year   = {2026},
  url    = {https://huggingface.co/tknika/gliner2-PII-basque}
}

This model builds on the following work, which you may also want to cite:

@misc{fastino2026gliner2pii,
  title   = {GLiNER2-PII: Multilingual PII Extraction via Synthetic Fine-Tuning},
  author  = {{Fastino AI Team}},
  year    = {2026},
  url     = {https://huggingface.co/fastino/gliner2-pii-v1}
}

@misc{zaratiana2026gliner2piimultilingualmodelpersonally,
  title  = {GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction},
  author = {Zaratiana, Urchade and Lewis, Ash and Hurn-Maloney, George},
  year   = {2026},
  eprint = {2605.09973},
  archivePrefix = {arXiv},
  url   = {https://arxiv.org/abs/2605.09973}
}

@InProceedings{urbizu2022basqueglue,
  author    = {Urbizu, Gorka and San Vicente, Iñaki and Saralegi, Xabier and Agerri, Rodrigo and Soroa, Aitor},
  title     = {BasqueGLUE: A Natural Language Understanding Benchmark for Basque},
  booktitle = {Proceedings of the Language Resources and Evaluation Conference},
  year      = {2022},
  address   = {Marseille, France},
  publisher = {European Language Resources Association},
  pages     = {1603--1612},
  url       = {https://aclanthology.org/2022.lrec-1.172}
}

@misc{he2021debertav3,
  title  = {DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing},
  author = {He, Pengcheng and Gao, Jianfeng and Chen, Weizhu},
  year   = {2021},
  eprint = {2111.09543},
  archivePrefix = {arXiv},
  url   = {https://arxiv.org/abs/2111.09543}
}
Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tknika/gliner2-PII-basque

Finetuned
(5)
this model

Datasets used to train tknika/gliner2-PII-basque

Papers for tknika/gliner2-PII-basque

Evaluation results