Instructions to use tknika/gliner2-PII-basque with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use tknika/gliner2-PII-basque with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
gliner2-PII-basque
A Basque-adapted version of fastino/gliner2-privacy-filter-PII-multi, a 205M-parameter multilingual PII (personally identifiable information) detection model built on GLiNER2 with an mDeBERTa-v3-base encoder.
The base model recognizes 42 PII entity types, but underperforms on Basque text: person names, declension suffixes, and agglutinative forms are frequently missed. This model fine-tunes the base on Basque named-entity data while preserving the original multilingual PII detection capabilities.
Training
Two data sources were combined (experience replay):
- Basque NER: the
nerc_idsplit of orai-nlp/basqueGLUE (EIEC + naiz.eus, manually annotated PER/LOC/ORG/MISC; 2,842 sentences), mapped to the base model's existing types:person name,location,organization,miscellaneous. - Synthetic PII: the train split of tknika/pii-synthetic-basque — 6,000 synthetic sentences in Basque, Spanish and English covering 14 common PII types, with checksum-valid values (Spanish national IDs mod-23, IBANs mod-97, Luhn-valid cards, phone formats). This follows the synthetic-data methodology of the GLiNER2-PII paper, since the original training corpus is not public.
Training used LoRA (r=16, alpha=32) on the boundary/task heads, merged into the base model after training; the resulting adapter is only ~13 MB. The model card repository contains the merged standalone model.
Results
Basque NER
Evaluation on BasqueGLUE validation sets (500 sentences per split, exact match on type + mention string):
| Metric | Base (zero-shot) | This model |
|---|---|---|
| nerc_id/val person name F1 | 0.664 | 0.850 |
| nerc_id/val micro F1 | 0.464 | 0.787 |
| nerc_od/val (Wikipedia) person name F1 | 0.721 | 0.813 |
PII detection
On a PII probe sentence combining several entity types:
| PII type | Base (zero-shot) | This model |
|---|---|---|
| person name | ✓ | ✓ |
| ✓ | ✓ | |
| phone number | ✓ | ✓ |
| bank account number | ✓ | ✓ |
| date of birth | ✓ | ✓ |
| zip code | ✓ | ✓ |
| national id | ✗ | ✓ |
| address | ✗ | ✓ |
PII detection on the v2 synthetic benchmark
This model was later re-evaluated on the eval split of pii-synthetic-basque-v2 (1,000 sentences, 17 types — a benchmark this model was not trained on, with three educational types added after its release). Exact (type, mention) matching; gold mentions use the exact surface form (declension suffixes included):
| Model | Precision | Recall | Micro F1 |
|---|---|---|---|
| Base (zero-shot) | 0.680 | 0.794 | 0.733 |
| This model (v1) | 0.774 | 0.882 | 0.824 |
| v2 (gliner2-PII-basque-v2, trained on this dataset) | 0.960 | 0.990 | 0.975 |
Per-type F1 of this model on the benchmark (types sorted by frequency):
| Type | F1 | Type | F1 | |
|---|---|---|---|---|
| person name | 0.899 | national id | 0.897 | |
| 0.774 | location | 0.695 | ||
| phone number | 0.990 | personal url | 0.670 | |
| user name | 0.615 (base 0.846) | city | 0.511 | |
| date of birth | 0.916 | zip code | 0.720 | |
| student id | 0.937 | bank account number | 0.844 | |
| address | 0.866 | age / date / credit card | 1.000 | |
| organization | 0.744 |
Two observations worth noting:
- The three educational types (
user name,personal url,student id) were introduced with the v2 dataset, after this model was trained.student idtransfers well (0.937) andpersonal urlimproves over the base (0.670 vs. 0.149), butuser namedegrades below the base model (0.615 vs. 0.846): fine-tuning on replay data that does not contain a type partially erases it — the motivation for covering all target types in the v2 replay data. city(0.511) is measured against gold mentions in their exact surface form (e.g. «Amorebietan»), while this model tends to output the undeclined form (e.g. «Amorebieta»), so part of that gap is an artifact of the stricter benchmark rather than pure model weakness.
Fine-tuning a PII model on new-language NER data alone typically causes catastrophic forgetting of the original entity types. The LoRA + experience replay combination used here preserves the base model's PII detection behavior (and even improves on national id and address) while achieving strong Basque NER results.
Usage
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque")
text = ("Kaixo, Ane Azkune naiz, Bilboko La Casilla kalean bizi naiz. "
"Posta: ane.azkune@euskaltel.eus, telefonoa 688 123 456.")
result = model.extract_entities(
text,
["person name", "location", "address", "email", "phone number",
"national id", "bank account number", "credit card number",
"date of birth", "city", "zip code", "age", "date"],
)
Limitations
- Research-quality evaluation (500 sentences per split); not production-tested.
- The synthetic replay data covers 14 of the 42 PII types of the base model; types not covered may still degrade after fine-tuning.
- The license of the BasqueGLUE/EIEC training corpus should be verified before redistribution of derivatives beyond this model.
- The training corpus text is reconstructed by joining tokens with whitespace, so punctuation spacing is slightly non-standard.
Citation
If you use this model, please cite:
@misc{tknika2026gliner2piibasque,
title = {gliner2-PII-basque: A Basque-Adapted Multilingual PII Detection Model},
author = {{TKNIKA} and Ezpeleta Mendikute, Xabier},
year = {2026},
url = {https://huggingface.co/tknika/gliner2-PII-basque}
}
This model builds on the following work, which you may also want to cite:
@misc{fastino2026gliner2pii,
title = {GLiNER2-PII: Multilingual PII Extraction via Synthetic Fine-Tuning},
author = {{Fastino AI Team}},
year = {2026},
url = {https://huggingface.co/fastino/gliner2-pii-v1}
}
@misc{zaratiana2026gliner2piimultilingualmodelpersonally,
title = {GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction},
author = {Zaratiana, Urchade and Lewis, Ash and Hurn-Maloney, George},
year = {2026},
eprint = {2605.09973},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.09973}
}
@InProceedings{urbizu2022basqueglue,
author = {Urbizu, Gorka and San Vicente, Iñaki and Saralegi, Xabier and Agerri, Rodrigo and Soroa, Aitor},
title = {BasqueGLUE: A Natural Language Understanding Benchmark for Basque},
booktitle = {Proceedings of the Language Resources and Evaluation Conference},
year = {2022},
address = {Marseille, France},
publisher = {European Language Resources Association},
pages = {1603--1612},
url = {https://aclanthology.org/2022.lrec-1.172}
}
@misc{he2021debertav3,
title = {DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing},
author = {He, Pengcheng and Gao, Jianfeng and Chen, Weizhu},
year = {2021},
eprint = {2111.09543},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2111.09543}
}
- Downloads last month
- -
Model tree for tknika/gliner2-PII-basque
Base model
fastino/gliner2-privacy-filter-PII-multiDatasets used to train tknika/gliner2-PII-basque
tknika/pii-synthetic-basque
Papers for tknika/gliner2-PII-basque
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
Evaluation results
- Micro F1 on pii-synthetic-basque-v2 (eval split)self-reported0.824
- Micro Precision on pii-synthetic-basque-v2 (eval split)self-reported0.774
- Micro Recall on pii-synthetic-basque-v2 (eval split)self-reported0.882
- Micro F1 on basqueGLUE (nerc_id/validation)validation set self-reported0.787
- Person name F1 on basqueGLUE (nerc_id/validation)validation set self-reported0.850