c5k-deberta_base-token_level-1-2

Superseded. This is an early checkpoint, kept because it has a DOI. For current pt-BR PII detection use arthrod/gliner-mmbert-small-ptbr-pii-full-3x-v1 or the OpenAI Privacy Filter fine-tune arthrod/gliner-opf-ptbr-pii-v1; results for both are in arthrod/gliner-opf-ptbr-pii-bench-v1.

Token-level GLiNER checkpoint fine-tuned for Brazilian Portuguese PII detection. Base model: microsoft/deberta-v3-base (~NoneM params).

Intended use

Brazilian Portuguese PII detection in legal, medical, and administrative text. Supports any GLiNER-compatible label set (CPF, RG, email, phone, person name, address, etc.).

Limitations

  • Trained primarily on Portuguese text; English/Spanish performance is not guaranteed.
  • Span boundaries depend on tokenization at inference time; take spans from the model's character offsets and never reconstruct raw text from tokens.
  • No benchmark results are published for this checkpoint.

Citation

@misc{arthrod_gliner_ptbr_pii,
  author = {arthrod},
  title  = {GLiNER Portuguese PII checkpoints},
  year   = {2026},
  url    = {https://huggingface.co/arthrod/c5k-deberta_base-token_level-1-2}
}
Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arthrod/c5k-deberta_base-token_level-1-2

Quantized
(30)
this model