obi-lookout

Version v0.6.0, built 2026-10-08. Earlier versions stay available by tag: revision="v0.4.0" or revision="v0.1.0". Since v0.4: business e-mails, chats and letters (first names used alone, honorifics kept out of names, banks as organisations) and fewer false alarms on documentation and encyclopedic text. The medical_or_sensitive_category label, which v0.1 listed but never had training data for, is no longer part of the label list.

Detects personal and business-sensitive spans in text, so they can be masked or blocked before the text reaches a language model or a log. It returns character positions and labels, not generated text. It is a detector, not a chat model.

Read this first

  • It is not a guarantee that every sensitive value is found. Do not use it as the only control.
  • The numbers below come from one evaluation set of 1569 documents, written by a language model from fake data. They show how the model does on that kind of text. Real documents will differ: on real public text it scores about 0.88 to 0.93, not 0.92 on every kind (see "Results on real documents" below), and it has not been tested on your own documents.
  • 36 of 340 documents that contain nothing sensitive had at least one false alarm (10.6%).

The data generators were not audited for generated values that could happen to match real people (for example Malaysian ID numbers, which have no check digit, or phone numbers in real ranges). The data is generated, not taken from real records, but a chance match cannot be ruled out.

Results

Overall strict F1 0.921 (95% interval 0.912 to 0.929). Partial-match F1 (spans overlapping by at least half) 0.925. Strict counts a hit only when start, end and label are all exact.

Label Precision Recall F1 (strict) F1 (partial) Examples
address 0.966 0.993 0.980 0.986 145
client_or_project_name 0.905 0.781 0.839 0.845 832
date_of_birth 0.890 0.990 0.937 0.937 98
email 1.000 0.996 0.998 0.998 560
financial_account 0.982 0.959 0.970 0.970 169
financial_figure 0.992 1.000 0.996 0.996 582
id_number 0.974 0.983 0.978 0.978 115
internal_host 0.974 0.803 0.880 0.891 421
organisation 0.621 0.880 0.728 0.735 474
person 0.980 0.988 0.984 0.986 604
phone 0.961 0.986 0.973 0.976 345
secret 0.986 0.983 0.984 0.990 348

By language:

Language Documents Precision Recall F1 (strict)
en 595 0.919 0.926 0.922
mixed 407 0.911 0.927 0.919
ms 567 0.920 0.922 0.921

Languages not listed were not measured. Results for a language are only as good as the examples in it.

Strict recall is 0.967 on values that also appear in the training data (823 values) and 0.916 on values that do not (3870 values). The second figure is the fairer guide to new text. The weakest label on unseen values is client_or_project_name (recall 0.599).

Results on real documents

We also tested on real public text: Wikipedia (English and Malay), Singapore government company registers, documentation files from open-source projects, and contracts filed with the US SEC. This is public text only; it contains no customer documents. The passages were labelled by a language model (Claude Sonnet 5.5) from a written guide, not by people, so treat small differences as uncertain. The passages are not published because they contain real names. Strict F1 on all twelve labels:

Real test set v0.1.0 v0.6.0
First set, 104 passages 0.55 0.88
Third set, 100 passages from sources not used to design the data (registers, prose, documentation) 0.53 0.90 (95% interval 0.84 to 0.94)
Business documents, 100: 50 SEC contracts and 50 e-mails, chats, invoices, letters and forms written for the test 0.76 0.93 (0.91 to 0.95)

False alarms on passages with nothing sensitive: 16% of 31 (v0.1: 65%) on the first set, 9% of 35 (v0.1: 29%) on the third set, and 10% of 72 passages of encyclopedic and documentation text drawn fresh (v0.1: 15%).

  • Read the business-document number with care. The 50 written documents (0.96) were written after earlier versions' mistakes were known, so they largely test what was fixed. The 50 real SEC contracts score 0.91; their remaining errors are organisation names split over lines and multi-line addresses.
  • Weak spots. Technical text (documentation, configs, command output) is the weakest type (0.49 on 42 first-set passages, and internal_host recall of 0.39 there); person names are sometimes missed in Malay prose (third set 0.81); addresses in registers (0.85) and client or project names versus ordinary companies are still confused.
  • Client or project names, bank accounts and secrets have few or no real examples in these sets (client names are only tested on the written documents), so those labels were not tested on real text.
  • Not measured: your own documents, scanned documents, or any language other than English and Malay. No other open model was re-run against v0.6; the earlier comparison (v0.1, nine labels, first set: 0.46 against 0.31 for the next best) is on the v0.1 page.
  • Labels are machine-made, and fax numbers were labelled as phone numbers.

How to use

pip install torch gliner2==2.0.0 transformers==4.57.6 peft==0.21.2

gliner2 does not install torch or peft, but fails to import without them. Tested with the versions above and torch 2.14. Other versions may work but were not tested.

Quick start: a ready-to-run guide with examples, long-text handling, an OpenAI-compatible Docker server and a Colab notebook is at https://github.com/obiguard/obi-lookout-quickstart

from gliner2 import GLiNER2

model = GLiNER2.from_pretrained("obiguard/obi-lookout")
labels = [
    "person",
    "id_number",
    "phone",
    "email",
    "address",
    "date_of_birth",
    "financial_account",
    "organisation",
    "secret",
    "internal_host",
    "financial_figure",
    "client_or_project_name"
    ]
result = model.extract_entities(text, labels, threshold=0.5, include_confidence=True, include_spans=True)
  • Ask for all the labels at once, as above. Every score on this card was produced that way, with a confidence threshold of 0.5.
  • The model runs offline once downloaded. It was loaded and run with no network and an empty cache to check this.
  • Raw output can contain overlapping spans with different labels for the same text (a hostname can come back as both internal_host and financial_account). Keep the one with the higher confidence. The scores here were computed after doing that.
  • Model, tokenizer and the span-width setting (24 tokens) are all in this repository.

How it was built

Fine-tuned from fastino/gliner2-privacy-filter-PII-multi. Training data: 22,317 documents: 6,879 synthetic (hand-written templates plus documents written by gpt-oss-20b, filled with generated values) and 15,438 generated for v0.2 to v0.6 (tables, directory lists, technical text, business e-mails and chats, hard negatives, public-context prose, multi-script names); no real personal data; 5 epochs, embedding table frozen, span width 24, 12 labels

Training code revision e2d09e0, in Obiguard's private repository. The training data and generators are not published; the evaluation data is.

No real personal data was used. Every identifier in the training and evaluation data was generated, and the generators avoid real mail providers and use reserved or test ranges where such ranges exist.

Evaluation data

1569 documents, SHA-256 2fc1fdfbc4a12989528ed4a77158bb6a9cd34bb3a81648dcdb439009ccc5d51c, published separately so the numbers can be reproduced.

Limitations

  • Trained and measured on synthetic text from the same generators, so the headline results describe similar text. On real public text the score is lower (see above); it has not been tested on your own contracts, logs or chats.
  • Spans longer than 24 model tokens cannot be predicted (the longest span in the training data was 22), so very long values such as a full private key are not covered.
  • medical_or_sensitive_category (listed by v0.1) is not detected and not part of this version's label list; asking for it returns nothing useful.
  • Labels the model confuses most: client_or_project_name found as organisation 171 times; organisation found as client_or_project_name 51 times. Where two labels look alike on the page (a client and an ordinary company are both company names), only the surrounding text separates them.
  • On logs and terminal output the model sometimes flags a public IP address as an internal host, a hex trace ID as an ID number, or user@host as an email, and misses some private IPs inside log lines. A table of bare numbers can swap ID numbers and bank accounts.
  • Internal URLs are labelled as a whole (https://host/path is one internal_host span); edges can differ by a few characters when a path follows.
  • Malaysian and Singapore formats are covered. Other regions are not.
  • Names in unusual positions, abbreviations and mixed-language text are the usual sources of errors.

Attributions

  • GLiNER2
  • mdeberta-v3-base
  • Copyright (c) Microsoft Corporation
  • gpt-oss
Downloads last month
61
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for obiguard/obi-lookout

Finetuned
(5)
this model