gliner2-small

Schema-guided named entity recognition using GLiNER2. Supply entity names and concise definitions at inference time; the output contains extracted spans and optional confidence scores.

Model and training provenance

Property Value
Model Siddharth63/gliner2-small
Evaluated checkpoint revision 756dba7c16ddc81baa99fed3b36f548bbd6cd254
Encoder jhu-clsp/ettin-encoder-68m
Stored parameter count 81,304,085
Extraction head count_lstm; token pooling first
Intended language English; multilingual quality not established here
Domain General and multiple domains

The author confirms this general model was trained on the audited combined corpus.

The combined source covers general text, patents, patent claims, materials, finance, environment, biomedicine and law, according to the author’s source-file inventory. The local source was final_gliner_file.csv.gz; this repository does not distribute training passages. The audit found 1,232,284 rows, 1,232,280 parseable NER records, 1,179,356 unique normalized texts, 630 entity labels, and 698 exact definition variants. Four records failed parsing; nine parseable records contained empty text. These are corpus-audit statistics, not inferred training hyperparameters.

Definitions contain 2–9 whitespace words, with median 4. The complete global label inventory, exact definition variants and observed counts are in training-labels.json and training-labels.tsv. Training steps, optimizer settings, checkpoint selection criteria and any model-specific data filtering were not independently recovered.

Quick start

Install the dependencies listed in the recorded environment, or install gliner2, transformers, torch, and peft in a compatible environment. The run uses GLiNER2 2.0.0 and Transformers 5.16.1. For exact reproduction use the lock file.

import torch
from gliner2 import GLiNER2

model = GLiNER2.from_pretrained("Siddharth63/gliner2-small").to("cuda").eval()
schema = {"person": "Individual human being", "company": "Business organization or corporation"}
with torch.inference_mode():
    result = model.batch_extract_entities(
        ['Alice joined Acme in London.'],
        schema,
        threshold=0.5,
        include_spans=True,
        include_confidence=True,
        max_len=3072,
    )[0]
print(result)

Definitions in this example are copied from the audited training corpus. The example is usage code; no fabricated expected prediction is supplied. Confidence scores are model scores, not calibrated probabilities. Keep entity definitions concise and distinguish neighboring types explicitly.

Context and entity length

Setting Value
Saved extractor max_len 8192
Saved extractor max_length 3072
Encoder max_position_embeddings 7999
Explicit max_len used in this evaluation 3072
Evaluation source window 64 whitespace tokens, overlap 16
Saved maximum candidate span width 8

The checkpoint contains conflicting length settings: max_len=8192, max_length=3072, and an encoder position setting of 7999. Do not interpret the 8192 field as a verified end-to-end context limit. This evaluation explicitly uses 3072 with short source windows. Schema tokens, special tokens and source subwords consume the processing budget; the advertised budget is not a guarantee of that many source words. GLiNER2 library defaults can differ from saved checkpoint values, so pass max_len explicitly.

Long documents should be split into overlapping windows and merged using original character offsets. The maximum candidate span width is distinct from the context budget and can limit recall on long entity mentions. No claim of validated extraction at the encoder’s maximum position setting is made.

Evaluation: two conditions

This is a time-bounded reproducible screening run, using up to 64 deterministically chosen test records per dataset, selected before inference with seed gliner-20260928-v1. These are sampled-panel results, not full official benchmark scores. The headline transfer suite is CrossNER’s five domains plus MIT Movie and MIT Restaurant.

A. Benchmark schema transfer: use the benchmark’s entity names with concise 2–9-word definitions. The attached manifest records each definition and its semantic relation to training labels. A different name alone does not establish unseen-type transfer.

B. Familiar labels on held-out text: replace only compatible benchmark types with the exact training label and an original training definition. Types with incompatible scope are excluded from this condition’s ontology; all selected texts, including negative examples for the retained ontology, remain in evaluation. A and B therefore have different target ontologies and their F1 values should not be compared as a pure prompt ablation.

The new-specific-type subset includes narrower types absent as equivalent explicit training labels, but related concepts such as person or occupation can still be familiar. It is not proof of completely unseen semantic concepts. For specialists, familiar-label assignments are verified against the combined corpus only. There is no benchmark-specific fitting, gradient update, few-shot demonstration or test-set threshold search.

A: benchmark-schema span F1 (%)

Benchmark This model GLiNER-bi-large-v2.0 GNER-T5-xxl NuExtract-2.0-8B
crossner_ai 13.77 47.71 51.13 36.70
crossner_literature 17.86 57.92 58.49 48.70
crossner_music 17.54 73.18 77.00 75.94
crossner_politics 26.02 50.64 68.06 46.40
crossner_science 17.12 63.84 64.64 52.47
mit_movie 25.60 41.38 53.54 51.46
mit_restaurant 3.11 42.24 41.59 30.07
bc5cdr 51.50 57.65 59.65 61.22
finer_ord 10.64 14.29 55.90 64.12
inlegalner 3.41 23.74 18.34 20.37
biodivner 8.76 26.25 27.77 17.11

B: exact training-schema span F1 (%)

Benchmark This model GLiNER-bi-large-v2.0 GNER-T5-xxl NuExtract-2.0-8B
crossner_ai 28.23 35.29 50.78 43.58
crossner_literature 39.38 56.53 64.33 61.14
crossner_music 38.92 75.00 71.70 75.09
crossner_politics 35.69 63.92 67.28 50.43
crossner_science 37.02 50.31 54.09 53.40
mit_movie 41.67 47.76 67.65 44.44
mit_restaurant 16.00 52.63 42.50 9.52
bc5cdr 55.43 58.96 63.53 57.14
finer_ord 42.25 18.82 57.38 44.71
inlegalner 24.31 27.59 21.56 30.26
biodivner 12.90 50.00 17.91 0.00

New specific types with related training concepts

Benchmark Gold mentions F1 on absent specific types
crossner_ai 120 8.89
crossner_literature 109 0.00
crossner_music 127 4.58
crossner_politics 51 0.00
crossner_science 118 4.79
mit_movie 130 24.71
mit_restaurant 42 2.11
inlegalner 4 2.30
biodivner 87 13.33

Precision, recall and uncertainty for this model

Benchmark / track Clean n Precision Recall F1 95% interval
crossner_ai / benchmark_schema 64 25.84 9.39 6.29–23.23
crossner_ai / familiar_training_schema 64 22.58 37.63 16.98–40.60
crossner_literature / benchmark_schema 64 34.45 12.06 12.56–23.44
crossner_literature / familiar_training_schema 64 37.07 41.99 31.18–46.97
crossner_music / benchmark_schema 63 44.76 10.90 10.43–25.22
crossner_music / familiar_training_schema 63 38.84 39.00 29.59–48.10
crossner_politics / benchmark_schema 64 32.36 21.76 20.25–32.31
crossner_politics / familiar_training_schema 64 29.95 44.16 29.94–41.36
crossner_science / benchmark_schema 64 24.74 13.09 9.89–25.05
crossner_science / familiar_training_schema 64 31.03 45.86 28.23–46.82
mit_movie / benchmark_schema 64 61.54 16.16 18.46–32.64
mit_movie / familiar_training_schema 64 34.09 53.57 26.31–56.34
mit_restaurant / benchmark_schema 64 4.62 2.34 0.00–6.56
mit_restaurant / familiar_training_schema 64 17.14 15.00 5.55–28.17
bc5cdr / benchmark_schema 61 52.44 50.59 39.52–62.28
bc5cdr / familiar_training_schema 61 51.52 60.00 43.31–66.35
finer_ord / benchmark_schema 64 21.74 7.04 2.20–20.41
finer_ord / familiar_training_schema 64 36.14 50.85 33.33–51.13
inlegalner / benchmark_schema 64 5.21 2.54 0.78–7.17
inlegalner / familiar_training_schema 64 20.35 30.17 15.52–36.46
biodivner / benchmark_schema 64 23.08 5.41 4.81–12.81
biodivner / familiar_training_schema 64 10.00 18.18 0.00–28.59

22 scored cells are available out of at most 22; some familiar-schema cells may be inapplicable. A dash means no completed applicable result, never zero performance.

Tables use the same fixed dataset samples and report results after removing detected training overlap. Evaluation is exact character span and entity type. Overlapping window predictions are deduplicated. Source windows contain 64 whitespace tokens with overlap 16; all gold entities remain in scoring, including entities longer than a window or model span limit. BIO datasets are rendered with a single space between source tokens. The strict offset protocol, windowing and concise-description prompts differ from some published benchmarks; published paper scores must not be mixed into these tables.

GLiNER2 receives a name-to-definition schema. GLiNER-bi receives label: definition strings, mapped back to benchmark types. GNER uses its published BIO instruction with a concise definition section. NuExtract uses its native extraction template, verbatim-string arrays and an instruction containing the same definitions. Generative models run greedily in bfloat16 with 512 generated tokens per window. Native outputs are aligned to exact text spans; unaligned entity strings count as false positives. No hidden synonym or substring credit is awarded. The adapters are an interface-aligned comparison, not identical native prompting. Generation failures and length-limit events are recorded in result JSON.

Document-level bootstrap intervals use 1,000 resamples with seed 20260928. Documents are sentence records for sentence-level corpora, not independent articles; intervals can understate uncertainty where sentences share source articles. A 64-record screening panel has limited statistical power, particularly for rare types. Do not claim state-of-the-art or a decisive ranking from overlapping intervals.

Comparator choice

The run selects three strong, locally executable open-checkpoint references: GLiNER-bi-large-v2.0 (recent NER-specific encoder), GNER-T5-xxl (generative zero-shot NER), and NuExtract-2.0-8B (schema-conditioned extraction). NuExtract is an extraction reference, not a claim that it tops every NER benchmark. This is a practical comparison set, not a verified universal top-three ranking across all NER tasks. Different baseline training corpora may include related benchmark material; those full corpora were not audited.

Training overlap and limitations

The full supplied training file was screened against the fixed evaluation panel. Audit completed: True; rows scanned: 1,232,284; panel records flagged: 4. Exact checks use NFKC/casefold word-normalized SHA-256. Near-duplicate screening uses five eight-word anchors for passages with at least twelve words, followed by substring or >=95 partial-ratio verification. Flagged records are excluded in the clean tables, while unfiltered metrics remain in JSON.

Anchor screening can miss short duplicates, paraphrases and other leakage. This establishes only that no match was detected by the recorded method for retained records; it does not prove absence from pretraining, synthetic-data prompts, source document families or comparator training. The broad domains themselves were already represented in the author’s training sources. Thus these results measure schema/data transfer, not necessarily unseen-domain generalization.

Unavailable datasets in this run:

All requested initial-panel datasets were available.

Materials and patent-specific gold benchmarks have not been scored by this initial panel. Domain expertise, long-context quality, nested/discontinuous entities, negation understanding and real-world safety are not established by these scores. Separately validate extraction before consequential legal, biomedical or financial use.

Reproducibility and artifacts

The scripts run on one NVIDIA A100 40 GB. To reproduce, place the authorized training CSV at the project root, create audit/labels.json with audit_training.py, and retain the pinned model inventory and dataset manifest. Run prepare_benchmarks.py, audit_holdout.py, and evaluate_model.py --model Siddharth63/gliner2-small --deadline <UTC Unix timestamp>. For an exact rerun, use the recorded manifest and panel IDs rather than refreshing revisions. Do not upload raw training data to a public repository.

License and publication status

The repository’s original license metadata is preserved: artistic-2.0. Dataset and backbone licenses have their own terms; this README does not grant additional rights. Repository visibility and model weights were not changed. No leaderboard submission is implied by these locally generated results.

Sources

Card generated from audited artifacts on 2026-09-28T20:27:32+00:00.

Follow-up investigation

The investigation report separates runtime checks, training-data findings and controlled development experiments. Development sweeps use 32 fixed records from each of seven official development splits; these records were not screened against the full training corpus. Threshold and schema changes are diagnostic and were not substituted into the original fixed-protocol test scores. No checkpoint weights were modified.

The original GLiNER2 extraction interface can return multiple types for the same span, while the GLiNER-bi adapter requests flat output. This decoding difference limits direct interpretation of the original native-output comparison. The secondary table below removes overlaps across labels and windows for flat NER. GLiNER2 candidates are ranked by native confidence; saved comparator outputs lack confidences, so their ties use longer spans, then offsets and label names. Generative BIO output is usually already flat. This is a post-hoc protocol check, not a leaderboard submission.

Benchmark: flat-output schema transfer This model GLiNER-bi-large-v2.0 GNER-T5-xxl NuExtract-2.0-8B
crossner_ai 13.86 47.71 51.13 33.88
crossner_literature 17.62 58.08 58.45 47.89
crossner_music 17.60 73.50 77.49 76.67
crossner_politics 24.88 50.72 68.32 44.14
crossner_science 17.12 63.94 64.73 49.45
mit_movie 25.60 41.38 53.54 50.00
mit_restaurant 2.12 42.24 41.59 30.07
bc5cdr 50.00 57.65 59.65 61.64
finer_ord 10.99 14.29 55.90 64.12
inlegalner 2.07 23.77 19.35 19.92
biodivner 8.12 26.25 27.30 16.00

The familiar-label science schema was corrected to use the exact training definitions “Chemical element from periodic table” and “Scientific theory or hypothesis”. Earlier selection by definition frequency chose unrelated senses. Corrected completed cells replace those two-definition results; the benchmark-schema condition was unaffected. Residual categories such as person/organization can still differ in scope between annotation schemes, so familiar-label results remain schema-transfer evidence rather than a proof of identical annotation policy.

Expanded evaluation and decision-layer experiments

A separate, larger holdout excludes normalized text duplicates from the earlier sampled panels and development set. The experiment report records both familiar-label transfer and new-specific-type transfer, current GLiNER2.5 comparisons, specialist in-domain results, and completion status. Missing cells remain pending.

Fresh seven-dataset core suite F1 (%)
Familiar labels: mean dataset F1 pending
Benchmark schema: mean dataset F1 pending
New specific types: pooled F1 pending

The repaired pipeline adds Mapika/decider-4b v2.1 to this checkpoint. Pipeline gains, correction damage, paired intervals and measured request latency are separate from this model’s own scores; weights in this repository were not updated. The primary held-out repair condition was fixed before fresh inference.

GLiNER2.5 on the original frozen core panel

These use the same earlier seven core-dataset samples, concise schemas, strict offsets and threshold 0.5. Boundary models use their native flat decoding; legacy GLiNER2 can return overlapping labels.

Model Familiar-label mean F1 Benchmark-schema mean F1
Siddharth63/gliner2-small 33.84 17.29
fastino/gliner2.5-small-v1 40.81 41.28
fastino/gliner2.5-base-v1 47.71 51.53
fastino/gliner2.5-multi-v1 47.04 50.67

JPT access and results, if available, are recorded separately in the report. Published paper scores are not substituted for measurements under this protocol. This expanded run makes no state-of-the-art or leaderboard-entry claim.

Downloads last month
122
Safetensors
Model size
81.3M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support