Instructions to use Siddharth63/gliner2-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use Siddharth63/gliner2-small with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("Siddharth63/gliner2-small") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
- gliner2-small
gliner2-small
Schema-guided named entity recognition using GLiNER2. Supply entity names and concise definitions at inference time; the output contains extracted spans and optional confidence scores.
Model and training provenance
| Property | Value |
|---|---|
| Model | Siddharth63/gliner2-small |
| Evaluated checkpoint revision | 756dba7c16ddc81baa99fed3b36f548bbd6cd254 |
| Encoder | jhu-clsp/ettin-encoder-68m |
| Stored parameter count | 81,304,085 |
| Extraction head | count_lstm; token pooling first |
| Intended language | English; multilingual quality not established here |
| Domain | General and multiple domains |
The author confirms this general model was trained on the audited combined corpus.
The combined source covers general text, patents, patent claims, materials, finance, environment, biomedicine and law, according to the author’s source-file inventory. The local source was final_gliner_file.csv.gz; this repository does not distribute training passages. The audit found 1,232,284 rows, 1,232,280 parseable NER records, 1,179,356 unique normalized texts, 630 entity labels, and 698 exact definition variants. Four records failed parsing; nine parseable records contained empty text. These are corpus-audit statistics, not inferred training hyperparameters.
Definitions contain 2–9 whitespace words, with median 4. The complete global label inventory, exact definition variants and observed counts are in training-labels.json and training-labels.tsv. Training steps, optimizer settings, checkpoint selection criteria and any model-specific data filtering were not independently recovered.
Quick start
Install the dependencies listed in the recorded environment, or install gliner2, transformers, torch, and peft in a compatible environment. The run uses GLiNER2 2.0.0 and Transformers 5.16.1. For exact reproduction use the lock file.
import torch
from gliner2 import GLiNER2
model = GLiNER2.from_pretrained("Siddharth63/gliner2-small").to("cuda").eval()
schema = {"person": "Individual human being", "company": "Business organization or corporation"}
with torch.inference_mode():
result = model.batch_extract_entities(
['Alice joined Acme in London.'],
schema,
threshold=0.5,
include_spans=True,
include_confidence=True,
max_len=3072,
)[0]
print(result)
Definitions in this example are copied from the audited training corpus. The example is usage code; no fabricated expected prediction is supplied. Confidence scores are model scores, not calibrated probabilities. Keep entity definitions concise and distinguish neighboring types explicitly.
Context and entity length
| Setting | Value |
|---|---|
Saved extractor max_len |
8192 |
Saved extractor max_length |
3072 |
Encoder max_position_embeddings |
7999 |
Explicit max_len used in this evaluation |
3072 |
| Evaluation source window | 64 whitespace tokens, overlap 16 |
| Saved maximum candidate span width | 8 |
The checkpoint contains conflicting length settings: max_len=8192, max_length=3072, and an encoder position setting of 7999. Do not interpret the 8192 field as a verified end-to-end context limit. This evaluation explicitly uses 3072 with short source windows. Schema tokens, special tokens and source subwords consume the processing budget; the advertised budget is not a guarantee of that many source words. GLiNER2 library defaults can differ from saved checkpoint values, so pass max_len explicitly.
Long documents should be split into overlapping windows and merged using original character offsets. The maximum candidate span width is distinct from the context budget and can limit recall on long entity mentions. No claim of validated extraction at the encoder’s maximum position setting is made.
Evaluation: two conditions
This is a time-bounded reproducible screening run, using up to 64 deterministically chosen test records per dataset, selected before inference with seed gliner-20260928-v1. These are sampled-panel results, not full official benchmark scores. The headline transfer suite is CrossNER’s five domains plus MIT Movie and MIT Restaurant.
A. Benchmark schema transfer: use the benchmark’s entity names with concise 2–9-word definitions. The attached manifest records each definition and its semantic relation to training labels. A different name alone does not establish unseen-type transfer.
B. Familiar labels on held-out text: replace only compatible benchmark types with the exact training label and an original training definition. Types with incompatible scope are excluded from this condition’s ontology; all selected texts, including negative examples for the retained ontology, remain in evaluation. A and B therefore have different target ontologies and their F1 values should not be compared as a pure prompt ablation.
The new-specific-type subset includes narrower types absent as equivalent explicit training labels, but related concepts such as person or occupation can still be familiar. It is not proof of completely unseen semantic concepts. For specialists, familiar-label assignments are verified against the combined corpus only. There is no benchmark-specific fitting, gradient update, few-shot demonstration or test-set threshold search.
A: benchmark-schema span F1 (%)
| Benchmark | This model | GLiNER-bi-large-v2.0 | GNER-T5-xxl | NuExtract-2.0-8B |
|---|---|---|---|---|
| crossner_ai | 13.77 | 47.71 | 51.13 | 36.70 |
| crossner_literature | 17.86 | 57.92 | 58.49 | 48.70 |
| crossner_music | 17.54 | 73.18 | 77.00 | 75.94 |
| crossner_politics | 26.02 | 50.64 | 68.06 | 46.40 |
| crossner_science | 17.12 | 63.84 | 64.64 | 52.47 |
| mit_movie | 25.60 | 41.38 | 53.54 | 51.46 |
| mit_restaurant | 3.11 | 42.24 | 41.59 | 30.07 |
| bc5cdr | 51.50 | 57.65 | 59.65 | 61.22 |
| finer_ord | 10.64 | 14.29 | 55.90 | 64.12 |
| inlegalner | 3.41 | 23.74 | 18.34 | 20.37 |
| biodivner | 8.76 | 26.25 | 27.77 | 17.11 |
B: exact training-schema span F1 (%)
| Benchmark | This model | GLiNER-bi-large-v2.0 | GNER-T5-xxl | NuExtract-2.0-8B |
|---|---|---|---|---|
| crossner_ai | 28.23 | 35.29 | 50.78 | 43.58 |
| crossner_literature | 39.38 | 56.53 | 64.33 | 61.14 |
| crossner_music | 38.92 | 75.00 | 71.70 | 75.09 |
| crossner_politics | 35.69 | 63.92 | 67.28 | 50.43 |
| crossner_science | 37.02 | 50.31 | 54.09 | 53.40 |
| mit_movie | 41.67 | 47.76 | 67.65 | 44.44 |
| mit_restaurant | 16.00 | 52.63 | 42.50 | 9.52 |
| bc5cdr | 55.43 | 58.96 | 63.53 | 57.14 |
| finer_ord | 42.25 | 18.82 | 57.38 | 44.71 |
| inlegalner | 24.31 | 27.59 | 21.56 | 30.26 |
| biodivner | 12.90 | 50.00 | 17.91 | 0.00 |
New specific types with related training concepts
| Benchmark | Gold mentions | F1 on absent specific types |
|---|---|---|
| crossner_ai | 120 | 8.89 |
| crossner_literature | 109 | 0.00 |
| crossner_music | 127 | 4.58 |
| crossner_politics | 51 | 0.00 |
| crossner_science | 118 | 4.79 |
| mit_movie | 130 | 24.71 |
| mit_restaurant | 42 | 2.11 |
| inlegalner | 4 | 2.30 |
| biodivner | 87 | 13.33 |
Precision, recall and uncertainty for this model
| Benchmark / track | Clean n | Precision | Recall | F1 95% interval |
|---|---|---|---|---|
| crossner_ai / benchmark_schema | 64 | 25.84 | 9.39 | 6.29–23.23 |
| crossner_ai / familiar_training_schema | 64 | 22.58 | 37.63 | 16.98–40.60 |
| crossner_literature / benchmark_schema | 64 | 34.45 | 12.06 | 12.56–23.44 |
| crossner_literature / familiar_training_schema | 64 | 37.07 | 41.99 | 31.18–46.97 |
| crossner_music / benchmark_schema | 63 | 44.76 | 10.90 | 10.43–25.22 |
| crossner_music / familiar_training_schema | 63 | 38.84 | 39.00 | 29.59–48.10 |
| crossner_politics / benchmark_schema | 64 | 32.36 | 21.76 | 20.25–32.31 |
| crossner_politics / familiar_training_schema | 64 | 29.95 | 44.16 | 29.94–41.36 |
| crossner_science / benchmark_schema | 64 | 24.74 | 13.09 | 9.89–25.05 |
| crossner_science / familiar_training_schema | 64 | 31.03 | 45.86 | 28.23–46.82 |
| mit_movie / benchmark_schema | 64 | 61.54 | 16.16 | 18.46–32.64 |
| mit_movie / familiar_training_schema | 64 | 34.09 | 53.57 | 26.31–56.34 |
| mit_restaurant / benchmark_schema | 64 | 4.62 | 2.34 | 0.00–6.56 |
| mit_restaurant / familiar_training_schema | 64 | 17.14 | 15.00 | 5.55–28.17 |
| bc5cdr / benchmark_schema | 61 | 52.44 | 50.59 | 39.52–62.28 |
| bc5cdr / familiar_training_schema | 61 | 51.52 | 60.00 | 43.31–66.35 |
| finer_ord / benchmark_schema | 64 | 21.74 | 7.04 | 2.20–20.41 |
| finer_ord / familiar_training_schema | 64 | 36.14 | 50.85 | 33.33–51.13 |
| inlegalner / benchmark_schema | 64 | 5.21 | 2.54 | 0.78–7.17 |
| inlegalner / familiar_training_schema | 64 | 20.35 | 30.17 | 15.52–36.46 |
| biodivner / benchmark_schema | 64 | 23.08 | 5.41 | 4.81–12.81 |
| biodivner / familiar_training_schema | 64 | 10.00 | 18.18 | 0.00–28.59 |
22 scored cells are available out of at most 22; some familiar-schema cells may be inapplicable. A dash means no completed applicable result, never zero performance.
Tables use the same fixed dataset samples and report results after removing detected training overlap. Evaluation is exact character span and entity type. Overlapping window predictions are deduplicated. Source windows contain 64 whitespace tokens with overlap 16; all gold entities remain in scoring, including entities longer than a window or model span limit. BIO datasets are rendered with a single space between source tokens. The strict offset protocol, windowing and concise-description prompts differ from some published benchmarks; published paper scores must not be mixed into these tables.
GLiNER2 receives a name-to-definition schema. GLiNER-bi receives label: definition strings, mapped back to benchmark types. GNER uses its published BIO instruction with a concise definition section. NuExtract uses its native extraction template, verbatim-string arrays and an instruction containing the same definitions. Generative models run greedily in bfloat16 with 512 generated tokens per window. Native outputs are aligned to exact text spans; unaligned entity strings count as false positives. No hidden synonym or substring credit is awarded. The adapters are an interface-aligned comparison, not identical native prompting. Generation failures and length-limit events are recorded in result JSON.
Document-level bootstrap intervals use 1,000 resamples with seed 20260928. Documents are sentence records for sentence-level corpora, not independent articles; intervals can understate uncertainty where sentences share source articles. A 64-record screening panel has limited statistical power, particularly for rare types. Do not claim state-of-the-art or a decisive ranking from overlapping intervals.
Comparator choice
The run selects three strong, locally executable open-checkpoint references: GLiNER-bi-large-v2.0 (recent NER-specific encoder), GNER-T5-xxl (generative zero-shot NER), and NuExtract-2.0-8B (schema-conditioned extraction). NuExtract is an extraction reference, not a claim that it tops every NER benchmark. This is a practical comparison set, not a verified universal top-three ranking across all NER tasks. Different baseline training corpora may include related benchmark material; those full corpora were not audited.
Training overlap and limitations
The full supplied training file was screened against the fixed evaluation panel. Audit completed: True; rows scanned: 1,232,284; panel records flagged: 4. Exact checks use NFKC/casefold word-normalized SHA-256. Near-duplicate screening uses five eight-word anchors for passages with at least twelve words, followed by substring or >=95 partial-ratio verification. Flagged records are excluded in the clean tables, while unfiltered metrics remain in JSON.
Anchor screening can miss short duplicates, paraphrases and other leakage. This establishes only that no match was detected by the recorded method for retained records; it does not prove absence from pretraining, synthetic-data prompts, source document families or comparator training. The broad domains themselves were already represented in the author’s training sources. Thus these results measure schema/data transfer, not necessarily unseen-domain generalization.
Unavailable datasets in this run:
All requested initial-panel datasets were available.
Materials and patent-specific gold benchmarks have not been scored by this initial panel. Domain expertise, long-context quality, nested/discontinuous entities, negation understanding and real-world safety are not established by these scores. Separately validate extraction before consequential legal, biomedical or financial use.
Reproducibility and artifacts
- Protocol and dataset revisions, including exact sample IDs and schema definitions.
- This model’s metrics, including precision/recall, counts, intervals and inference time.
- All completed comparison metrics; missing results remain missing.
- Corpus audit summary and panel overlap summary.
- Runtime observations, environment versions, and run status.
- Evaluation source files in evaluation/scripts.
The scripts run on one NVIDIA A100 40 GB. To reproduce, place the authorized training CSV at the project root, create audit/labels.json with audit_training.py, and retain the pinned model inventory and dataset manifest. Run prepare_benchmarks.py, audit_holdout.py, and evaluate_model.py --model Siddharth63/gliner2-small --deadline <UTC Unix timestamp>. For an exact rerun, use the recorded manifest and panel IDs rather than refreshing revisions. Do not upload raw training data to a public repository.
License and publication status
The repository’s original license metadata is preserved: artistic-2.0. Dataset and backbone licenses have their own terms; this README does not grant additional rights. Repository visibility and model weights were not changed. No leaderboard submission is implied by these locally generated results.
Sources
- GLiNER: generalist NER, NAACL 2024.
- CrossNER official data.
- MIT movie and MIT restaurant.
- BC5CDR, FiNER-ORD, InLegalNER, BiodivNER.
- GLiNER-bi-large-v2.0, GNER-T5-xxl, NuExtract-2.0-8B.
- Just Pass Twice, ACL 2026 was considered; its accessible service requires separate credentials/quota and was not silently substituted for a reproducible local checkpoint.
Card generated from audited artifacts on 2026-09-28T20:27:32+00:00.
Follow-up investigation
The investigation report separates runtime checks, training-data findings and controlled development experiments. Development sweeps use 32 fixed records from each of seven official development splits; these records were not screened against the full training corpus. Threshold and schema changes are diagnostic and were not substituted into the original fixed-protocol test scores. No checkpoint weights were modified.
The original GLiNER2 extraction interface can return multiple types for the same span, while the GLiNER-bi adapter requests flat output. This decoding difference limits direct interpretation of the original native-output comparison. The secondary table below removes overlaps across labels and windows for flat NER. GLiNER2 candidates are ranked by native confidence; saved comparator outputs lack confidences, so their ties use longer spans, then offsets and label names. Generative BIO output is usually already flat. This is a post-hoc protocol check, not a leaderboard submission.
| Benchmark: flat-output schema transfer | This model | GLiNER-bi-large-v2.0 | GNER-T5-xxl | NuExtract-2.0-8B |
|---|---|---|---|---|
| crossner_ai | 13.86 | 47.71 | 51.13 | 33.88 |
| crossner_literature | 17.62 | 58.08 | 58.45 | 47.89 |
| crossner_music | 17.60 | 73.50 | 77.49 | 76.67 |
| crossner_politics | 24.88 | 50.72 | 68.32 | 44.14 |
| crossner_science | 17.12 | 63.94 | 64.73 | 49.45 |
| mit_movie | 25.60 | 41.38 | 53.54 | 50.00 |
| mit_restaurant | 2.12 | 42.24 | 41.59 | 30.07 |
| bc5cdr | 50.00 | 57.65 | 59.65 | 61.64 |
| finer_ord | 10.99 | 14.29 | 55.90 | 64.12 |
| inlegalner | 2.07 | 23.77 | 19.35 | 19.92 |
| biodivner | 8.12 | 26.25 | 27.30 | 16.00 |
The familiar-label science schema was corrected to use the exact training definitions “Chemical element from periodic table” and “Scientific theory or hypothesis”. Earlier selection by definition frequency chose unrelated senses. Corrected completed cells replace those two-definition results; the benchmark-schema condition was unaffected. Residual categories such as person/organization can still differ in scope between annotation schemes, so familiar-label results remain schema-transfer evidence rather than a proof of identical annotation policy.
Expanded evaluation and decision-layer experiments
A separate, larger holdout excludes normalized text duplicates from the earlier sampled panels and development set. The experiment report records both familiar-label transfer and new-specific-type transfer, current GLiNER2.5 comparisons, specialist in-domain results, and completion status. Missing cells remain pending.
| Fresh seven-dataset core suite | F1 (%) |
|---|---|
| Familiar labels: mean dataset F1 | pending |
| Benchmark schema: mean dataset F1 | pending |
| New specific types: pooled F1 | pending |
The repaired pipeline adds Mapika/decider-4b v2.1 to this checkpoint. Pipeline gains, correction damage, paired intervals and measured request latency are separate from this model’s own scores; weights in this repository were not updated. The primary held-out repair condition was fixed before fresh inference.
GLiNER2.5 on the original frozen core panel
These use the same earlier seven core-dataset samples, concise schemas, strict offsets and threshold 0.5. Boundary models use their native flat decoding; legacy GLiNER2 can return overlapping labels.
| Model | Familiar-label mean F1 | Benchmark-schema mean F1 |
|---|---|---|
| Siddharth63/gliner2-small | 33.84 | 17.29 |
| fastino/gliner2.5-small-v1 | 40.81 | 41.28 |
| fastino/gliner2.5-base-v1 | 47.71 | 51.53 |
| fastino/gliner2.5-multi-v1 | 47.04 | 50.67 |
JPT access and results, if available, are recorded separately in the report. Published paper scores are not substituted for measurements under this protocol. This expanded run makes no state-of-the-art or leaderboard-entry claim.
- Downloads last month
- 122