evalexplorer-classify-gliner2.5-base

Classifies an international development evaluation report from its first pages into a JSON object with evaluation_approach, evaluation_type, temporality, themes and countries. One of the runs documented in baobabtech/evalexplorer-classify-experiments; code in code/.

Base model fastino/gliner2.5-base-v1
Method Full fine-tune
Data baobabtech/evalexplorer-data, config classify_codes, 1148 training documents
Hardware 1× NVIDIA A100-SXM4-80GB (79 GB)
Training time 13.8 min
Run report gliner2.5-base--test

Training

Full fine-tuning of the encoder and task heads: 5 epochs, batch 16, encoder lr 1e-05, task lr 0.0005, best checkpoint by validation loss.

Scores

Metric Test (134 documents)
JSON valid 1.000
Exact match (all five fields) 0.022
Mean field score 0.578
evaluation_approach accuracy 0.313
evaluation_type accuracy 0.604
temporality accuracy 0.552
themes micro F1 0.610
countries micro F1 0.756

Greedy decoding, thinking off. mean_field_score is the per-document mean of the five field scores (1/0 for the single-code fields, F1 for the lists). The model was trained on, and is scored here against, the EvalExplorer ingestion pipeline's LLM labels, which no person has reviewed, so these numbers measure agreement with that pipeline, not correctness. The run report also scores the same answers against an independent GLM-5.3-Flash relabelling the model never saw. This is an exploratory model; the intended next version is trained on the GLM labels.

Usage

from gliner2 import AutoExtractor

model = AutoExtractor.from_pretrained("baobabtech/evalexplorer-classify-gliner2.5-base")
tasks = {"evaluation_approach": [...], "evaluation_type": [...], "temporality": [...],
         "themes": {"labels": [...], "multi_label": True}}   # codes as listed in the dataset prompt
print(model.classify_text_long(first_pages_text, tasks, chunk_size=384, chunk_overlap=64))
print(model.extract_entities_long(first_pages_text, {"country": "Country the evaluation focuses on"}))
Downloads last month
1
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for baobabtech/evalexplorer-classify-gliner2.5-base

Finetuned
(3)
this model

Collection including baobabtech/evalexplorer-classify-gliner2.5-base