Instructions to use baobabtech/evalexplorer-classify-gliner2.5-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use baobabtech/evalexplorer-classify-gliner2.5-small with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("baobabtech/evalexplorer-classify-gliner2.5-small") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
evalexplorer-classify-gliner2.5-small
Classifies an international development evaluation report from its first pages into a JSON object with evaluation_approach, evaluation_type, temporality, themes and countries. One of the runs documented in baobabtech/evalexplorer-classify-experiments; code in code/.
| Base model | fastino/gliner2.5-small-v1 |
| Method | Full fine-tune |
| Data | baobabtech/evalexplorer-data, config classify_codes, 1148 training documents |
| Hardware | 1× NVIDIA A100-SXM4-80GB (79 GB) |
| Training time | 8.2 min |
| Run report | gliner2.5-small--test |
Training
Full fine-tuning of the encoder and task heads: 5 epochs, batch 16, encoder lr 1e-05, task lr 0.0005, best checkpoint by validation loss.
Scores
| Metric | Test (134 documents) |
|---|---|
| JSON valid | 1.000 |
| Exact match (all five fields) | 0.022 |
| Mean field score | 0.573 |
| evaluation_approach accuracy | 0.254 |
| evaluation_type accuracy | 0.575 |
| temporality accuracy | 0.597 |
| themes micro F1 | 0.634 |
| countries micro F1 | 0.758 |
Greedy decoding, thinking off. mean_field_score is the per-document mean of the five field scores (1/0 for the single-code fields, F1 for the lists). The model was trained on, and is scored here against, the EvalExplorer ingestion pipeline's LLM labels, which no person has reviewed, so these numbers measure agreement with that pipeline, not correctness. The run report also scores the same answers against an independent GLM-5.3-Flash relabelling the model never saw. This is an exploratory model; the intended next version is trained on the GLM labels.
Usage
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("baobabtech/evalexplorer-classify-gliner2.5-small")
tasks = {"evaluation_approach": [...], "evaluation_type": [...], "temporality": [...],
"themes": {"labels": [...], "multi_label": True}} # codes as listed in the dataset prompt
print(model.classify_text_long(first_pages_text, tasks, chunk_size=384, chunk_overlap=64))
print(model.extract_entities_long(first_pages_text, {"country": "Country the evaluation focuses on"}))
- Downloads last month
- 3
Model tree for baobabtech/evalexplorer-classify-gliner2.5-small
Base model
fastino/gliner2.5-small-v1