GLiNER2.5-Decide-TR

Turkish-focused fine-tune of fastino/GLiNER2.5-multi-Decide (287M, mDeBERTa-v3-base). It is a fast, calibrated first-stage decision layer for agent/LLM systems. It decides directly when confident and hands uncertain cases to an LLM.

  • Question types: choice (single or multi-label), yes/no, and ordinal score (computed from choice probabilities).
  • Tasks: intent routing, agent/tool/MCP selection with "none", guardrails (prompt injection, violence, fraud), RAG passage relevance and answerability, injection inside documents, PII presence, language ID, tone/urgency, multi-turn routing.
  • Languages: Turkish first, English, plus 15 other languages.
  • ~22 ms per decision on a consumer GPU · fp16, 564 MB · Apache-2.0.

Base vs. this model

Same test sets and fixed thresholds for both models. None of the test items were seen in training.

Task base this model
General Turkish decisions (1,568) 63.3 82.1
Tool/agent catalog, 9–41 options (TR / EN) 64.5 / 68.4 92.6 / 97.9
Correct answer is "none" 6.1 97.0
MCP server + tool selection 55 83
Banking intent routing 51 89
Guardrail: correct without LLM fallback 10 90
RAG passage relevance AUROC .63 .98
RAG answerability AUROC .58 .95
Prompt injection inside documents, AUROC .58 .99
PII present (F1 / false-alarm rate) .39 / .10 .94 / .00
Language identification (12 languages) 31 98
Claim grounded in passage (3-way) 34 78
Tone / urgency (ordinal) 56 90
Multi-turn routing 53 79
General intent, non-banking domains 73 91

Usage

from gliner2 import AutoExtractor
import json, math

model = AutoExtractor.from_pretrained("sevket09/GLiNER2.5-Decide-TR")
cal = json.load(open("calibration.json"))["params"]   # from this repo
logit = lambda p: math.log(max(p, 1e-6) / max(1 - p, 1e-6))

def choice(text, labels, prompt, descriptions=None):
    schema = model.create_schema().classification(
        "q", descriptions or labels, multi_label=True, cls_threshold=0.0, prompt=prompt)
    out = model.extract(text, schema, include_confidence=True)["q"]
    s = {x["label"]: logit(x["confidence"]) / cal["T"] for x in out}
    z = sum(math.exp(v) for v in s.values())
    return {k: math.exp(v) / z for k, v in s.items()}           # calibrated probabilities

def yes_no(text, question):
    schema = model.create_schema().classification("q", ["yes", "no"], multi_label=True, cls_threshold=0.0, prompt=question)
    c = {x["label"]: x["confidence"] for x in model.extract(text, schema, include_confidence=True)["q"]}
    return 1 / (1 + math.exp(-cal["platt_a"] * (logit(c["yes"]) - logit(c["no"]))))   # P(yes)

teams = {"card": "Credit and debit cards (not money transfers)",
         "transfer": "EFT, FAST and havale transfers, including transfer limits",
         "loan": "Loans",
         "none": "Unrelated to banking"}
p = choice("FAST limitimi öğrenmek istiyorum", list(teams), "Which team should handle the customer's request?", teams)
label, conf = max(p.items(), key=lambda kv: kv[1])
decision = label if conf >= 0.70 else "ESCALATE_TO_LLM"

# Score: ordinal levels as a choice, then the probability-weighted mean
levels = ["calm", "annoyed", "angry", "furious"]
p = choice("Üç gündür kimse dönmedi, rezalet!", levels, "How angry is the customer?")
score = sum(i * p[l] for i, l in enumerate(levels)) / (len(levels) - 1)   # 0..1

Ask unrelated questions in separate calls: mixing different question types in one pass lowers accuracy.

Writing label descriptions

Descriptions matter most where two labels overlap. Keep them short and state the boundary:

agents = {
    "card_agent":     "Card limit, statement, card debt, PIN, new card",
    "security_agent": "Lost or stolen card, transactions the customer did not make, suspicious login",
    "none":           "Request unrelated to all agents above",
}
  • Mention who handles the overlapping case (lost card → security, not card).
  • Short and distinctive beats long and exhaustive: adding many keywords to one label can pull neighbouring requests to it.
  • Keep "none" for truly out-of-scope requests and say so in its description.

Limitations

  • Test sets are synthetic or hand-written, mostly Turkish banking and general domains. Evaluate on your own data and re-fit thresholds before automating consequential decisions.

Training data and licenses

Base model: Apache-2.0. Training data:

  • FineWeb / FineWeb-2: ODC-By.
  • MASSIVE, Banking77: CC BY 4.0.
  • CLINC150: CC BY 3.0.
  • deepset prompt-injections, toxic-TR, fast-decisions: Apache-2.0.
  • Synthetic data generated with deepseek-v4-flash (MIT).

No share-alike or non-commercial data. Outputs of proprietary models (Gemini, Claude, GPT) were used only for test sets.

Downloads last month
10
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sevket09/GLiNER2.5-Decide-TR

Finetuned
(3)
this model