GLiClass Knowledge Classifier v2

Tag English passages with five knowledge facets: decisions, constraints, procedures, mechanisms and traps.

A GLiClass Base v3 fine-tune that tags English passages by what they contain: a trap, decision, constraint, mechanism or procedure. It scores all five facets in one pass, allows several per passage, and runs locally through ONNX Runtime. Use it to organize notes and documentation, or to see what kind of material a search returned.

On 1,300 generated passages from three document families excluded from training, macro average precision is 0.9676 for the fine-tune, 0.9384 for a TF-IDF baseline trained on the same rows, and 0.7709 for the zero-shot upstream checkpoint.

This revision's ONNX graph adds Vulkan compatibility. Learned weights, prompts, calibration and thresholds are unchanged.

  • Input: one passage, up to 768 tokens including the label prompts
  • Output: five calibrated scores, with two threshold tables for turning them into labels
  • Runtimes: ONNX Runtime 1.24.4 on CPU and CUDA; Vulkan through the ONNX Runtime WebGPU plugin
  • Size: about 461.5 MB for the package, excluding runtime libraries; the ONNX graph is 452.8 MB
  • License: Apache-2.0

Quick start

Install the runtime dependencies and download the package:

pip install onnxruntime==1.24.4 tokenizers numpy huggingface_hub
hf download Daecore/gliclass-knowledge-classifier-v2 --local-dir downloaded-model

This CPU example uses the package's published prompts and calibration rather than new label wording:

import json
from pathlib import Path
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer

root = Path("downloaded-model")
metadata = json.load(open(f"{root}/classifier-metadata.json", encoding="utf-8"))
facets = metadata["facet_order"]
prep = metadata["preprocessing"]
prefix = "".join(
    prep["label_token"] + prep["label_prompts"][facet] for facet in facets
) + prep["separator_token"]
tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json")
tokenizer.enable_truncation(max_length=prep["max_length"])
encoded = tokenizer.encode(prefix + "We chose weekly releases to reduce rollout risk.")
session = ort.InferenceSession(
    f"{root}/model.onnx", providers=["CPUExecutionProvider"]
)
logits = session.run(["logits"], {
    "input_ids": np.array([encoded.ids], dtype=np.int64),
    "attention_mask": np.array([encoded.attention_mask], dtype=np.int64),
})[0][0].astype(np.float64)
calibration = metadata["probability_calibration"]
clip = calibration["probability_clip"]
raw = np.clip(1 / (1 + np.exp(-np.clip(logits, -80, 80))), clip, 1 - clip)
temperatures = np.array([calibration["temperatures"][f] for f in facets])
probabilities = 1 / (1 + np.exp(-np.log(raw / (1 - raw)) / temperatures))
print(dict(zip(facets, probabilities.tolist())))

The output is five independent calibrated scores, not a distribution that sums to one. For applications that need labels, the metadata supplies two threshold tables: contract favors precision and recall_leaning retains more candidates. Choose between them by the relative cost of missed and incorrect labels.

What the labels mean

Facet The passage…
trap Warns about a specific mistake, hazard or failure mode
decision Records a choice or commitment among alternatives
constraint States a rule or condition that materially restricts action
mechanism Explains how something is organized, connected or works
procedure Gives sequenced actions for carrying out a task

Mentioning a decision is different from making one; training labels treat “mentions only” as negative for that facet. Facets overlap: a procedure may also contain a constraint and warn about a trap. Use the scores to describe passages, not to judge them: they do not establish relevance, truth or authority and should not serve as a hard retrieval filter.

Before and after fine-tuning

“Upstream” is the exact knowledgator/gliclass-base-v3.0 checkpoint used to start training, applied zero-shot with the same label definitions and no Daecore fine-tuning. The word-feature baseline is TF-IDF plus logistic regression, trained on the same labeled rows as the fine-tune. All three use the same 1,300-passage panel; unresolved labels are excluded per facet for every model.

Metric, macro average over five facets Upstream GLiClass Word-feature baseline Daecore fine-tune
Average precision 0.7709 0.9384 0.9676
Precision at 90% recall 66.11% 83.35% 90.21%
Precision at 95% recall 64.51% 78.62% 86.02%

Average precision summarizes how well scores rank positives across thresholds. Precision at 90% recall is the cleanest measured prediction set that still keeps at least 90% of positive labels; it comes from this evaluation's curve, not from a threshold chosen in advance.

Facet Resolved passages Positive prevalence Upstream AP Fine-tuned AP
Trap 1,300 71.77% 0.9028 0.9864
Decision 1,299 46.04% 0.5219 0.9135
Constraint 1,300 91.31% 0.9544 0.9986
Mechanism 1,297 68.77% 0.7692 0.9817
Procedure 1,293 37.51% 0.7064 0.9578

Upstream, word-feature and fine-tuned classifier results

As the baseline shows, much of this task can be learned from word cues; the fine-tune's clearest benefit is at high recall. Constraint is common in this panel, so its near-perfect AP says less than the decision and procedure results. The panel was deliberately enriched for difficult decision and procedure cases and is not a sample of natural traffic. The evaluation companion contains the text-free labels and scores, model identities and calculation code.

Data and training choices

The final fit used 59,886 passages from 9,376 documents: 54,803 generated and 5,083 real. The real portion is 4,583 passages from self-owned workspaces and 500 from public-domain US Federal Register documents; all real rows were used for training, none for evaluation. Workspace documents may also be AI-written; real describes their source, not human authorship. Generated material was written as complete organizational documents and then parsed into passages, covering more settings and document types than the available workspaces.

Models applied fixed facet definitions to produce the labels. The main label campaigns used independent judgments with conflict resolution; 3,750 inherited rows used a single primary judge with a separate audit. Splits separate whole source families: the three evaluation families and four calibration families are disjoint from training and from each other. Calibration uses 1,700 passages to set per-facet temperatures and the two threshold tables.

Training choice Value and reason
Starting model GLiClass Base v3 on DeBERTa-v3, 186.5M parameters
Input length 768 tokens including label prompts; context-length experiments favored it over 512
Objective Weighted binary cross-entropy; mentions-only negatives receive 2× weight
Optimizer schedule Learning rate 2e-5; batch 4, accumulated over 8 steps
Duration Four fixed epochs; equal-weight average of epochs 2, 3 and 4, chosen before the final fit
Serving compression INT8 token embeddings with the transformer body in FP32

Quantizing only the token embeddings shrinks the large vocabulary table while keeping transformer calculations in FP32; broader quantization reduced quality.

Runtime and limits

On the 1,300 evaluation passages, the derived graph produced the same labels under both threshold tables on CPU, CUDA and Vulkan as the previous graph on CPU, with maximum score differences below 5.7e-6.

Vulkan uses onnxruntime==1.24.4 with the native WebGPU plugin onnxruntime-ep-webgpu==0.4.0, registered explicitly, with dawnBackendType=Vulkan and no providers list at session creation. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.

  • The model is trained and evaluated only on English text.
  • Labels are model judgments without human validation.
  • The evaluation covers generated text from three held-out families and was reused during development; performance on unrelated real documents is unmeasured.
  • Training was run once; variation across seeds is unmeasured.

License

Apache-2.0. The package includes upstream attribution and a modification notice.

Downloads last month
72
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Daecore/gliclass-knowledge-classifier-v2

Finetuned
(1)
this model