--- license: apache-2.0 base_model: knowledgator/gliclass-base-v3.0 base_model_relation: finetune pipeline_tag: text-classification language: - en tags: - gliclass - deberta-v3 - multi-label-classification - fine-tuned - onnx - int8 - vulkan --- # GLiClass Knowledge Classifier v2 **Tag English passages with five knowledge facets: decisions, constraints, procedures, mechanisms and traps.** A GLiClass Base v3 fine-tune that tags English passages by what they contain: a **trap, decision, constraint, mechanism or procedure**. It scores all five facets in one pass, allows several per passage, and runs locally through ONNX Runtime. Use it to organize notes and documentation, or to see what kind of material a search returned. On 1,300 generated passages from three document families excluded from training, macro average precision is **0.9676** for the fine-tune, 0.9384 for a TF-IDF baseline trained on the same rows, and 0.7709 for the zero-shot upstream checkpoint. This revision's ONNX graph adds Vulkan compatibility. Learned weights, prompts, calibration and thresholds are unchanged. - **Input:** one passage, up to 768 tokens including the label prompts - **Output:** five calibrated scores, with two threshold tables for turning them into labels - **Runtimes:** ONNX Runtime 1.24.4 on CPU and CUDA; Vulkan through the ONNX Runtime WebGPU plugin - **Size:** about 461.5 MB for the package, excluding runtime libraries; the ONNX graph is 452.8 MB - **License:** Apache-2.0 ## Quick start Install the runtime dependencies and download the package: ```sh pip install onnxruntime==1.24.4 tokenizers numpy huggingface_hub hf download Daecore/gliclass-knowledge-classifier-v2 --local-dir downloaded-model ``` This CPU example uses the package's published prompts and calibration rather than new label wording: ```python import json from pathlib import Path import numpy as np import onnxruntime as ort from tokenizers import Tokenizer root = Path("downloaded-model") metadata = json.load(open(f"{root}/classifier-metadata.json", encoding="utf-8")) facets = metadata["facet_order"] prep = metadata["preprocessing"] prefix = "".join( prep["label_token"] + prep["label_prompts"][facet] for facet in facets ) + prep["separator_token"] tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json") tokenizer.enable_truncation(max_length=prep["max_length"]) encoded = tokenizer.encode(prefix + "We chose weekly releases to reduce rollout risk.") session = ort.InferenceSession( f"{root}/model.onnx", providers=["CPUExecutionProvider"] ) logits = session.run(["logits"], { "input_ids": np.array([encoded.ids], dtype=np.int64), "attention_mask": np.array([encoded.attention_mask], dtype=np.int64), })[0][0].astype(np.float64) calibration = metadata["probability_calibration"] clip = calibration["probability_clip"] raw = np.clip(1 / (1 + np.exp(-np.clip(logits, -80, 80))), clip, 1 - clip) temperatures = np.array([calibration["temperatures"][f] for f in facets]) probabilities = 1 / (1 + np.exp(-np.log(raw / (1 - raw)) / temperatures)) print(dict(zip(facets, probabilities.tolist()))) ``` The output is five independent calibrated scores, not a distribution that sums to one. For applications that need labels, the metadata supplies two threshold tables: `contract` favors precision and `recall_leaning` retains more candidates. Choose between them by the relative cost of missed and incorrect labels. ## What the labels mean | Facet | The passage… | |---|---| | `trap` | Warns about a specific mistake, hazard or failure mode | | `decision` | Records a choice or commitment among alternatives | | `constraint` | States a rule or condition that materially restricts action | | `mechanism` | Explains how something is organized, connected or works | | `procedure` | Gives sequenced actions for carrying out a task | Mentioning a decision is different from making one; training labels treat “mentions only” as negative for that facet. Facets overlap: a procedure may also contain a constraint and warn about a trap. Use the scores to describe passages, not to judge them: they do not establish relevance, truth or authority and should not serve as a hard retrieval filter. ## Before and after fine-tuning “Upstream” is the exact `knowledgator/gliclass-base-v3.0` checkpoint used to start training, applied zero-shot with the same label definitions and **no Daecore fine-tuning**. The word-feature baseline is TF-IDF plus logistic regression, trained on the same labeled rows as the fine-tune. All three use the same 1,300-passage panel; unresolved labels are excluded per facet for every model. | Metric, macro average over five facets | Upstream GLiClass | Word-feature baseline | Daecore fine-tune | |---|---:|---:|---:| | Average precision | 0.7709 | 0.9384 | **0.9676** | | Precision at 90% recall | 66.11% | 83.35% | **90.21%** | | Precision at 95% recall | 64.51% | 78.62% | **86.02%** | Average precision summarizes how well scores rank positives across thresholds. Precision at 90% recall is the cleanest measured prediction set that still keeps at least 90% of positive labels; it comes from this evaluation's curve, not from a threshold chosen in advance. | Facet | Resolved passages | Positive prevalence | Upstream AP | Fine-tuned AP | |---|---:|---:|---:|---:| | Trap | 1,300 | 71.77% | 0.9028 | **0.9864** | | Decision | 1,299 | 46.04% | 0.5219 | **0.9135** | | Constraint | 1,300 | 91.31% | 0.9544 | **0.9986** | | Mechanism | 1,297 | 68.77% | 0.7692 | **0.9817** | | Procedure | 1,293 | 37.51% | 0.7064 | **0.9578** | ![Upstream, word-feature and fine-tuned classifier results](figures/classifier-comparison.svg) As the baseline shows, much of this task can be learned from word cues; the fine-tune's clearest benefit is at high recall. Constraint is common in this panel, so its near-perfect AP says less than the decision and procedure results. The panel was deliberately enriched for difficult decision and procedure cases and is not a sample of natural traffic. The [evaluation companion](evaluation/README.md) contains the text-free labels and scores, model identities and calculation code. ## Data and training choices The final fit used **59,886 passages from 9,376 documents**: 54,803 generated and 5,083 real. The real portion is 4,583 passages from self-owned workspaces and 500 from public-domain US Federal Register documents; all real rows were used for training, none for evaluation. Workspace documents may also be AI-written; real describes their source, not human authorship. Generated material was written as complete organizational documents and then parsed into passages, covering more settings and document types than the available workspaces. Models applied fixed facet definitions to produce the labels. The main label campaigns used independent judgments with conflict resolution; 3,750 inherited rows used a single primary judge with a separate audit. Splits separate whole source families: the three evaluation families and four calibration families are disjoint from training and from each other. Calibration uses 1,700 passages to set per-facet temperatures and the two threshold tables. | Training choice | Value and reason | |---|---| | Starting model | GLiClass Base v3 on DeBERTa-v3, 186.5M parameters | | Input length | 768 tokens including label prompts; context-length experiments favored it over 512 | | Objective | Weighted binary cross-entropy; mentions-only negatives receive 2× weight | | Optimizer schedule | Learning rate 2e-5; batch 4, accumulated over 8 steps | | Duration | Four fixed epochs; equal-weight average of epochs 2, 3 and 4, chosen before the final fit | | Serving compression | INT8 token embeddings with the transformer body in FP32 | Quantizing only the token embeddings shrinks the large vocabulary table while keeping transformer calculations in FP32; broader quantization reduced quality. ## Runtime and limits On the 1,300 evaluation passages, the derived graph produced the same labels under both threshold tables on CPU, CUDA and Vulkan as the previous graph on CPU, with maximum score differences below 5.7e-6. Vulkan uses `onnxruntime==1.24.4` with the native WebGPU plugin [`onnxruntime-ep-webgpu==0.4.0`](https://pypi.org/project/onnxruntime-ep-webgpu/), registered explicitly, with `dawnBackendType=Vulkan` and no providers list at session creation. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available. - The model is trained and evaluated only on English text. - Labels are model judgments without human validation. - The evaluation covers generated text from three held-out families and was reused during development; performance on unrelated real documents is unmeasured. - Training was run once; variation across seeds is unmeasured. ## License Apache-2.0. The package includes upstream attribution and a modification notice.