Download evaluation/README.md from Daecore/gliclass-knowledge-classifier-v2: direct link, hf CLI and curl.
- Browser
- Download file 3.23 kB
-
https://huggingface.co/Daecore/gliclass-knowledge-classifier-v2/resolve/main/evaluation/README.md
- Command line
-
hf download hf://Daecore/gliclass-knowledge-classifier-v2/evaluation/README.md
-
curl -L -o README.md https://huggingface.co/Daecore/gliclass-knowledge-classifier-v2/resolve/main/evaluation/README.md
Five-facet classifier evaluation
These records support the model card's results, comparing the fine-tuned classifier with upstream GLiClass and a word-feature baseline on trap, decision, constraint, mechanism and procedure labels. The classifier describes a passage's content; it does not score retrieval relevance or truth.
| File | Contents |
|---|---|
classifier.json |
Labels and all three models' scores for 1,300 passages |
classifier-summary.json |
The summary that metrics.py reproduces |
serving-qualification.json |
CPU, CUDA and Vulkan serving checks |
metrics.py, figures.py, svg_figures.py |
Metric and figure code |
Reproduce the results
From this directory, run with Python 3.11 or later:
python metrics.py --directory .
python figures.py --output ../figures
No packages or network access are needed. classifier.json contains anonymized labels, scores and exact model identities. The first command reproduces classifier-summary.json; the second redraws the model card's figure. No private passage text, source identifiers or workspace paths are included. Re-running inference would require the private passages.
Comparison
All three models use the same 1,300 generated passages from three source families excluded from training. Unresolved labels are excluded identically for every model, separately for each facet. The panel was enriched for difficult decision and procedure cases and does not represent natural traffic.
Upstream uses the same five label definitions and 768-token limit as the fine-tune. Raw upstream logits and calibrated fine-tune probabilities rank examples within each facet; their numerical scales are not compared. The word-feature baseline uses unigram and bigram TF-IDF and one balanced logistic-regression model per facet, fitted on the same 59,886 training passages. It sees full passage text, while the neural models apply their token limit.
Average precision is the area under the stepwise precision–recall curve, grouping tied scores at a single threshold. Macro average precision weights the five facets equally. Precision at 90% or 95% recall is the best measured precision at any threshold reaching that recall; the evaluation curve supplies that threshold, so it is not an advance operating-point guarantee.
Serving checks
serving-qualification.json records checks on the released graph across CPU, CUDA and Vulkan. All 1,300 predictions matched the reference within 5.62e-6, and neither threshold table changed any labels. Learned weights, prompts, calibration and thresholds are unchanged by the Vulkan graph rewrite.
Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB). Other GPUs and Linux remain untested. Classification is an independent function; using this model does not require Daecore's embedder or reranker.
Limits
Labels are model judgments without a human reference. The evaluation covers generated text, was reused during development and has unusually common constraint labels. A strong word-feature baseline shows that much of the task can be learned from wording cues. Performance on unrelated documents and variation across training seeds are unmeasured.