Instructions to use evalengine/decision-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use evalengine/decision-0.8b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-0.8B") model = PeftModel.from_pretrained(base_model, "evalengine/decision-0.8b") - Notebooks
- Google Colab
- Kaggle
Download evaluation/EXPORT_VERIFICATION.md from evalengine/decision-0.8b: direct link, hf CLI and curl.
- Browser
- Download file 1.74 kB
-
https://huggingface.co/evalengine/decision-0.8b/resolve/main/evaluation/EXPORT_VERIFICATION.md
- Command line
-
hf download hf://evalengine/decision-0.8b/evaluation/EXPORT_VERIFICATION.md
-
curl -L -o EXPORT_VERIFICATION.md https://huggingface.co/evalengine/decision-0.8b/resolve/main/evaluation/EXPORT_VERIFICATION.md
Export verification
All deployment checks use the same fixed 892-case development panel, maximum context 2,048 tokens, raw option probabilities, and complete finite logits. Prompt token IDs, option boundaries and EOS IDs are checked for every GGUF case. This is a workstation fidelity check, not a new blind test or a phone measurement.
The established release gates allow at most 0.5 percentage points of accuracy loss for the merged BF16 model versus the adapter, and at most 1 point for GGUF versus merged BF16. The same thresholds were retained for every attempt.
| Artifact | Dev accuracy |
|---|---|
| BF16 adapter | 75.90% |
| Merged BF16 | 76.23% |
| F16 | 75.45% |
| Q8_0 | 75.78% |
| Q4_K_M (failed; not published) | 72.65% |
Merged BF16, F16 and Q8_0 passed. Q4_K_M failed with a 3.59-point drop versus merged BF16 and is not published. Q8_0 was tested as a higher-precision fallback after that failure; format selection used only development results.
Two earlier 32-case merge diagnostics failed a separate, stricter maximum absolute probability-difference tolerance of 0.02. Direct BF16 merge: 0.0200787; FP32 merge arithmetic followed by BF16 output: 0.0399304. Neither changed the sampled argmax predictions. Those failures remain failures. The established full-panel release procedure checks classification accuracy, not identical probabilities; passing it does not resolve the failed probability-equivalence diagnostics.
Probabilities remain uncalibrated, and calibration does not transfer automatically between formats. Historical and fresh test scores on the model card belong to the adapter. F16 and Q8_0 have not been evaluated on those test panels. Physical Mac, iPhone and Ollama latency have not been measured.