Instructions to use evalengine/decision-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use evalengine/decision-0.8b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-0.8B") model = PeftModel.from_pretrained(base_model, "evalengine/decision-0.8b") - Notebooks
- Google Colab
- Kaggle
Download evaluation/RESULTS.md from evalengine/decision-0.8b: direct link, hf CLI and curl.
- Browser
- Download file 4.89 kB
-
https://huggingface.co/evalengine/decision-0.8b/resolve/main/evaluation/RESULTS.md
- Command line
-
hf download hf://evalengine/decision-0.8b/evaluation/RESULTS.md
-
curl -L -o RESULTS.md https://huggingface.co/evalengine/decision-0.8b/resolve/main/evaluation/RESULTS.md
Decision-0.8B: measured comparison with Tev-0.8B
Selected adapter: runs/decision-0.8b-v2/adapter (step 9,289).
Selection used only the 892-case development panel. Model identities were frozen before final test inference.
| Panel | Cases | Qwen-0.8B | Tev-0.8B | Decision-0.8B |
|---|---|---|---|---|
| Fresh public holdout: accuracy | 500 | 56.20% | 68.20% | 79.60% |
| Fresh public holdout: family macro | 500 | 56.20% | 68.20% | 79.60% |
| Historical nine-family suite: accuracy | 2800 | 43.89% | 55.54% | 68.82% |
| Historical nine-family suite: family macro | 2800 | 39.81% | 51.42% | 62.61% |
| Reused development: accuracy | 892 | 44.62% | 71.86% | 75.90% |
| Reused original-core development: accuracy | 442 | — | 84.39% | 83.94% |
Paired uncertainty (Decision minus Tev; percentage points):
- fresh: family macro +11.40 pp; paired group-bootstrap 95% interval [+7.40, +15.40] pp.
- historical: family macro +11.19 pp; paired group-bootstrap 95% interval [+9.42, +13.01] pp.
Fresh-holdout source results
| Source | Cases | Tev | Decision | Difference |
|---|---|---|---|---|
| clinc150 | 100 | 83.00% | 89.00% | +6.00 pp |
| go_emotions | 100 | 69.00% | 82.00% | +13.00 pp |
| helpsteer2 | 100 | 56.00% | 47.00% | -9.00 pp |
| paws | 100 | 57.00% | 92.00% | +35.00 pp |
| vitaminc | 100 | 76.00% | 88.00% | +12.00 pp |
All development source results (selection diagnostics)
| Source | Cases | Tev | Decision | Difference |
|---|---|---|---|---|
| ag_news | 50 | 94.00% | 84.00% | -10.00 pp |
| banking77 | 50 | 80.00% | 92.00% | +12.00 pp |
| boolq | 50 | 86.00% | 86.00% | +0.00 pp |
| clinc150 | 50 | 74.00% | 94.00% | +20.00 pp |
| go_emotions | 50 | 70.00% | 88.00% | +18.00 pp |
| helpsteer2 | 50 | 54.00% | 54.00% | +0.00 pp |
| mnli | 50 | 82.00% | 86.00% | +4.00 pp |
| paws | 50 | 66.00% | 96.00% | +30.00 pp |
| policy | 48 | 100.00% | 100.00% | +0.00 pp |
| policy_v2 | 48 | 85.42% | 87.50% | +2.08 pp |
| research_taxonomy_v21 | 48 | 100.00% | 100.00% | +0.00 pp |
| routing_v2 | 48 | 70.83% | 75.00% | +4.17 pp |
| sst5 | 50 | 62.00% | 46.00% | -16.00 pp |
| vitaminc | 50 | 68.00% | 88.00% | +20.00 pp |
| workflow.simple_quorum | 100 | 56.00% | 52.00% | -4.00 pp |
| workflow.veto_alternative | 100 | 46.00% | 44.00% | -2.00 pp |
Training and artifacts
The v2 full run completed 9,289 optimizer steps and 29,185,870 tokens in 26.1 minutes (training/evaluation loop only), with 51.53 GiB peak allocated VRAM.
Recipe: Qwen3.5-0.8B, BF16, rank-8 LoRA, alpha 16, learning rate 5e-5, one epoch, microbatch 8, accumulation 1, seed 42, 2048-token context. Batch-eight numerical parity passed the pre-existing thresholds before training.
Selected identity: cb80a8e8a5ffd1232abaecd37a3f91a58c0f268d218069ce4cb779f70158aff8.
Model checkpoints, tokenizer files, input datasets and evaluation reports have content hashes. See selection.json, final/freeze.json, final/comparison.json and the candidate preflight/summary.
The earlier v1 run was intentionally interrupted after preserving a 500-step diagnostic checkpoint; it is not counted as a completed epoch.
Warm local latency
Warm local Python decision call: prompt rendering, tokenization, host/device transfer, forward, option scoring, CPU result; excludes model load and HTTP/server/Ollama overhead. CUDA workstation, not Mac.
| Model | E2E median | E2E p95 | Forward median |
|---|---|---|---|
| tev | 14.23 ms | 41.18 ms | 13.02 ms |
| decision | 20.48 ms | 48.67 ms | 19.31 ms |
Export status
The adapter, F16 and Q8_0 GGUF builds are published. Q4_K_M failed the unchanged deployment accuracy gate and is not published. Earlier probability-equivalence failures remain documented in export verification. All benchmark and timing scores here belong to the adapter.
Scope and limitations
The primary 500-case holdout has 100 previously unused local groups from each of five public task families. It was constructed before candidate test inference and excludes previously used local IDs, groups, normalized states and prompts. This does not rule out pretraining or competitor-training overlap. The historical 2800-case suite was previously used for 4B reporting. Its scores are regression evidence, not a newly blind test. Development results, including the 442 original-core cases, were reused for selection. The comparison scores every declared option using a common prompt, disabled thinking and local BF16 inference. It evaluates downloaded checkpoints, not hosted-endpoint parity or unconstrained-generation formatting. These results do not establish universal superiority. Report every source regression. Bootstrap intervals use 2000 paired whole-group draws and exclude training-seed uncertainty. HelpSteer2 uses held-out original training prompts; VitaminC grouping is not article separation. No physical Mac/Ollama or iPhone latency was measured.