decision-0.8b / evaluation /RESULTS.md
johnsonchromia's picture
Publish Decision-0.8B and verified benchmarks
cb17a40 verified
|
Raw History Blame Contribute Delete
4.89 kB

Decision-0.8B: measured comparison with Tev-0.8B

Selected adapter: runs/decision-0.8b-v2/adapter (step 9,289). Selection used only the 892-case development panel. Model identities were frozen before final test inference.

Panel Cases Qwen-0.8B Tev-0.8B Decision-0.8B
Fresh public holdout: accuracy 500 56.20% 68.20% 79.60%
Fresh public holdout: family macro 500 56.20% 68.20% 79.60%
Historical nine-family suite: accuracy 2800 43.89% 55.54% 68.82%
Historical nine-family suite: family macro 2800 39.81% 51.42% 62.61%
Reused development: accuracy 892 44.62% 71.86% 75.90%
Reused original-core development: accuracy 442 — 84.39% 83.94%

Paired uncertainty (Decision minus Tev; percentage points):

  • fresh: family macro +11.40 pp; paired group-bootstrap 95% interval [+7.40, +15.40] pp.
  • historical: family macro +11.19 pp; paired group-bootstrap 95% interval [+9.42, +13.01] pp.

Fresh-holdout source results

Source Cases Tev Decision Difference
clinc150 100 83.00% 89.00% +6.00 pp
go_emotions 100 69.00% 82.00% +13.00 pp
helpsteer2 100 56.00% 47.00% -9.00 pp
paws 100 57.00% 92.00% +35.00 pp
vitaminc 100 76.00% 88.00% +12.00 pp

All development source results (selection diagnostics)

Source Cases Tev Decision Difference
ag_news 50 94.00% 84.00% -10.00 pp
banking77 50 80.00% 92.00% +12.00 pp
boolq 50 86.00% 86.00% +0.00 pp
clinc150 50 74.00% 94.00% +20.00 pp
go_emotions 50 70.00% 88.00% +18.00 pp
helpsteer2 50 54.00% 54.00% +0.00 pp
mnli 50 82.00% 86.00% +4.00 pp
paws 50 66.00% 96.00% +30.00 pp
policy 48 100.00% 100.00% +0.00 pp
policy_v2 48 85.42% 87.50% +2.08 pp
research_taxonomy_v21 48 100.00% 100.00% +0.00 pp
routing_v2 48 70.83% 75.00% +4.17 pp
sst5 50 62.00% 46.00% -16.00 pp
vitaminc 50 68.00% 88.00% +20.00 pp
workflow.simple_quorum 100 56.00% 52.00% -4.00 pp
workflow.veto_alternative 100 46.00% 44.00% -2.00 pp

Training and artifacts

The v2 full run completed 9,289 optimizer steps and 29,185,870 tokens in 26.1 minutes (training/evaluation loop only), with 51.53 GiB peak allocated VRAM. Recipe: Qwen3.5-0.8B, BF16, rank-8 LoRA, alpha 16, learning rate 5e-5, one epoch, microbatch 8, accumulation 1, seed 42, 2048-token context. Batch-eight numerical parity passed the pre-existing thresholds before training. Selected identity: cb80a8e8a5ffd1232abaecd37a3f91a58c0f268d218069ce4cb779f70158aff8. Model checkpoints, tokenizer files, input datasets and evaluation reports have content hashes. See selection.json, final/freeze.json, final/comparison.json and the candidate preflight/summary. The earlier v1 run was intentionally interrupted after preserving a 500-step diagnostic checkpoint; it is not counted as a completed epoch.

Warm local latency

Warm local Python decision call: prompt rendering, tokenization, host/device transfer, forward, option scoring, CPU result; excludes model load and HTTP/server/Ollama overhead. CUDA workstation, not Mac.

Model E2E median E2E p95 Forward median
tev 14.23 ms 41.18 ms 13.02 ms
decision 20.48 ms 48.67 ms 19.31 ms

Export status

The adapter, F16 and Q8_0 GGUF builds are published. Q4_K_M failed the unchanged deployment accuracy gate and is not published. Earlier probability-equivalence failures remain documented in export verification. All benchmark and timing scores here belong to the adapter.

Scope and limitations

The primary 500-case holdout has 100 previously unused local groups from each of five public task families. It was constructed before candidate test inference and excludes previously used local IDs, groups, normalized states and prompts. This does not rule out pretraining or competitor-training overlap. The historical 2800-case suite was previously used for 4B reporting. Its scores are regression evidence, not a newly blind test. Development results, including the 442 original-core cases, were reused for selection. The comparison scores every declared option using a common prompt, disabled thinking and local BF16 inference. It evaluates downloaded checkpoints, not hosted-endpoint parity or unconstrained-generation formatting. These results do not establish universal superiority. Report every source regression. Bootstrap intervals use 2000 paired whole-group draws and exclude training-seed uncertainty. HelpSteer2 uses held-out original training prompts; VitaminC grouping is not article separation. No physical Mac/Ollama or iPhone latency was measured.