decision-0.8b / evaluation /EXPORT_VERIFICATION.md
johnsonchromia's picture
Publish Decision-0.8B and verified benchmarks
cb17a40 verified
|
Raw History Blame Contribute Delete
1.74 kB

Export verification

All deployment checks use the same fixed 892-case development panel, maximum context 2,048 tokens, raw option probabilities, and complete finite logits. Prompt token IDs, option boundaries and EOS IDs are checked for every GGUF case. This is a workstation fidelity check, not a new blind test or a phone measurement.

The established release gates allow at most 0.5 percentage points of accuracy loss for the merged BF16 model versus the adapter, and at most 1 point for GGUF versus merged BF16. The same thresholds were retained for every attempt.

Artifact Dev accuracy
BF16 adapter 75.90%
Merged BF16 76.23%
F16 75.45%
Q8_0 75.78%
Q4_K_M (failed; not published) 72.65%

Merged BF16, F16 and Q8_0 passed. Q4_K_M failed with a 3.59-point drop versus merged BF16 and is not published. Q8_0 was tested as a higher-precision fallback after that failure; format selection used only development results.

Two earlier 32-case merge diagnostics failed a separate, stricter maximum absolute probability-difference tolerance of 0.02. Direct BF16 merge: 0.0200787; FP32 merge arithmetic followed by BF16 output: 0.0399304. Neither changed the sampled argmax predictions. Those failures remain failures. The established full-panel release procedure checks classification accuracy, not identical probabilities; passing it does not resolve the failed probability-equivalence diagnostics.

Probabilities remain uncalibrated, and calibration does not transfer automatically between formats. Historical and fresh test scores on the model card belong to the adapter. F16 and Q8_0 have not been evaluated on those test panels. Physical Mac, iPhone and Ollama latency have not been measured.