decision-0.8b / evaluation /EXPORT_VERIFICATION.md
johnsonchromia's picture
Publish Decision-0.8B and verified benchmarks
cb17a40 verified
|
Raw History Blame Contribute Delete
1.74 kB
# Export verification
All deployment checks use the same fixed 892-case development panel, maximum context 2,048 tokens, raw option probabilities, and complete finite logits. Prompt token IDs, option boundaries and EOS IDs are checked for every GGUF case. This is a workstation fidelity check, not a new blind test or a phone measurement.
The established release gates allow at most 0.5 percentage points of accuracy loss for the merged BF16 model versus the adapter, and at most 1 point for GGUF versus merged BF16. The same thresholds were retained for every attempt.
| Artifact | Dev accuracy |
|---|---:|
| BF16 adapter | 75.90% |
| Merged BF16 | 76.23% |
| F16 | 75.45% |
| Q8_0 | 75.78% |
| Q4_K_M (failed; not published) | 72.65% |
Merged BF16, F16 and Q8_0 passed. Q4_K_M failed with a 3.59-point drop versus merged BF16 and is not published. Q8_0 was tested as a higher-precision fallback after that failure; format selection used only development results.
Two earlier 32-case merge diagnostics failed a separate, stricter maximum absolute probability-difference tolerance of 0.02. Direct BF16 merge: 0.0200787; FP32 merge arithmetic followed by BF16 output: 0.0399304. Neither changed the sampled argmax predictions. Those failures remain failures. The established full-panel release procedure checks classification accuracy, not identical probabilities; passing it does not resolve the failed probability-equivalence diagnostics.
Probabilities remain uncalibrated, and calibration does not transfer automatically between formats. Historical and fresh test scores on the model card belong to the adapter. F16 and Q8_0 have not been evaluated on those test panels. Physical Mac, iPhone and Ollama latency have not been measured.