# Export verification All deployment checks use the same fixed 892-case development panel, maximum context 2,048 tokens, raw option probabilities, and complete finite logits. Prompt token IDs, option boundaries and EOS IDs are checked for every GGUF case. This is a workstation fidelity check, not a new blind test or a phone measurement. The established release gates allow at most 0.5 percentage points of accuracy loss for the merged BF16 model versus the adapter, and at most 1 point for GGUF versus merged BF16. The same thresholds were retained for every attempt. | Artifact | Dev accuracy | |---|---:| | BF16 adapter | 75.90% | | Merged BF16 | 76.23% | | F16 | 75.45% | | Q8_0 | 75.78% | | Q4_K_M (failed; not published) | 72.65% | Merged BF16, F16 and Q8_0 passed. Q4_K_M failed with a 3.59-point drop versus merged BF16 and is not published. Q8_0 was tested as a higher-precision fallback after that failure; format selection used only development results. Two earlier 32-case merge diagnostics failed a separate, stricter maximum absolute probability-difference tolerance of 0.02. Direct BF16 merge: 0.0200787; FP32 merge arithmetic followed by BF16 output: 0.0399304. Neither changed the sampled argmax predictions. Those failures remain failures. The established full-panel release procedure checks classification accuracy, not identical probabilities; passing it does not resolve the failed probability-equivalence diagnostics. Probabilities remain uncalibrated, and calibration does not transfer automatically between formats. Historical and fresh test scores on the model card belong to the adapter. F16 and Q8_0 have not been evaluated on those test panels. Physical Mac, iPhone and Ollama latency have not been measured.