Publish independent quantization fidelity measurement
Browse files
README.md
CHANGED
|
@@ -69,6 +69,18 @@ Use K=8 for list/JSON-heavy workloads (`ONE_SPARK_K=8`). Raw data in the GitHub
|
|
| 69 |
|
| 70 |
Every cold request had zero prefix-cache hits; every warm checksum response was correct.
|
| 71 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
## Links
|
| 73 |
|
| 74 |
- **GitHub source and instructions:** `https://github.com/gitcommit90/glm-5.3-one-spark`
|
|
|
|
| 69 |
|
| 70 |
Every cold request had zero prefix-cache hits; every warm checksum response was correct.
|
| 71 |
|
| 72 |
+
## Quantization Analysis
|
| 73 |
+
|
| 74 |
+
This deployment uses [Turboderp's GLM-5.3-Flash EXL3 2.05-bpw quant](https://huggingface.co/turboderp/GLM-5.3-Flash-exl3/tree/2.05bpw), revision `51058cd551c7e570d87bd32a4adee720edce2349`. The exact checkpoint is **85.23 GB (79.38 GiB)**.
|
| 75 |
+
|
| 76 |
+
An [independent full-vocabulary measurement](https://github.com/malaiwah/quant-fidelity-suite/blob/794d80fa79db4d30cd0fa8140a07c665dd363251/registry/receipts/malaiwah/stream-turbo-2.05bpw-kld.json) of this exact revision against BF16 teacher logits reported:
|
| 77 |
+
|
| 78 |
+
| Quant | Size | Top-1 agreement | Mean KLD | Scored positions |
|
| 79 |
+
|---|---:|---:|---:|---:|
|
| 80 |
+
| Turboderp EXL3 2.05 | 85.23 GB | **88.92%** | **0.121638** | **51,175** |
|
| 81 |
+
|
| 82 |
+
The measurement used 25 windows, the full 154,880-token vocabulary, teacher forcing, FP64 accumulation, and two cold runs with identical results. The checkpoint was quantized from the official FP8 release; the measurement reference is BF16. Machine-readable summary: [`quantization-analysis.json`](https://github.com/gitcommit90/glm-5.3-one-spark/blob/main/benchmarks/quality/quantization-analysis.json).
|
| 83 |
+
|
| 84 |
## Links
|
| 85 |
|
| 86 |
- **GitHub source and instructions:** `https://github.com/gitcommit90/glm-5.3-one-spark`
|