Third-party fidelity measurement: KL(reference ‖ NVFP4) = 0.0548 nats
What this is. A third-party fidelity measurement of nvidia/GLM-5.2-NVFP4 @ 53e0691e21895a3863a606dfd12910c69eba94ab (nvfp4, 4 bits per weight as declared) against its unquantized reference, made with quant-fidelity-suite. Measured by malaiwah, not by the model's author. Every number below is read from a sealed receipt named at the bottom.
| KL(reference ‖ candidate), mean tokenwise, nats | 0.05483693836564808 |
| top-1 agreement | 0.9343820224719102 |
| KL median / p95 / p99 / max | 0.003298336777704599 / 0.2462004655492725 / 0.8396656210393004 / 9.295262576779585 |
| reference root dataset | malaiwah/glm52-fidelity-root-v1 @ 5977559307ee9fb7d6478e81a875faa10ffee9b8 (capture a544e029a0392c2a…) |
| panel | panel--glm53.malaiwah.corpus5x5-v1, 25 contexts, 51175 scored positions |
| direction / vocabulary / accumulation | KL(reference |
| method | dequantize-and-run, weights only: nvfp4-modelopt-dequant-to-bf16 -- the ModelOpt NVFP4 dialect (quant_algo NVFP4, producer modelopt 0.46.0.dev65+g977d34dc3): each routed-expert projection's packed e2m1 nibbles (group 16 along the input axis) are decoded to exact fp32 as e2m1 x weight_scale.f32 x weight_scale_2 on the capture device and cast once to bf16 under the official tensor name, bitwise the compressed-tensors reference on real fetched rows; routed experts only -- every non-routed tensor is carried as shipped (plain bf16 under the official names); the per-tensor input_scale is an activation quantity and is NOT applied (static-nvfp4-not-applied), so the measurement is weights-only; same engine, schedule and device as the reference capture |
| determinism | two fresh processes captured the candidate; both sealed captures carry content digest 574db348a8e2df11… (self-comparison 0.0) |
| comparability class | advisory |
Scope (scope_digest): attn.o=native:bf16@16|attn.other=native:bf16@16|attn.qkv=native:bf16@16|embed_tokens=native:bf16@16|lm_head=native:bf16@16|mlp.down=native:bf16@16|mlp.gate=native:bf16@16|mlp.up=native:bf16@16|moe.experts=quantized:nvfp4@4|moe.router=native:mixed|moe.shared_expert=native:bf16@16|mtp=native:mixed|norm=native:bf16@16|head=native|kv=bf16
Disclosures on the comparison receipt:
native_head_replay(info): HEAD-1d: each side replayed through its own sealed head (reference a012be05e771, candidate a012be05e771); head error is inside the measurement, as under HEAD-2, and nothing is substituted. The heads are content-identical.activation_quantization_not_captured(caveat): candidate was captured from a bf16 materialisation of its weights (nvfp4-modelopt-dequant-to-bf16); the checkpoint declares activation quantization (static-nvfp4-not-applied) that a weights-only capture does not apply, so a served deployment also quantizes activations at runtime. That term is not in this number, which is expected to understate the served divergence; it is not a mathematical bound. The comparison is advisory.
What this number does not do. It is a same-lane distance from one reference capture on one panel. It does not rank this artifact against numbers measured on another panel, lane or reference, and a same-lane root does not retroactively upgrade rows measured against another teacher. Per-window scatter exceeds the gap between adjacent bit-widths; compare only within a group whose comparability keys match.
Receipts.
- comparison receipt
receipt_sha25691787525e2defcda0d2263393d7edc8f338e9984dda160e78ff91fc8e9a5a2fd - root-qualification receipt
receipt_sha256c35eb145b02430f6dacf96b7b6fe10d8aab9bbef1a307e0f9c9f7ce702f5ea40(canonical dataset_sha2560668268f240438f80b39d80da4f65d4ac2f6f13a6a45bd66cb80de6b46e5b922, repeateb8656635d8f9b26f061e71a62fc002412007c92b33d0f2afa0f912cefbb59ea) - job
job_id_fullebb4e6f72bca211e05dc833021da93cba80eaa190cd20320129ac42f3922b9a8
Reproduce: fetch the two datasets named above and run fidelity-dataset compare --reference <root> --candidate <this> --own-heads; the receipt's estimator block is the exact recipe. Questions and corrections are welcome here; the registry files this as a third-party row (measured_by enumerated, never conflated with author-reported numbers).