Update model card: NFS-served TP4 lane (validated 2026-09-09), supersedes all-local-NVMe lane
be70d30 |
Download README.md from vikasclawd/GLM-5.3-Int4-Int8Mix: direct link, hf CLI and curl.
- Browser
- Download file 7.24 kB
-
https://huggingface.co/vikasclawd/GLM-5.3-Int4-Int8Mix/resolve/main/README.md
- Command line
-
hf download hf://vikasclawd/GLM-5.3-Int4-Int8Mix/README.md
-
curl -L -o README.md https://huggingface.co/vikasclawd/GLM-5.3-Int4-Int8Mix/resolve/main/README.md
7.24 kB
| license: other | |
| base_model: zai-org/GLM-5.3-BF16 | |
| tags: [glm-5.3, quantization, int4, int8, mixed-precision, compressed-tensors, dgx-spark, quanttrio, vllm, mtp, speculative-decoding, nfs, rdma] | |
| pipeline_tag: text-generation | |
| # GLM-5.3-Int4-Int8Mix | |
| The first Int4/Int8Mix quantization of big **GLM-5.3 (743B)** — produced with the | |
| QuantTrio recipe, verified with fail-closed structural + numerical gates, and served at | |
| TP4 on four NVIDIA DGX Spark (GB10 / sm121) nodes over a RoCE fabric. | |
| - **Base:** `zai-org/GLM-5.3-BF16` (1507 GB, genuine BF16 — chosen over the fp8 repo so the router, DSA indexer, lm_head and layer 0 stay true BF16) | |
| - **Size:** 377.4 GiB, 282 shards, `compressed-tensors` / `pack-quantized` | |
| - **Per-rank at TP4:** 94.6 GiB | |
| - **Companion repo (scripts + raw results):** https://github.com/chishiki37/glm-5.3-int4-int8mix-4x-dgx-spark | |
| ## Serving (TP4 on 4× DGX Spark, NFS-served weights) | |
| **Status:** validated 2026-09-09 · **Nodes:** 4 (gx10-141d head + 3 NFSoRDMA workers) · | |
| **Context:** 200K (1M in config, KV-pinned pool 200,064 tok) | |
| Weights live in **one copy** on the head node (gx10-141d, 200G IB) and are | |
| NFSoRDMA-exported to the three workers — no per-node NVMe copies. This topology | |
| **supersedes the all-local-NVMe lane** (27.2/59.1/86.9 on 1d49+3b24+cb98+04af, | |
| 2026-08-30): it matches-or-beats it while keeping 3× ~405 GB of NVMe copies off the | |
| workers. Head reads weights locally (the NFS server never mounts its own export); | |
| workers mount the same path over RDMA (proto=rdma, vers=3, port 20049), so one uniform | |
| launcher runs on all ranks. Measured NFS: single-stream O_DIRECT 507 MB/s, 4-stream | |
| aggregate 3.6 GB/s; weight-load phase ≈ 7 min for 3 concurrent NFS ranks + 1 local rank. | |
| Sweep-optimized winner config (autoresearch over spec k, max-seqs, NCCL channels, | |
| Marlin atomic-add, cudagraph mode, rail count — full table in the companion repo): | |
| ``` | |
| --speculative-config '{"method":"mtp","num_speculative_tokens":3,"draft_tensor_parallel_size":1,"attention_backend":"FLASHMLA_SPARSE"}' | |
| --tensor-parallel-size 4 --nnodes 4 --master-port 29552 --max-model-len 200000 | |
| --max-num-seqs 16 --max-num-batched-tokens 8192 | |
| --gpu-memory-utilization 0.91 --kv-cache-memory-bytes 10950000000 --kv-cache-dtype fp8_ds_mla | |
| --compilation-config '{"cudagraph_mode":"FULL"}' --enable-prefix-caching --async-scheduling | |
| --reasoning-parser glm45 --tool-call-parser glm47 --enable-auto-tool-choice --trust-remote-code | |
| ``` | |
| The three winning knobs vs the local-weights lane: **dual-rail NCCL** | |
| (`NCCL_IB_HCA=rocep1s0f0,roceP2p1s0f0`, GID index omitted), **Marlin atomic-add** | |
| (`VLLM_MARLIN_USE_ATOMIC_ADD=1`), and **max-num-seqs 16** (was 6/8). Doctrine unchanged: | |
| NCCL_MAX/MIN_NCHANNELS=4, HF_HUB_OFFLINE=1, pinned KV (never unpinned), swappiness=10, | |
| threshold-flushers only (drop iff MemAvailable<8 GiB). Boot ≈ 17–20 min. | |
| ### Measured (e2e wall-clock, 256-token prose, temp 0, warm engine) | |
| | concurrency | tok/s (2-run mean) | best run | | |
| |---|---|---| | |
| | C1 | **26.3** | 27.6 | | |
| | C4 | **56.1** | 56.7 | | |
| | C8 | **89.2** | 92.9 | | |
| Winner = dual-rail + Marlin atomic-add + seqs16 (Σ 171.6). vs the all-local lane: | |
| C1 −3%/+1.6% (run-dependent), C4 −5%, C8 +3–7% — NFS parity at C1/C8 with 1.2 TB of | |
| worker NVMe freed. TTFT ≈ 0.6 s at ~780-token prompt · KV pool 200,064 tokens · MTP mean | |
| acceptance length 2.3–3.1 of 3. | |
| *Variance note: single-shot batteries swing ±10% on C1/C8 at identical config | |
| (24.94–27.64 C1). Marlin's C4 edge was the only consistently reproducible gain (3/3 runs | |
| ≥55.6). Run ≥2 batteries per arm before believing C1/C8 deltas under ~10%.* | |
| ### Integrity probes (NFS-lane health) | |
| - **Corruption probe:** PASS — ~6K-char generations × 3 passes each in ko/en/tr, 0 bad | |
| tokens across ~54K chars. | |
| - **Long-context needle:** PASS — 34,996-token prompt, all 3 keys (begin/middle/end) | |
| retrieved verbatim; prefill ≈550 tok/s, 63.7 s e2e. | |
| ### Benchmarks (gx10:8000, single stream) | |
| - **GSM8K 98%** (49/50, canonical set — 1 arithmetic slip) | |
| - **HumanEval-chat 92% pass@1** (46/50) | |
| ### Quality (10 samples/task, temp 0, served via vLLM chat endpoint) | |
| | task | local i4mix | hosted GLM-5.2 | | |
| |---|---|---| | |
| | GSM8K | 90% | 80% | | |
| | HumanEval | 100% | 100% | | |
| | MBPP | 80% | 80% | | |
| | IFEval | 70% | 90% | | |
| | MMLU-STEM | 100% | 100% | | |
| | ARC-C | 100% | 100% | | |
| | HellaSwag | 100% | 100% | | |
| Ties-or-better on 6/7 vs the hosted GLM-5.2 reference (a cross-generation point, not the | |
| same model); the IFEval deficit is 2 questions at n=10 — noise-level, not a quant signal. | |
| Decode during the battery: 22–29 tok/s single-stream mixed content. | |
| ## Quantization recipe | |
| Config groups transferred verbatim from `QuantTrio/GLM-5.2-Int4-Int8Mix` — GLM-5.2 and | |
| GLM-5.3 are structurally identical (78 layers, hidden 6144, kv_lora_rank 512, 256 routed | |
| experts, first_k_dense_replace 3). Data-free RTN: no calibration, static, symmetric, | |
| weight-only. | |
| | group | targets | bits | strategy | modules | | |
| |---|---|---|---|---| | |
| | `w4a16_experts` | layers 3–77 `mlp.experts.N.{gate,up,down}_proj` | 4 | group/128 | 57,600 | | |
| | `w8a16_linears` | layers 1–77 attention projs + dense & shared-expert MLP | 8 | group/128 | 616 | | |
| | `w8a16_mtp_channel` | layer 78 (MTP) attention + MLP + experts | 8 | channel (gs −1) | 776 | | |
| **Kept full precision:** all of layer 0, every `mlp.gate` (MoE router), the DSA | |
| `indexer`/`indexers_proj`, MTP `eh_proj`/`enorm`/`hnorm`/`shared_head`, and `lm_head`. | |
| Quantized by shard streaming (read one BF16 shard → quantize → pack → write one output | |
| shard; ~10 GiB peak RAM, resumable), using compressed-tensors' own `calculate_qparams` / | |
| `quantize` / `pack_to_int32`, with scales rounded to bf16 before quantizing so the | |
| persisted scale is exactly the one quantization assumed. Runtime: ~30–57 min CPU-only on | |
| one DGX Spark depending on source medium. | |
| ## Verification (fail-closed gates) | |
| - **Structural:** 282/282 shards open clean; exact 1:1 key mapping (59,585 source tensors | |
| → 58,992 quantized modules + 593 BF16/FP32 keeps); census 57,600/616/776 matches the | |
| pre-quant dry-run exactly; sacred modules plain BF16, zero packed leaks. | |
| - **Numerical:** dequant vs BF16 on sampled tensors across all three groups — | |
| W4-g128 nrmse ≈ 0.111 (exactly theoretical for data-free absmax RTN), W8-g128 ≈ 0.0072, | |
| W8-channelwise MTP ≈ 0.009. Packing convention proven by discrimination | |
| (wrong unpack scores cos ≈ −0.58). | |
| - **Integrity:** full 282-entry sha256 manifest (`.sha256` per shard in this repo); | |
| distributed to 4 cluster nodes with per-target verification, 282/282 OK each. | |
| ## Provenance | |
| Quantized, verified, and benchmarked on a fleet of DGX Spark nodes over a RoCE fabric by | |
| an agent-assisted pipeline (Aug 29–30 2026; NFS serving lane validated Sep 9 2026). | |
| Companion repo with the verification gate script, launch recipe, autoresearch sweep | |
| orchestrator, quality battery, and all raw JSONs: | |
| https://github.com/chishiki37/glm-5.3-int4-int8mix-4x-dgx-spark | |
| *Note: the one-off streaming quantizer driver was ad-hoc and is not preserved; the | |
| verifier encodes the full recipe invariants (group census, sacred-module allowlist, | |
| packing convention) needed to reproduce or validate a re-quantization.* | |