GLM 5.3 Flash / DFlash2 inference runtime 0.5.4
This is a Docker-save runtime archive for two NVIDIA Thor T5000 nodes with driver 595.78 and SM 11.0. It contains no model weights. Download the target and DFlash2 snapshots separately from the manateelazycat model repositories. The corresponding LPK uses the selected Hugging Face / ModelScope source for both model files and this archive. Public downloads require no hub login.
The frozen reference environment is TensorFold 0.6.1 plus Mia 1.3.2 with the validated Thor decoder and communication patches. Default capacity is 1,048,576 tokens per request, a 2,684,928-token shared KV pool, and four streams. Reference acceptance includes empty-cache CUDA JIT, vision, structured output, tools, exact serial/drafted/concurrent token equality, restart, three-run benchmarks, and full 1M three-position retrieval with format-only guided regex. Version 0.5.1 repeats the local two-rank draft timing calibration in three independent batches and selects one complete measured table. It does not embed timing constants from the reference machines. Version 0.5.4 bundles the Thor Q4 prompt tiling and a four-rail TCP transport for 1–16 MiB prompt exchanges. The native library and launch integration are inside the image; no experiment directory or source bind mount is required. Both ranks discover their existing optical IPv4 addresses and match four distinct subnets through the control store. Incomplete four-link topology retains NCCL; decode and smaller/larger exchanges retain the existing path. Transport faults are reported as errors, never silently replaced with stale results. The sparse-attention experiment is not part of the released profile. The regex constrains labels and hex syntax, never the random expected values. An earlier 0.5.0 unguided 1M run retrieved all three values but omitted the labels, so its strict output-format check failed and its receipt is retained. See runtime-manifest.json for exact identities and measured performance; capacity is not a speed claim.
Check SHA256SUMS before docker load. Model volumes are mounted read-only; compiler caches are separate. Framework/code notices are preserved inside the archive. Component licenses apply independently; this archive grants no additional rights to model weights. DFlash2 weights are CC BY-NC-ND 4.0 and must be used without modification in the authorized noncommercial deployment.