GLM-5.3-NVFP4-CSF

A lossless NVFP4-CSF container of local-inference-lab/GLM-5.3-NVFP4 at revision b472e4ee53f6a9862da5486c56c6ca21be3dab70, the GLM-5.3 NVFP4 checkpoint (744B MoE; NVFP4 routed experts, BF16 dense layers).

CSF stores the routed-expert block-scale planes in a compressed form. Nothing else changes: FP4 weight nibbles, every other tensor, the tensor names, dtypes, shapes and the source metadata are all byte-identical. Decoding is integer arithmetic on bytes, with no requantization or fitting. The original safetensors shards can be restored bit-exactly.

Routed experts (layers 3-77 x 256) - NVFP4 (E2M1, E4M3 scale per 16), ModelOpt PTQ (max calibration)
MTP (layer 78) routed experts - BF16
Attention (MLA, DSA indexer), shared experts, dense layers 0-2 - BF16

Routed-expert block scales - lossless NVFP4-CSF (byte-window4-fixed-stream-u24-exceptions/1)

Sizes

  • Weight files: 464.82 GB in the source, 442.83 GB here (21.99 GB saved).
  • Compressed scales: 57,600 matrices (E4M3 block scales of the 75 x 256 x 3 main-layer routed-expert projections): 45.30 GB -> 23.30 GB (51.4%).

Provenance

  • Source: local-inference-lab/GLM-5.3-NVFP4 revision b472e4ee53f6a9862da5486c56c6ca21be3dab70, 85 index-referenced safetensors shards. During export, each source shard's SHA-256 was checked against its Hugging Face LFS hash.
  • Built with trellis-quant trellis_quant.lossless_scale_checkpoint (commit 6ecc90456275, family glm53_744b_nvfp4).
  • verification.json: every original shard was rebuilt from this container, and both the rebuilt and the stored files matched their SHA-256 (57,600 scale matrices, 85 shards, passed).
  • Files that the source index does not reference (amax_checkpoint.safetensors, amax_new.safetensors) and the quantization log finalize.log are not carried. vLLM never loads them.

Layout

lil-nvfp4-csf-checkpoint/1, codec byte-window4-fixed-stream-u24-exceptions/1:

  • tensors/ - the source shard names; each routed-expert scale <name> is stored as <name>.nvfp4_csf_fixed (uint8) plus <name>.nvfp4_csf_exceptions (uint32)
  • metadata/ - byte copies of the source's config, tokenizer, index, README and LICENSE
  • config.json - a copy of metadata/config.json at the top of the repository, where the Hub counts downloads; runtimes read metadata/
  • manifest.json, build-contract.json, receipts/ (per-shard source headers and hashes), verification.json
  • LICENSE, NOTICE, LICENSES/ (upstream license texts), CITATION.cff, REUSE.toml, and the SHA-256 of every file in SHA256SUMS and lil-manifest.json

There is no top-level model.safetensors.index.json, so a plain safetensors loader will not open this directory by mistake.

Serving

Use vLLM with the NVFP4-CSF reader: --quantization nvfp4_csf --load-format nvfp4_csf. Point vLLM at a serving directory that holds the files of metadata/, with config.json's quantization_config replaced by:

{
  "quant_method": "nvfp4_csf",
  "format_version": 1,
  "checkpoint_root": "/path/to/this/checkpoint",
  "source_quantization_config": { "...": "metadata/config.json quantization_config" }
}

The weights are read from checkpoint_root; the serving directory holds only metadata.

Serving needs a runtime whose NVFP4-CSF reader knows the glm53_744b_nvfp4 family.

Verify or restore

PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint verify \
  --checkpoint /path/to/this/checkpoint --workers 16
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint restore \
  --checkpoint /path/to/this/checkpoint --destination /path/to/fresh/dir --workers 16

verify rebuilds every original shard and compares SHA-256 values. restore writes the original shards and metadata files into a fresh directory, checking every hash before a file is published. The result is an ordinary copy of the source checkpoint.

Upstream license

GLM-5.3 (zai-org/GLM-5.3-BF16) is licensed by Z.AI under the GLM-5.3 License, and its conditions, including the Model as a Service security review, also apply to this model. Its text is in LICENSES/LicenseRef-glm-5.3.txt, metadata/LICENSE keeps the source checkpoint's copy unchanged, and NOTICE names the upstream materials. This container is licensed under the Local Inference Lab License, Version 1.0 (LICENSE); see License and attribution below.

License and attribution

GLM-5.3-NVFP4-CSF is licensed under the Local Inference Lab License, Version 1.0, which reproduces the terms and conditions of the Apache License, Version 2.0, and adds conditions that restrict them. It is not the Apache License, and these files are not open source. Copyright (c) 2026 Local Inference Lab, Inc., a non-profit organization, 311 West Main Street, Grayson, Kentucky 41143.

  • No reuploads. Do not upload, mirror, or redistribute these files, or any Substantially Similar Copy of them (including renamed, re-sharded, re-packaged, metadata-stripped, converted, or dequantized copies). Link to https://huggingface.co/local-inference-lab/GLM-5.3-NVFP4-CSF instead.
  • Attribution at the top. Every README, model card, or other landing page for a project, model, dataset, application, service, or distribution that contains, is derived from, or runs GLM-5.3-NVFP4-CSF must begin with this Attribution Notice as its first paragraph (a single title line may come before it):

This model is based on GLM-5.3-NVFP4-CSF by Local Inference Lab, Inc., a non-profit organization, available at https://huggingface.co/local-inference-lab/GLM-5.3-NVFP4-CSF. GLM-5.3-NVFP4-CSF is licensed under the Local Inference Lab License, Version 1.0.

Plain-text form:

This model is based on GLM-5.3-NVFP4-CSF by Local Inference Lab, Inc., a non-profit organization, available at https://huggingface.co/local-inference-lab/GLM-5.3-NVFP4-CSF. GLM-5.3-NVFP4-CSF is licensed under the Local Inference Lab License, Version 1.0.
  • Keep the marks. Do not remove the license metadata, the Identifying Marks ("Ні пуху, ні пера", LIL-CANARY-BAE5-5E6E-9CF9-11D5), or the copyright notices embedded in these files.
  • Breach. A breach of the no-reupload or attribution terms ends the license immediately (Section 5.1), and Local Inference Lab, Inc. may ask hosting services to remove the material (Section 5.2).

File integrity: SHA256SUMS and lil-manifest.json list the SHA-256 of every file.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for local-inference-lab/GLM-5.3-NVFP4-CSF

Base model

zai-org/GLM-5.3
Quantized
(1)
this model