AI2Apps SSD-ready checkpoint

This is a byte-preserving storage-layout conversion of Vontra/GLM-5.3-Flash-MLX-4bit-MTP at immutable revision 06d6c7530e8290e20fabdc37a825ce07bdfc490c. Original tensor values and quantization are retained. Original routed experts are externalized into the experts/ directory rather than duplicated in the backbone safetensors.

Requires an AI2Apps Runtime with explicit glm5-next-affine-q4-gate-up-fused-v2 support. This candidate is not a drop-in checkpoint for unmodified Transformers, mlx-lm or mlx-vlm. Do not use the backbone safetensors alone. The corresponding Runtime and model Package have not yet completed release acceptance. Full/Cached engine compatibility is recorded separately in the Runtime release receipt; the presence of a reversible tensor map alone is not an end-to-end engine guarantee.

  • ssd-checkpoint.json: format, provenance and file digests.
  • external-tensors.json: original tensor names and external byte locations.
  • source-tensor-sha256.json: original tensor payload digests verified during export.
  • model.safetensors.index.json: ordinary/vision/other retained tensor index.
  • experts/: complete routed expert payloads, with no re-quantization.

See LICENSE and README.upstream.md for upstream terms and attribution. This storage format changes installation space and data access; it is not a new model training or a claim of improved model accuracy.

Downloads last month
75
Safetensors
Model size
17B params
Tensor type
U32
路
F32
路
BF16
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support