GLM-5-0.88B-MTP

Small BF16 test model using zai-org/GLM-5, with a toy-trained backbone and synthetic MTP weights.

Quantization recipe

Uses LLM Compressor model-free quantization. MTP inference has not been verified for this fixture.

from llmcompressor import model_free_ptq

MODEL_ID = "inference-optimization/GLM-5-0.88B-MTP"
SAVE_DIR = "GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQ"

model_free_ptq(
    model_stub=MODEL_ID,
    save_directory=SAVE_DIR,
    scheme="FP8_DYNAMIC",
    ignore=[
        "lm_head",
        "model.embed_tokens",
        r"re:.*\.mlp\.gate$",
        r"re:.*\.eh_proj$",
        r"re:.*\.indexer\..*",
    ],
    max_workers=2,
    device="cuda",
)
Downloads last month
5
Safetensors
Model size
0.9B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-optimization/GLM-5-0.88B-MTP

Quantizations
1 model

Collections including inference-optimization/GLM-5-0.88B-MTP