GLM-4.5-Air-0.72B-MTP

Small BF16 test model using zai-org/GLM-4.5-Air, with a toy-trained backbone and synthetic MTP weights.

Quantization recipe

Uses Transformers 5.17.0 and LLM Compressor PR #3225.

import torch
from compressed_tensors.offload import set_onload_device
from compressed_tensors.quantization import preset_name_to_scheme
from transformers import AutoTokenizer, Glm4MoeForCausalLM

from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.utils import load_context

MODEL_ID = "inference-optimization/GLM-4.5-Air-0.72B-MTP"
SAVE_DIR = "GLM-4.5-Air-0.72B-MTP-FP8-Dynamic"

with load_context(Glm4MoeForCausalLM, load_mtp=True):
    model = Glm4MoeForCausalLM.from_pretrained(
        MODEL_ID, dtype=torch.bfloat16, device_map="cpu",
    )
set_onload_device(model, "cuda")

recipe = QuantizationModifier(
    config_groups={
        "mtp": preset_name_to_scheme("FP8_DYNAMIC", targets=[r"re:^mtp\.layers\."]),
        "backbone": preset_name_to_scheme("FP8_DYNAMIC", targets=["Linear"]),
    },
    ignore=["lm_head", r"re:.*\.eh_proj$", r"re:.*\.indexer\..*"],
)

oneshot(model=model, recipe=recipe)
model.save_pretrained(SAVE_DIR)
AutoTokenizer.from_pretrained(MODEL_ID).save_pretrained(SAVE_DIR)
Downloads last month
6
Safetensors
Model size
0.7B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-optimization/GLM-4.5-Air-0.72B-MTP

Quantizations
1 model

Collections including inference-optimization/GLM-4.5-Air-0.72B-MTP