GLM-4.7-Flash-RAZOR-24B-A3B-E48of64

GLM-4.7-Flash with a quarter of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 48 of its original 64 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights.

Base This model
Routed experts per MoE layer 64 48
Total stored parameters, including MTP 31.221B 24.123B
Stored MTP parameters 1.278B 1.127B
Stored parameters excluding MTP 29.943B 22.996B
Active parameters per token ~3B ~3B
Active experts per token (top-k) 4 4 (unchanged)
Shared experts 1 1 (unchanged)
Decoder layers 47 47 (unchanged)
Context length 202,752 202,752 (unchanged)

Pruning touches only the routed expert pool. Attention, the shared expert, the embedding, the LM head and the router's remaining rows are untouched, so the compute per token is essentially unchanged while expert storage shrinks by a quarter.

A 50% variant is available as GLM-4.7-Flash-RAZOR-17B-A3B-E32of64.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, dtype="bfloat16", device_map="auto",
)

messages = [{"role": "user", "content": "Explain mixture-of-experts routing."}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0]))

Requires a Transformers build that contains the native glm4_moe_lite implementation. The checkpoint stores bfloat16 weights in seven safetensors shards.

How the experts were selected

RAZOR asks whether the surviving computation can replace an expert's function, rather than how often or how strongly the expert fires. For a token routed to the set $S$ with normalized weights $w_j$, let $c=\sum_{j\in S} w_j f_j$ be the routed mixture and $r_j=f_j-c$ the consensus residual of expert $j$. Deleting a selected expert $i$ makes the router promote its highest-ranked unselected expert $r$, whose pseudo-weight $w_r$ is its score divided by the original selected score sum. With $\lambda$ the routed-output scale, the exact local output change is

δi=λ∥c−c~−i∥2=λ ∥wiri−wrrr∥21−wi+wr. \delta_i=\lambda\lVert c-\tilde c^{-i}\rVert_2 =\lambda\,\frac{\lVert w_i r_i - w_r r_r\rVert_2}{1-w_i+w_r}.

Scores are aggregated by conditional root mean square over the calibration tokens routed to each expert, and the highest-scoring experts are retained per layer. See the paper for the derivation and the code for the implementation.

Reproducing this checkpoint

Calibration used 32,768-token rows from RazorCal, the 2,048-sample multi-domain corpus released with RAZOR.

pip install razor-moe

razor saliency --model <path-to-GLM-4.7-Flash> \
               --data data/RazorCal.json --max-len 32768 --out out/sal
razor prune    --model <path-to-GLM-4.7-Flash> --saliency out/sal \
               --method razor --ratio 0.25 --out out/pruned

kept_expert_indices.json in this repository is the keep-set manifest this checkpoint was built from, in the format razor verify expects:

razor verify --pruned . --src <path-to-GLM-4.7-Flash>

Selection depends on the calibration draw, so an independent run reproduces the procedure rather than this exact expert set.

Limitations

Expert pruning is lossy. Benchmark retention and predictive fidelity do not guarantee stable generation: the paper reports that responses change in diversity, formatting and termination behaviour even where task accuracy is largely preserved. Evaluate on your own workload before deploying.

The calibration corpus is multi-domain but finite, so behaviour on domains far from it is not characterised by the released measurements.

License and attribution

This is a derivative of GLM-4.7-Flash, released by Z.ai under the MIT license, and is distributed under the same license. The retained weights are the base model's own weights; RAZOR contributes the expert selection, not new parameters.

The RAZOR code is Apache-2.0. RazorCal records remain subject to their applicable upstream terms; see data/LICENSE-DATA.

Citation

@misc{song2026razorpruningreplaceableexperts,
      title={RAZOR: Pruning Replaceable Experts in LLMs},
      author={Mingyang Song and Mao Zheng},
      year={2026},
      eprint={2609.30465},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.30465},
}
Downloads last month
721
Safetensors
Model size
24B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64

Finetuned
(80)
this model

Collection including Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64

Paper for Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64