Instructions to use Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64") model = AutoModelForCausalLM.from_pretrained("Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64
- SGLang
How to use Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64 with Docker Model Runner:
docker model run hf.co/Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64
GLM-4.7-Flash-RAZOR-24B-A3B-E48of64
GLM-4.7-Flash with a quarter of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 48 of its original 64 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights.
| Base | This model | |
|---|---|---|
| Routed experts per MoE layer | 64 | 48 |
| Total stored parameters, including MTP | 31.221B | 24.123B |
| Stored MTP parameters | 1.278B | 1.127B |
| Stored parameters excluding MTP | 29.943B | 22.996B |
| Active parameters per token | ~3B | ~3B |
| Active experts per token (top-k) | 4 | 4 (unchanged) |
| Shared experts | 1 | 1 (unchanged) |
| Decoder layers | 47 | 47 (unchanged) |
| Context length | 202,752 | 202,752 (unchanged) |
Pruning touches only the routed expert pool. Attention, the shared expert, the embedding, the LM head and the router's remaining rows are untouched, so the compute per token is essentially unchanged while expert storage shrinks by a quarter.
A 50% variant is available as GLM-4.7-Flash-RAZOR-17B-A3B-E32of64.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, dtype="bfloat16", device_map="auto",
)
messages = [{"role": "user", "content": "Explain mixture-of-experts routing."}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0]))
Requires a Transformers build that contains the native glm4_moe_lite
implementation. The checkpoint stores bfloat16 weights in seven safetensors
shards.
How the experts were selected
RAZOR asks whether the surviving computation can replace an expert's function, rather than how often or how strongly the expert fires. For a token routed to the set $S$ with normalized weights $w_j$, let $c=\sum_{j\in S} w_j f_j$ be the routed mixture and $r_j=f_j-c$ the consensus residual of expert $j$. Deleting a selected expert $i$ makes the router promote its highest-ranked unselected expert $r$, whose pseudo-weight $w_r$ is its score divided by the original selected score sum. With $\lambda$ the routed-output scale, the exact local output change is
Scores are aggregated by conditional root mean square over the calibration tokens routed to each expert, and the highest-scoring experts are retained per layer. See the paper for the derivation and the code for the implementation.
Reproducing this checkpoint
Calibration used 32,768-token rows from RazorCal, the 2,048-sample multi-domain corpus released with RAZOR.
pip install razor-moe
razor saliency --model <path-to-GLM-4.7-Flash> \
--data data/RazorCal.json --max-len 32768 --out out/sal
razor prune --model <path-to-GLM-4.7-Flash> --saliency out/sal \
--method razor --ratio 0.25 --out out/pruned
kept_expert_indices.json in this repository is the keep-set manifest this
checkpoint was built from, in the format razor verify expects:
razor verify --pruned . --src <path-to-GLM-4.7-Flash>
Selection depends on the calibration draw, so an independent run reproduces the procedure rather than this exact expert set.
Limitations
Expert pruning is lossy. Benchmark retention and predictive fidelity do not guarantee stable generation: the paper reports that responses change in diversity, formatting and termination behaviour even where task accuracy is largely preserved. Evaluate on your own workload before deploying.
The calibration corpus is multi-domain but finite, so behaviour on domains far from it is not characterised by the released measurements.
License and attribution
This is a derivative of GLM-4.7-Flash, released by Z.ai under the MIT license, and is distributed under the same license. The retained weights are the base model's own weights; RAZOR contributes the expert selection, not new parameters.
The RAZOR code is Apache-2.0. RazorCal records remain subject to their applicable upstream terms; see data/LICENSE-DATA.
Citation
@misc{song2026razorpruningreplaceableexperts,
title={RAZOR: Pruning Replaceable Experts in LLMs},
author={Mingyang Song and Mao Zheng},
year={2026},
eprint={2609.30465},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.30465},
}
- Downloads last month
- 721
Model tree for Nickyang/GLM-4.7-Flash-RAZOR-24B-A3B-E48of64
Base model
zai-org/GLM-4.7-Flash