Hy3-RAZOR-154B-A21B-E96of192

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights.

Base This model
Routed experts per MoE layer 192 96
Total parameters 295B 154B
Active parameters per token ~21B ~21B
Active experts per token (top-k) 8 8 (unchanged)
MoE decoder layers 80 80 (unchanged)

Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed.

The other budget is Hy3-RAZOR-226B-A21B-E144of192.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Nickyang/Hy3-RAZOR-154B-A21B-E96of192"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, dtype="bfloat16", device_map="auto",
)

messages = [{"role": "user", "content": "Explain mixture-of-experts routing."}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0]))

Requires a Transformers build containing the native hy_v3 implementation. Weights are bfloat16.

How the experts were selected

RAZOR asks whether the surviving computation can replace an expert's function, rather than how often or how strongly the expert fires. For a token routed to the set $S$ with normalized weights $w_j$, let $c=\sum_{j\in S} w_j f_j$ be the routed mixture and $r_j=f_j-c$ the consensus residual of expert $j$. Deleting a selected expert $i$ makes the router promote its highest-ranked unselected expert $r$, whose pseudo-weight $w_r$ is its score divided by the original selected score sum. With $\lambda$ the routed-output scale, the exact local output change is

δi=λ∥c−c~−i∥2=λ ∥wiri−wrrr∥21−wi+wr. \delta_i=\lambda\lVert c-\tilde c^{-i}\rVert_2 =\lambda\,\frac{\lVert w_i r_i - w_r r_r\rVert_2}{1-w_i+w_r}.

Scores are aggregated by conditional root mean square over the calibration tokens routed to each expert, and the highest-scoring experts are retained per layer. See the paper for the derivation and the code for the implementation.

Reproducing this checkpoint

Calibration used 32,768-token rows from RazorCal, the 2,048-sample multi-domain corpus released with RAZOR.

pip install razor-moe

razor saliency --model <path-to-Hy3> \
               --data data/RazorCal.json --max-len 32768 --out out/sal
razor prune    --model <path-to-Hy3> --saliency out/sal \
               --method razor --ratio 0.5 --out out/pruned

kept_expert_indices.json is the keep-set manifest this checkpoint was built from, in the format razor verify expects:

razor verify --pruned . --src <path-to-Hy3>

Selection depends on the calibration draw, so an independent run reproduces the procedure rather than this exact expert set.

Limitations

Expert pruning is lossy. Benchmark retention and predictive fidelity do not guarantee stable generation: the paper reports that responses change in diversity, formatting and termination behaviour even where task accuracy is largely preserved. Evaluate on your own workload before deploying.

The calibration corpus is multi-domain but finite, so behaviour on domains far from it is not characterised by the released measurements.

License and attribution

This is a derivative of Hy3, released under the Apache-2.0 license, and is distributed under the same license. The retained weights are the base model's own weights; RAZOR contributes the expert selection, not new parameters.

The RAZOR code is Apache-2.0. RazorCal records remain subject to their applicable upstream terms; see data/LICENSE-DATA.

Citation

@misc{song2026razorpruningreplaceableexperts,
      title={RAZOR: Pruning Replaceable Experts in LLMs},
      author={Mingyang Song and Mao Zheng},
      year={2026},
      eprint={2609.30465},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.30465},
}
Downloads last month
304
Safetensors
Model size
154B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nickyang/Hy3-RAZOR-154B-A21B-E96of192

Base model

tencent/Hy3
Finetuned
(13)
this model

Collection including Nickyang/Hy3-RAZOR-154B-A21B-E96of192

Paper for Nickyang/Hy3-RAZOR-154B-A21B-E96of192