Instructions to use Nickyang/Hy3-RAZOR-154B-A21B-E96of192 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nickyang/Hy3-RAZOR-154B-A21B-E96of192 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Nickyang/Hy3-RAZOR-154B-A21B-E96of192") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Nickyang/Hy3-RAZOR-154B-A21B-E96of192") model = AutoModelForCausalLM.from_pretrained("Nickyang/Hy3-RAZOR-154B-A21B-E96of192", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Nickyang/Hy3-RAZOR-154B-A21B-E96of192 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nickyang/Hy3-RAZOR-154B-A21B-E96of192" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nickyang/Hy3-RAZOR-154B-A21B-E96of192", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Nickyang/Hy3-RAZOR-154B-A21B-E96of192
- SGLang
How to use Nickyang/Hy3-RAZOR-154B-A21B-E96of192 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nickyang/Hy3-RAZOR-154B-A21B-E96of192" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nickyang/Hy3-RAZOR-154B-A21B-E96of192", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nickyang/Hy3-RAZOR-154B-A21B-E96of192" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nickyang/Hy3-RAZOR-154B-A21B-E96of192", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Nickyang/Hy3-RAZOR-154B-A21B-E96of192 with Docker Model Runner:
docker model run hf.co/Nickyang/Hy3-RAZOR-154B-A21B-E96of192
Hy3-RAZOR-154B-A21B-E96of192
Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights.
| Base | This model | |
|---|---|---|
| Routed experts per MoE layer | 192 | 96 |
| Total parameters | 295B | 154B |
| Active parameters per token | ~21B | ~21B |
| Active experts per token (top-k) | 8 | 8 (unchanged) |
| MoE decoder layers | 80 | 80 (unchanged) |
Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed.
The other budget is Hy3-RAZOR-226B-A21B-E144of192.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Nickyang/Hy3-RAZOR-154B-A21B-E96of192"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, dtype="bfloat16", device_map="auto",
)
messages = [{"role": "user", "content": "Explain mixture-of-experts routing."}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0]))
Requires a Transformers build containing the native hy_v3 implementation.
Weights are bfloat16.
How the experts were selected
RAZOR asks whether the surviving computation can replace an expert's function, rather than how often or how strongly the expert fires. For a token routed to the set $S$ with normalized weights $w_j$, let $c=\sum_{j\in S} w_j f_j$ be the routed mixture and $r_j=f_j-c$ the consensus residual of expert $j$. Deleting a selected expert $i$ makes the router promote its highest-ranked unselected expert $r$, whose pseudo-weight $w_r$ is its score divided by the original selected score sum. With $\lambda$ the routed-output scale, the exact local output change is
Scores are aggregated by conditional root mean square over the calibration tokens routed to each expert, and the highest-scoring experts are retained per layer. See the paper for the derivation and the code for the implementation.
Reproducing this checkpoint
Calibration used 32,768-token rows from RazorCal, the 2,048-sample multi-domain corpus released with RAZOR.
pip install razor-moe
razor saliency --model <path-to-Hy3> \
--data data/RazorCal.json --max-len 32768 --out out/sal
razor prune --model <path-to-Hy3> --saliency out/sal \
--method razor --ratio 0.5 --out out/pruned
kept_expert_indices.json is the keep-set manifest this checkpoint was built
from, in the format razor verify expects:
razor verify --pruned . --src <path-to-Hy3>
Selection depends on the calibration draw, so an independent run reproduces the procedure rather than this exact expert set.
Limitations
Expert pruning is lossy. Benchmark retention and predictive fidelity do not guarantee stable generation: the paper reports that responses change in diversity, formatting and termination behaviour even where task accuracy is largely preserved. Evaluate on your own workload before deploying.
The calibration corpus is multi-domain but finite, so behaviour on domains far from it is not characterised by the released measurements.
License and attribution
This is a derivative of Hy3, released under the Apache-2.0 license, and is distributed under the same license. The retained weights are the base model's own weights; RAZOR contributes the expert selection, not new parameters.
The RAZOR code is Apache-2.0. RazorCal records remain subject to their applicable upstream terms; see data/LICENSE-DATA.
Citation
@misc{song2026razorpruningreplaceableexperts,
title={RAZOR: Pruning Replaceable Experts in LLMs},
author={Mingyang Song and Mao Zheng},
year={2026},
eprint={2609.30465},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.30465},
}
- Downloads last month
- 304
Model tree for Nickyang/Hy3-RAZOR-154B-A21B-E96of192
Base model
tencent/Hy3