Marco-Nano-Instruct-REAP15

Marco-Nano-Instruct (8.6B total, ~0.6B active, 232 experts, 8 active per token) with 15% of its experts removed by REAP: 198 of 232 experts kept. These are the full-precision (bf16) safetensors weights, for fine-tuning, re-quantising or running with transformers.

For ready-to-run quantised files at every pruning ratio, see kueizen/Marco-Nano-Instruct-REAP-GGUF.

Load it

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "kueizen/Marco-Nano-Instruct-REAP15"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")

How this was made

  • Pruning: REAP (Router-weighted Expert Activation Pruning, Cerebras Research), using the reference implementation. REAP scores each expert by its router weight times the size of its output over a calibration set, then removes the lowest-scoring experts whole. It is one-shot: no retraining. Seed 42.
  • Calibration set: theblackcat102/evol-codealpaca-v1 (train split, shuffled, seed 42), 8 batches per category, batch size 2, max sequence length 512 tokens.
  • Experts removed: 15% (198 of 232 kept).

We have not run task benchmarks on these checkpoints (MMLU, coding, multilingual). Test them on your own workload before relying on them.

License and credits

Derived from ATH-MaaS/Marco-Nano-Instruct and released under the same Apache 2.0 licence. What we changed: removed experts with REAP. Nothing else was modified or retrained. REAP is by Cerebras Research: paper, code. All credit for the base model goes to its authors.

Published by Kueizen.

Downloads last month
193
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kueizen/Marco-Nano-Instruct-REAP15

Finetuned
(4)
this model

Paper for kueizen/Marco-Nano-Instruct-REAP15