Marco-Mini-Instruct-REAP70

Marco-Mini-Instruct (17.3B total, ~0.86B active, 256 experts, 8 active per token) with 70% of its experts removed by REAP: 77 of 256 experts kept. These are the full-precision (bf16) safetensors weights, for fine-tuning, re-quantising or running with transformers.

For ready-to-run quantised files at every pruning ratio, see kueizen/Marco-Mini-Instruct-REAP-GGUF.

Load it

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "kueizen/Marco-Mini-Instruct-REAP70"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")

How this was made

  • Pruning: REAP (Router-weighted Expert Activation Pruning, Cerebras Research), using the reference implementation. REAP scores each expert by its router weight times the size of its output over a calibration set, then removes the lowest-scoring experts whole. It is one-shot: no retraining. Seed 42.
  • Calibration set: theblackcat102/evol-codealpaca-v1 (train split, shuffled, seed 42), 64 samples per category, batch size 1, max sequence length 2048 tokens.
  • Experts removed: 70% (77 of 256 kept).

We have not run task benchmarks on these checkpoints (MMLU, coding, multilingual). Test them on your own workload before relying on them.

License and credits

Derived from ATH-MaaS/Marco-Mini-Instruct and released under the same Apache 2.0 licence. What we changed: removed experts with REAP. Nothing else was modified or retrained. REAP is by Cerebras Research: paper, code. All credit for the base model goes to its authors.

Published by Kueizen.

Downloads last month
189
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kueizen/Marco-Mini-Instruct-REAP70

Finetuned
(7)
this model

Paper for kueizen/Marco-Mini-Instruct-REAP70