OverRep / README.md
SoongE's picture
Initial upload
d31a62f verified
|
Raw History Blame Contribute Delete
3.86 kB
---
license: other
license_name: llama-community-licenses
license_link: LICENSE.md
base_model:
- meta-llama/Llama-2-7b-hf
- meta-llama/Llama-2-13b-hf
- meta-llama/Llama-3.2-3B
- meta-llama/Llama-3.1-8B
pipeline_tag: text-generation
tags:
- structured-pruning
- layer-pruning
- model-compression
- knowledge-distillation
- llama
---
# OverRep
Recovery checkpoints for **Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning**
(EMNLP 2026). [Paper](https://arxiv.org/abs/2609.06974) · [Code](https://github.com/mmai-laboratory/overrep)
Built with Llama.
Each folder holds the two recovery blocks of one pruned model, already merged into plain Hugging Face decoder-layer weights. For a mask `T:I`, layers `T+1 … T+I` of the base model are removed and the blocks at `T` and `T+I+1` are replaced by the weights in `layer_T.safetensors`. The result is a standard pruned model with no extra parameters.
| Backbone | Base model | Pruning | Folder | Reasoning RP (%) | Generation RP (%) |
| --- | --- | --- | --- | ---: | ---: |
| LLaMA2-7B | `meta-llama/Llama-2-7b-hf` | 25% | `llama-2-7b/25pct` | 92.6 | 61.9 |
| LLaMA2-7B | `meta-llama/Llama-2-7b-hf` | 50% | `llama-2-7b/50pct` | 76.9 | 29.4 |
| LLaMA2-13B | `meta-llama/Llama-2-13b-hf` | 25% | `llama-2-13b/25pct` | 95.3 | 67.2 |
| LLaMA2-13B | `meta-llama/Llama-2-13b-hf` | 50% | `llama-2-13b/50pct` | 85.5 | 42.1 |
| LLaMA3.2-3B | `meta-llama/Llama-3.2-3B` | 25% | `llama-3.2-3b/25pct` | 90.2 | 47.1 |
| LLaMA3.2-3B | `meta-llama/Llama-3.2-3B` | 50% | `llama-3.2-3b/50pct` | 72.3 | 10.7 |
| LLaMA3.2-3B (OverRep-KD) | `meta-llama/Llama-3.2-3B` | 25% | `llama-3.2-3b/25pct-kd` | 91.8 | 61.0 |
| LLaMA3.2-3B (OverRep-KD) | `meta-llama/Llama-3.2-3B` | 50% | `llama-3.2-3b/50pct-kd` | 73.4 | 26.7 |
| LLaMA3.1-8B | `meta-llama/Llama-3.1-8B` | 25% | `llama-3.1-8b/25pct` | 91.3 | 50.0 |
| LLaMA3.1-8B | `meta-llama/Llama-3.1-8B` | 50% | `llama-3.1-8b/50pct` | 69.5 | 14.7 |
**The last two columns are retained performance (RP), not accuracy:** the pruned model's average score divided by
the dense base model's average score, × 100 (100 = no loss from pruning). Reasoning RP averages ten 0-shot tasks
(ARC-Easy, ARC-Challenge, BoolQ, HellaSwag, MathQA, MMLU, OpenBookQA, PIQA, RACE, WinoGrande); Generation RP averages
CoQA (F1, 0-shot), GSM8K (EM, 8-shot) and TriviaQA (EM, 5-shot). Values are those reported in the paper.
Every folder contains `layer_T.safetensors` and `overrep_config.json` (`base_model`, `include_layer`).
## Usage
Install the code from GitHub, download one folder, and evaluate it with the mask `T:I` from its
`overrep_config.json` (`include_layer`):
```bash
hf download SoongE/OverRep --include "llama-3.2-3b/25pct/*" --local-dir checkpoints
overrep-eval deploy_model hellaswag --custom \
--tokenizer-name meta-llama/Llama-3.2-3B --include-layer 17:9 \
--checkpoint checkpoints/llama-3.2-3b/25pct
```
The base models are gated on Hugging Face; accept their licenses before downloading them.
## License
The checkpoints are derivatives of the Llama base models and are distributed under their licenses:
`llama-2-*` under the [Llama 2 Community License](LICENSE_LLAMA2), `llama-3.1-8b` under the
[Llama 3.1 Community License](LICENSE_LLAMA3_1), and `llama-3.2-3b` under the
[Llama 3.2 Community License](LICENSE_LLAMA3_2). See [LICENSE.md](LICENSE.md) and [NOTICE](NOTICE).
## Citation
```bibtex
@inproceedings{oh2026overrep,
title = {Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured {LLM} Pruning},
author = {Oh, Seungmin and Lee, Donggeon and Ryu, Jongbin},
booktitle = {Proceedings of the Conference on Empirical Methods in Natural Language Processing},
year = {2026},
publisher = {Association for Computational Linguistics},
url = {https://arxiv.org/abs/2609.06974}
}
```