--- license: other license_name: llama-community-licenses license_link: LICENSE.md base_model: - meta-llama/Llama-2-7b-hf - meta-llama/Llama-2-13b-hf - meta-llama/Llama-3.2-3B - meta-llama/Llama-3.1-8B pipeline_tag: text-generation tags: - structured-pruning - layer-pruning - model-compression - knowledge-distillation - llama --- # OverRep Recovery checkpoints for **Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning** (EMNLP 2026). [Paper](https://arxiv.org/abs/2609.06974) · [Code](https://github.com/mmai-laboratory/overrep) Built with Llama. Each folder holds the two recovery blocks of one pruned model, already merged into plain Hugging Face decoder-layer weights. For a mask `T:I`, layers `T+1 … T+I` of the base model are removed and the blocks at `T` and `T+I+1` are replaced by the weights in `layer_T.safetensors`. The result is a standard pruned model with no extra parameters. | Backbone | Base model | Pruning | Folder | Reasoning RP (%) | Generation RP (%) | | --- | --- | --- | --- | ---: | ---: | | LLaMA2-7B | `meta-llama/Llama-2-7b-hf` | 25% | `llama-2-7b/25pct` | 92.6 | 61.9 | | LLaMA2-7B | `meta-llama/Llama-2-7b-hf` | 50% | `llama-2-7b/50pct` | 76.9 | 29.4 | | LLaMA2-13B | `meta-llama/Llama-2-13b-hf` | 25% | `llama-2-13b/25pct` | 95.3 | 67.2 | | LLaMA2-13B | `meta-llama/Llama-2-13b-hf` | 50% | `llama-2-13b/50pct` | 85.5 | 42.1 | | LLaMA3.2-3B | `meta-llama/Llama-3.2-3B` | 25% | `llama-3.2-3b/25pct` | 90.2 | 47.1 | | LLaMA3.2-3B | `meta-llama/Llama-3.2-3B` | 50% | `llama-3.2-3b/50pct` | 72.3 | 10.7 | | LLaMA3.2-3B (OverRep-KD) | `meta-llama/Llama-3.2-3B` | 25% | `llama-3.2-3b/25pct-kd` | 91.8 | 61.0 | | LLaMA3.2-3B (OverRep-KD) | `meta-llama/Llama-3.2-3B` | 50% | `llama-3.2-3b/50pct-kd` | 73.4 | 26.7 | | LLaMA3.1-8B | `meta-llama/Llama-3.1-8B` | 25% | `llama-3.1-8b/25pct` | 91.3 | 50.0 | | LLaMA3.1-8B | `meta-llama/Llama-3.1-8B` | 50% | `llama-3.1-8b/50pct` | 69.5 | 14.7 | **The last two columns are retained performance (RP), not accuracy:** the pruned model's average score divided by the dense base model's average score, × 100 (100 = no loss from pruning). Reasoning RP averages ten 0-shot tasks (ARC-Easy, ARC-Challenge, BoolQ, HellaSwag, MathQA, MMLU, OpenBookQA, PIQA, RACE, WinoGrande); Generation RP averages CoQA (F1, 0-shot), GSM8K (EM, 8-shot) and TriviaQA (EM, 5-shot). Values are those reported in the paper. Every folder contains `layer_T.safetensors` and `overrep_config.json` (`base_model`, `include_layer`). ## Usage Install the code from GitHub, download one folder, and evaluate it with the mask `T:I` from its `overrep_config.json` (`include_layer`): ```bash hf download SoongE/OverRep --include "llama-3.2-3b/25pct/*" --local-dir checkpoints overrep-eval deploy_model hellaswag --custom \ --tokenizer-name meta-llama/Llama-3.2-3B --include-layer 17:9 \ --checkpoint checkpoints/llama-3.2-3b/25pct ``` The base models are gated on Hugging Face; accept their licenses before downloading them. ## License The checkpoints are derivatives of the Llama base models and are distributed under their licenses: `llama-2-*` under the [Llama 2 Community License](LICENSE_LLAMA2), `llama-3.1-8b` under the [Llama 3.1 Community License](LICENSE_LLAMA3_1), and `llama-3.2-3b` under the [Llama 3.2 Community License](LICENSE_LLAMA3_2). See [LICENSE.md](LICENSE.md) and [NOTICE](NOTICE). ## Citation ```bibtex @inproceedings{oh2026overrep, title = {Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured {LLM} Pruning}, author = {Oh, Seungmin and Lee, Donggeon and Ryu, Jongbin}, booktitle = {Proceedings of the Conference on Empirical Methods in Natural Language Processing}, year = {2026}, publisher = {Association for Computational Linguistics}, url = {https://arxiv.org/abs/2609.06974} } ```