Fast Code Pruner
Task-aware context pruning for coding agents with a 17-layer Qwen2.5-Coder-0.5B backbone and a native vLLM serving path.
GitHub repository · Original SWE-Pruner
Validation quality
Evaluation uses the fixed 6,119-example seed-43 validation split.
| Model | Accuracy ↑ | Precision ↑ | Recall ↑ | F1 ↑ |
|---|---|---|---|---|
| fast-code-pruner | 85.94% | 81.49% | 83.49% | 82.48% |
| code-pruner | 84.07% | 80.02% | 80.91% | 80.46% |
Architecture
Original Code-Pruner Fast Code Pruner
Qwen3 layers 1–28 Qwen2.5-Coder layers 1–17
frozen backbone frozen weights + rank-8 LoRA
│ │
hidden layers 7 + 14 + 28 normalized hidden layer 17
3 × 1024 → concatenate 3072 896 → gated PolyNorm → 2432
│ │
1× bidirectional full attention 1× bidirectional full attention
8 heads, width 3072 16 heads, width 2432
│ │
CRF emissions: 3072 → 256 → 2 CRF emissions: 2432 → 128 → 2
Fast Code Pruner uses only the normalized final-layer representation and physically skips the three removed attention branches. Rank-8 LoRA updates are merged into dense weights during export. Prefix caching is disabled because the bidirectional fusion head requires the complete input.
Serving
Install the native vLLM package directly from this model repository:
pip install \
"fast-code-pruner[vllm] @ https://huggingface.co/irotem98/fast-code-pruner/resolve/main/package/fast_code_pruner-0.1.0-py3-none-any.whl"
fast-code-pruner serve
Inputs longer than 8,192 tokens are split with overlap and merged automatically.
Serving performance
| Model | Backend | Concurrency 1 ↑ | Concurrency 16 ↑ |
|---|---|---|---|
| fast-code-pruner | vLLM 0.13.0 | 79.0 req/s | 143.4 req/s |
| fast-code-pruner | Hugging Face | 16.01 req/s | 16.03 req/s |
| code-pruner | Hugging Face | 9.83 req/s | 10.03 req/s |
The vLLM results are medians of three 1,000-request end-to-end HTTP runs over
the same 100 validation examples on one NVIDIA RTX PRO 6000 Blackwell Server
Edition. Prefix caching was disabled. Median model-input throughput was 56.1k
tokens/s at concurrency 1 and 101.8k tokens/s at concurrency 16.
fast-code-pruner Hugging Face results are medians of three 100-request runs
with an 8,192-token model length, matching the serialized legacy endpoint. The
original code-pruner row uses its Hugging Face backend benchmark.
Training
The public repository includes a lightweight, hard-label launcher matching the released architecture:
bash train/train_fast_code_pruner.sh
It contains no teacher-target generation or knowledge-distillation stage.
Attribution
This model builds on SWE-Pruner and Qwen2.5-Coder-0.5B.
@misc{wang2026sweprunerselfadaptivecontextpruning,
title={SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents},
author={Yuhang Wang and Yuling Shi and Mo Yang and Rongrui Zhang and
Shilin He and Heng Lian and Yuting Chen and Siyu Ye and Kai Cai and
Xiaodong Gu},
year={2026},
eprint={2601.16746},
archivePrefix={arXiv},
primaryClass={cs.SE}
}
- Downloads last month
- -