CacheRepair
Repair networks for CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion. Each network predicts residual corrections to independently precomputed document KV caches for one frozen target LLM. The same repairer serves MuSiQue, HotpotQA, MultiHop-RAG, and TriviaQA.
The Qwen repairers are Built with Qwen. The Llama repairers are Built with Llama and are named Llama-CacheRepair S/M/L.
| Model directory | S | M | L |
|---|---|---|---|
Qwen2.5-3B-Instruct |
9M | 30M | 51M |
Llama-3.1-8B-Instruct |
23M | 59M | 93M |
Qwen2.5-14B-Instruct |
42M | 104M | 159M |
Each target/size directory contains model.safetensors and config.json.
The weight file holds repair parameters and the sigma_stale / sigma_delta
normalization tensors in FP32. The target LLM supplies its own frozen embedding.
manifest.json lists exact parameter counts and file identities.
Use
Install the CacheRepair code, then load one repairer with its matching target LLM:
import torch
from transformers import AutoModelForCausalLM
from cacherepair import load_repairer
target = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-3B-Instruct", torch_dtype=torch.bfloat16,
attn_implementation="eager").to("cuda").eval()
repairer = load_repairer(
"gwang3456/cacherepair", target,
subfolder="Qwen2.5-3B-Instruct/M").to(dtype=torch.bfloat16)
Obtain the target LLM from its official source under its access terms. To use
another repairer, change the target LLM and subfolder together. The code
README shows cache construction, residual repair, and query processing.
For the paper's evaluation inputs, build a manifest using the supplied fixed IDs, then generate answers and measure F1/TTFT:
python -m cacherepair.evaluate --model Qwen/Qwen2.5-3B-Instruct \
--manifest inputs/qwen3/musique.json --method repair --size M \
--weights gwang3456/cacherepair --output outputs/qwen3/musique/repair-M
Use --method all for Full Prefill, Stale KV, and all three CacheRepair sizes.
The code's data-preparation guide describes rebuilding inputs from the original
sources; its plotting entry point turns generated results into quality–TTFT
figures.
Training
The repairers use 50,000 generic retrieval examples from ELI5, FEVER, Natural Questions, Wizard of Wikipedia, T-REx, and Structured Zeroshot through CoRAG/KILT. Training supervision is the normalized residual between stale and joint document KV. The target LLM is frozen. Training ends after 75,000 updates of a 100,000-update learning-rate schedule; each configuration records its settings.
Authors
Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu, Guoliang Xing, and Zhenyu Yan.
License
| Repairer family | Applicable terms |
|---|---|
| Qwen2.5-3B-Instruct | Qwen Research License; research and evaluation use |
| Llama-3.1-8B-Instruct | Llama 3.1 Community License and Acceptable Use Policy |
| Qwen2.5-14B-Instruct | Apache-2.0 |
See LICENSE and NOTICE. The target LLMs are obtained separately from their official repositories. The CacheRepair code is Apache-2.0.
Version and citation
This release pairs code version 1.0.0 with the nine repairers
listed in manifest.json in the model repository and
arXiv v1 (September 28, 2026).
For reproducible runs, retain the code commit and the model repository revision.
@misc{wang2026cacherepair,
title={CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion},
author={Genglin Wang and Wangsong Yin and Yeerzhati Abudunuer and Haoxuan Xu and Guoliang Xing and Zhenyu Yan},
year={2026},
eprint={2609.35139},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.35139}
}