Mistral-7B-NovaKV
NovaKV compresses the KV cache of a frozen Mistral-7B-Instruct-v0.3 by a post-hoc low-rank factorisation of the K and V projections. The cache is 45.4% smaller, measured as a real allocation delta rather than estimated from ranks. The base model is not retrained.
Code, full results and reproduction commands: https://github.com/declercqcharles/NovaKV-public
This is a delta, not a full model
novakv_mistral7b_v03_joint_delta.pt (2.45 GB) carries only the tensors that differ from
mistralai/Mistral-7B-Instruct-v0.3 β the compressed K/V factors and the q/o projections. Every
other weight was verified bit-identical to the base model and is not redistributed here. It cannot
be loaded with from_pretrained; it needs the loader from the repository above, which fetches the
base model and applies the delta on top.
git clone https://github.com/declercqcharles/NovaKV-public
cd NovaKV-public && pip install -r requirements.txt
hf download declercqcharles/Mistral-7B-NovaKV novakv_mistral7b_v03_joint_delta.pt --local-dir .
PYTHONPATH=. python scripts/eval.py --model mistralai/Mistral-7B-Instruct-v0.3 \
--weights novakv_mistral7b_v03_joint_delta.pt \
--ppl --ppl-datasets wikitext2,c4 --ppl-seqlen 2048
Results
ββββββββββββββββββββββββ¬βββββββββ¬βββββββββ
β β Dense β NovaKV β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β PPL WikiText-2 @2048 β 5.50 β 5.76 β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β PPL C4 @2048 β 8.85 β 9.47 β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β RULER β 97.12% β 96.35% β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β LongBench β 43.44% β 39.72% β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β MMLU β 59.98% β 57.44% β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β HumanEval-instruct β 42.07% β 32.93% β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β MBPP-instruct β 42.60% β 35.20% β
ββββββββββββββββββββββββΌβββββββββΌβββββββββ€
β Cache, KB/token β 131.07 β 71.51 β
ββββββββββββββββββββββββ΄βββββββββ΄βββββββββ
Code generation is where the cost concentrates: HumanEval and MBPP lose 7 to 9 points, while
perplexity, long-context retrieval and general knowledge stay close to dense.
Single-stream decode latency, fixed-shape cache with a captured CUDA graph, against the stock
Hugging Face dense path: 10.29 ms/token at context 256, 12.50 at 2048, 18.43 at 8192, versus a
flat ~27 ms for dense. The dense baseline is not itself CUDA-graph captured. Above batch 1 the
compressed path is slower than dense.
RULER is averaged over 12 of its 13 subtasks; ruler_qa_hotpot fetches its data from a host that
has been unreachable since July 2026, and is excluded from the dense baseline identically.
Mistral's optional sliding-window attention is not implemented in this pipeline. It is harmless
with sliding_window: null, which is the case for this checkpoint, but would be silently wrong
otherwise.
License
Apache 2.0, as the base model. The code in the linked repository is MIT.
Model tree for declercqcharles/Mistral-7B-NovaKV
Base model
mistralai/Mistral-7B-v0.3