PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving
Abstract
PEEK improves online LLM serving by clustering queued requests via a radix tree to maximize prefix cache reuse, reduce latency, and boost throughput.
We present PEEK, a lightweight scheduling and eviction framework for both online (streaming) and offline (batch) LLM serving; this paper focuses on the online regime. PEEK maintains an incremental radix tree over the pending queue, exposing prefix-sharing clusters no existing engine surfaces. A low-overhead dual-walk matches the tree against the engine's prefix cache to yield longest-prefix-match for every waiting request; PEEK then admits cluster pioneers first so siblings inherit the freshly cached prefix, a co-designed eviction hook protects blocks ancestral to queued demand, and a multi-lane stride scheduler bounds starvation. On SGLang and vLLM across five workloads up to 4timesH100 (DP=2 over TP=2), PEEK delivers up to 3.0times/2.6times cache hit, 7.9times/7.1times TTFT, 6.7times/5.5times E2E, and 3.6times/4.5times throughput gains over each engine's strongest stock baseline (SGLang/vLLM), while matching baselines within noise on workloads with no exploitable prefix structure. Wins hold as KV-cache pressure and inference parallelism scale.
Get this paper in your agent:
hf papers read 2607.02525 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper