Papers
arxiv:2609.32259

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

Published on Sep 29
· Submitted by
Vincent-Daniel Yun
on Oct 2
Authors:
,
,
,
,
,

Abstract

Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose HeteroFold, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8BrightarrowMinistral-3-14B transfer is 10.7times faster than Native Prefill and 1.18--1.47times faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.

Community

Paper submitter

Can LLMs from different model families directly share KV caches, without receiver-side prefill?

Our answer is HeteroFold.

I am so excited to share our new paper: Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs!

Recent work on prefill-free KV cache transfer has shown that LLMs can reuse previously computed context instead of repeatedly performing receiver-side prefill. However, existing approaches have largely focused on models within the same family or with compatible tokenization.

We introduce HeteroFold, a method for prefill-free cross-family KV cache transfer that enables models such as Llama, Qwen, and Ministral to directly share KV caches despite differences in tokenizers, model depth, KV structure, and representation spaces.

HeteroFold aligns tokens across different tokenizers and layers across different architectures, maps the sender’s K/V states into the receiver’s representation space, and calibrates the transferred cache to preserve the receiver’s attention patterns and outputs. Both the sender and receiver remain frozen, and the learned mappings allow the receiver to directly decode from the transferred cache without processing the original context again.

At 32K context length, HeteroFold achieves about 10.7× faster transfer than native receiver prefill. We evaluate all six transfer directions among Llama, Qwen, and Ministral, with strong results across long-context, short-context, and multi-agent settings.

What excites me most is the possibility of making KV caches reusable across model-family boundaries. Instead of different LLMs repeatedly recomputing the same context, heterogeneous models can directly reuse computation from one another.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.32259
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.32259 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32259 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32259 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.