Spaces:
Running
Running
File size: 23,290 Bytes
64850b5 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c cfcfbaf 29aa94c 64850b5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | <!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="theme-color" content="#fafafa">
<meta name="description" content="CacheBack: receiver-conditioned communication between LLM agents. Paper, code, results, and demonstrations.">
<meta property="og:title" content="CacheBack | Receiver-conditioned communication">
<meta property="og:description" content="The receiver asks for what it needs. CacheBack uses that request to select which internal state gets sent, without additional training.">
<meta property="og:type" content="website">
<meta property="og:url" content="https://agentcacheback.github.io/">
<meta property="og:image" content="https://agentcacheback.github.io/social-preview.jpg">
<meta property="og:image:alt" content="CacheBack: receiver-conditioned communication between LLM agents, with accuracy and completion-time results across four model families.">
<meta property="og:image:type" content="image/jpeg">
<meta property="og:image:width" content="1200">
<meta property="og:image:height" content="727">
<meta name="twitter:image" content="https://agentcacheback.github.io/social-preview.jpg">
<meta name="twitter:card" content="summary_large_image">
<link rel="canonical" href="https://agentcacheback.github.io/">
<title>Receiver-Conditioned Latent Communication gives 94% CacheBack</title>
<link rel="icon" type="image/svg+xml" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 32 32'%3E%3Crect width='32' height='32' rx='6' fill='%231c211e'/%3E%3Cpath d='M7 7h7v7H7zm11 0h7v7h-7zM7 18h7v7H7z' fill='%230f766e'/%3E%3Cpath d='M18 18h7v7h-7z' fill='%23fff'/%3E%3C/svg%3E">
<link rel="stylesheet" href="website/assets/styles.css?v=method-order-1">
<link rel="stylesheet" href="demo/communication.css?v=inline-1">
<script type="module" src="demo/communication.js?v=inline-1"></script>
<script type="module" src="website/assets/app.js"></script>
</head>
<body>
<a class="skip-link" href="#main">Skip to content</a>
<header class="site-header wrap">
<a class="brand" href="#" aria-label="CacheBack home"><span class="brand-mark" aria-hidden="true"><i></i><i></i><i></i><i></i></span>CacheBack</a>
<nav aria-label="Main navigation"><a href="#abstract">Overview</a><a href="#results">Results</a><a href="#demo">Demos</a></nav>
</header>
<main id="main">
<section class="hero wrap" aria-labelledby="hero-title">
<div class="hero-copy">
<p class="eyebrow">September 2026</p>
<h1 id="hero-title">Receiver-Conditioned Latent Communication gives <span class="rclc-emphasis">94% CacheBack</span></h1>
<div class="hero-authors">
<ul aria-label="Authors"><li>Maximillian Rossi</li><li>Prajwal Raghunath</li><li>Haoqing Xuan</li><li>Yusen Zhang</li><li>Eugene Wu</li></ul>
<p class="affiliations"><a href="https://daplab.cs.columbia.edu/"><img src="https://daplab.cs.columbia.edu/files/images/daplab_logo_horiz.png" width="1508" height="331" alt="DAP Lab"></a><span>Columbia University</span></p>
</div>
<div class="tldr"><p><b>TLDR;</b> Agents can share <strong>internal model state</strong> instead of generating text messages. Sending all of it, however, overwhelms the receiver.</p><p><strong>The receiver asks for what it needs.</strong> CacheBack uses attention to that request to select which state gets sent. Simple, robust, and <strong>training-free.</strong></p><p>At 16× compression, CacheBack passes on just 6.25% of sender positions.</p></div>
<nav class="resource-links" aria-label="Project resources">
<a href="https://arxiv.org/pdf/2609.32046">Paper <span>PDF</span></a>
<a href="https://github.com/agentcacheback/cacheback">Code</a>
<a href="#demo">Watch demo</a>
<a href="https://arxiv.org/abs/2609.32046">arXiv</a>
<a href="#citation" id="cite-link">Cite</a>
</nav>
</div>
<aside class="hero-results" aria-labelledby="hero-results-title">
<h2 id="hero-results-title" class="eyebrow">FANOUTQA</h2><p class="visual-note">FanOutQA asks agents to combine evidence across documents. Each row compares same-size models: higher accuracy and lower time are better. pp = percentage points.</p>
<figure id="result-summary" class="comparison-plot" aria-labelledby="hero-results-title">
<div class="plot-legend"><span><i></i>Text · same-size senders</span><span class="rclc-emphasis"><i></i>CacheBack</span></div>
<div class="paired-metrics chart-heading"><div><strong>Strict accuracy (%) ↑</strong></div><div><strong>Median time (s) ↓</strong></div></div>
<div class="model-row" data-model="qwen"><div class="model-heading paired-metrics"><div class="model-label"><h3>Qwen 3 <small>8B</small></h3></div></div><div class="paired-metrics"><div class="paired-bars accuracy-bars"><div class="accuracy-guide" style="--bar-end:55.3300%"><span class="gain-bracket" aria-hidden="true"></span><strong class="accuracy-gain">+14.7<abbr title="percentage points">pp</abbr></strong></div><div class="bar-row" role="img" aria-label="Text: 40.7%"><div class="bar-track"><i style="width:40.6700%"></i></div><b>40.7%</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 55.3%"><div class="bar-track"><i style="width:55.3300%"></i></div><b>55.3%</b></div></div><div class="paired-bars latency-bars"><div class="speed-guide" style="--cache-end:11.9938%"><strong class="speedup">3.2× faster</strong></div><div class="bar-row" role="img" aria-label="Text: 612 s"><div class="bar-track"><i style="width:38.2437%"></i></div><b>612 s</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 192 s"><div class="bar-track"><i style="width:11.9938%"></i></div><b>192 s</b></div></div></div></div>
<div class="model-row" data-model="nemotron"><div class="model-heading paired-metrics"><div class="model-label"><h3>Nemotron <small>12B</small></h3></div></div><div class="paired-metrics"><div class="paired-bars accuracy-bars"><div class="accuracy-guide" style="--bar-end:50.0000%"><span class="gain-bracket" aria-hidden="true"></span><strong class="accuracy-gain">+11.3<abbr title="percentage points">pp</abbr></strong></div><div class="bar-row" role="img" aria-label="Text: 38.7%"><div class="bar-track"><i style="width:38.6667%"></i></div><b>38.7%</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 50.0%"><div class="bar-track"><i style="width:50.0000%"></i></div><b>50.0%</b></div></div><div class="paired-bars latency-bars"><div class="speed-guide" style="--cache-end:6.3516%"><strong class="speedup">1.3× faster</strong></div><div class="bar-row" role="img" aria-label="Text: 131 s"><div class="bar-track"><i style="width:8.2023%"></i></div><b>131 s</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 102 s"><div class="bar-track"><i style="width:6.3516%"></i></div><b>102 s</b></div></div></div></div>
<div class="model-row" data-model="gemma"><div class="model-heading paired-metrics"><div class="model-label"><h3>Gemma 4 <small>12B</small></h3></div></div><div class="paired-metrics"><div class="paired-bars accuracy-bars"><div class="accuracy-guide" style="--bar-end:59.3300%"><span class="gain-bracket" aria-hidden="true"></span><strong class="accuracy-gain">+7.3<abbr title="percentage points">pp</abbr></strong></div><div class="bar-row" role="img" aria-label="Text: 52.0%"><div class="bar-track"><i style="width:52.0000%"></i></div><b>52.0%</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 59.3%"><div class="bar-track"><i style="width:59.3300%"></i></div><b>59.3%</b></div></div><div class="paired-bars latency-bars"><div class="speed-guide" style="--cache-end:12.5812%"><strong class="speedup">3× faster</strong></div><div class="bar-row" role="img" aria-label="Text: 603 s"><div class="bar-track"><i style="width:37.6938%"></i></div><b>603 s</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 201 s"><div class="bar-track"><i style="width:12.5812%"></i></div><b>201 s</b></div></div></div></div>
<div class="model-row" data-model="ministral"><div class="model-heading paired-metrics"><div class="model-label"><h3>Ministral 3 <small>14B</small></h3></div></div><div class="paired-metrics"><div class="paired-bars accuracy-bars"><div class="accuracy-guide" style="--bar-end:59.3300%"><span class="gain-bracket" aria-hidden="true"></span><strong class="accuracy-gain">+20.7<abbr title="percentage points">pp</abbr></strong></div><div class="bar-row" role="img" aria-label="Text: 38.7%"><div class="bar-track"><i style="width:38.6700%"></i></div><b>38.7%</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 59.3%"><div class="bar-track"><i style="width:59.3300%"></i></div><b>59.3%</b></div></div><div class="paired-bars latency-bars"><div class="speed-guide" style="--cache-end:12.0813%"><strong class="speedup">8× faster</strong></div><div class="bar-row" role="img" aria-label="Text: 1,537 s"><div class="bar-track"><i style="width:96.0563%"></i></div><b>1,537 s</b></div><div class="bar-row rclc-bar" role="img" aria-label="CacheBack: 193 s"><div class="bar-track"><i style="width:12.0813%"></i></div><b>193 s</b></div></div></div></div>
<div class="paired-metrics chart-axes" aria-hidden="true"><div class="chart-axis"><span>0</span><span>50</span><span>100</span></div><div class="chart-axis"><span>0</span><span>800</span><span>1,600</span></div></div>
</figure>
</aside>
</section>
<section id="abstract" class="abstract-section wrap" aria-label="Communication overview">
<figure class="overview-figure">
<div class="communication-animation">
<button id="animation-play" aria-label="Play animation" title="Play animation">▶</button>
<svg id="communication-svg" viewBox="0 0 1180 548" role="img" aria-label="Animated comparison of full KV transfer and CacheBack"></svg>
<div id="animation-caption">Selected source positions can be sent as <span style="color:#6a4fb8"><b>token IDs</b></span>; latent steps must be sent as <span style="color:#3f669f"><b>continuous vectors</b></span>. The receiver prefills both to build its own state.</div>
</div>
</figure>
</section>
<section id="demo" class="demo-section" aria-labelledby="demo-title">
<div class="wrap">
<div class="demo-heading"><h2 id="demo-title">CacheBack on a coding task</h2></div>
<article class="recorded-demo coding-demo" id="coding-demo" aria-labelledby="coding-title">
<div class="demo-description"><div><p class="demo-type">36 seconds · 4K · Recorded coding task</p><h3 id="coding-title">Fixing a Django bug with seven coding agents</h3><p>Seven Qwen3-8B workers inspect the code and send information to a coordinator, which writes the Django fix. The video compares selected latent state with generated text messages. CacheBack produces the patch <strong class="rclc-emphasis">4.41× faster</strong> than with text communication on this case. Both runs make the same one-line code change and pass all 88 tests.</p></div></div>
<div class="demo-player"><video id="coding-video" controls playsinline preload="auto" tabindex="0" width="3840" height="2160" poster="demo/coding/cacheback-coding-poster.jpg" aria-labelledby="coding-title" aria-describedby="coding-caption">
<source src="demo/coding/cacheback-coding-4k.mp4" type="video/mp4">
<p><a href="demo/coding/cacheback-coding-4k.mp4">Watch the demo as an MP4</a>.</p>
</video><button class="play-demo" id="play-demo" type="button" hidden><span aria-hidden="true">▶</span> Watch the full demo</button></div>
<div class="demo-caption" id="coding-caption"><p><strong>One recorded case:</strong> CacheBack 25.66 s; text 113.21 s. Separate recorded runs are aligned at their start. Startup and test grading excluded. Silent video; playback speed varies. <a href="demo/coding/cacheback-coding-4k.mp4">Open in full window</a></p>
</div>
</article>
</div>
</section>
<section id="results" class="section wrap" aria-labelledby="results-title">
<h2 id="results-title">Results</h2>
<div class="paper-curves" id="paper-curves">
<div class="benchmark-setup"><div class="curves-heading"><h3>FanOutQA</h3><p id="compression-guide">Questions require combining evidence across Wikipedia pages. We split at least 120K tokens among three parallel senders; a receiver answers from their messages. Strict accuracy requires every reference-answer group. <strong>Higher is more accurate; left is faster.</strong> Each panel fixes the receiver model. Text labels give sender size; CacheBack uses same-size senders. <strong>4× compression retains ¼ of sender positions; 16× retains ¹⁄₁₆.</strong></p></div>
<div class="topology-diagram parallel" role="img" aria-label="Three senders read separate Wikipedia pages in parallel. Their messages converge on one receiver, which answers."><div class="senders"><div class="agent-node"><small>Wikipedia pages A</small><strong>Sender 1</strong></div><div class="agent-node"><small>Wikipedia pages B</small><strong>Sender 2</strong></div><div class="agent-node"><small>Wikipedia pages C</small><strong>Sender 3</strong></div></div><div class="merge" aria-hidden="true"><span>messages</span></div><div class="agent-node receiver"><small>Combines 3 messages</small><strong>Receiver</strong><small>↓ Answer</small></div></div></div>
<div class="curve-legend" aria-label="Curve legend"><span class="rclc-emphasis">● ━ CacheBack · labels show compression factor</span><span class="text-emphasis">▲ ┄ Text · labels show sender size</span><span>× Question only</span></div>
<div class="curve-grid">
<figure><img src="website/assets/figures/fanoutqa-qwen.svg?v=compression-labels" width="560" height="460" loading="lazy" alt="FanOutQA, Qwen 3 8B: CacheBack 4× compression reaches 55.3% strict accuracy at 192 seconds, compared with 40.7% at 612 seconds for same-size text; weaker settings are also plotted."><figcaption>Qwen: 4× compression improves accuracy and completion time over 8B text; 1.7B text is slightly faster but less accurate.</figcaption></figure>
<figure><img src="website/assets/figures/fanoutqa-nemotron.svg?v=compression-labels" width="560" height="460" loading="lazy" alt="FanOutQA, Nemotron Nano 2 12B: CacheBack 8× compression reaches 50% accuracy near 102 seconds versus 38.7% near 131 seconds for same-size text. A smaller text sender is faster at lower accuracy."><figcaption>Nemotron: 8× compression improves on 12B text; 4B text is faster but less accurate.</figcaption></figure>
</div>
<div class="chain-results" id="chain-results"><div class="benchmark-setup"><div class="curves-heading"><h3>LongBench v2 Easy</h3><p>Long-document questions test whether evidence survives repeated handoffs. We split each 100K-246K-token document into four equal parts. Each sender reads its part plus the previous message; a final receiver answers from the fourth message. Compression factors retain a fraction of accumulated context (4× keeps ¼); fixed budgets (32K, 64K) cap message positions. Axes and text-size labels follow FanOutQA.</p></div>
<div class="topology-diagram sequential" role="img" aria-label="Four agents read document parts in order. Each passes a message to the next. A final receiver answers from the fourth message."><div class="agent-node"><small>Document part 1</small><strong>Agent 1</strong></div><div class="agent-node"><small>Document part 2</small><strong>Agent 2</strong></div><div class="agent-node"><small>Document part 3</small><strong>Agent 3</strong></div><div class="agent-node"><small>Document part 4</small><strong>Agent 4</strong></div><div class="agent-node receiver"><small>Final message</small><strong>Receiver</strong><small>↓ Answer</small></div></div></div>
<div class="curve-legend" aria-label="LongBench curve legend"><span class="rclc-emphasis">● ━ CacheBack · labels show compression factor</span><span class="fixed-emphasis">■ ━ CacheBack · labels show fixed position budget</span><span class="text-emphasis">▲ ┄ Text · labels show sender size</span><span>× Question only</span></div>
<div class="curve-grid">
<figure><img src="website/assets/figures/longbench-qwen.svg?v=compression-labels" width="560" height="460" loading="lazy" alt="LongBench v2 Easy, Qwen 3 8B: both relative and fixed CacheBack budgets include settings faster and more accurate than same-size text. Stronger compression eventually lowers accuracy."><figcaption>Qwen: 4× compression gives the highest plotted accuracy; stronger compression loses accuracy.</figcaption></figure>
<figure><img src="website/assets/figures/longbench-nemotron.svg?v=compression-labels" width="560" height="460" loading="lazy" alt="LongBench v2 Easy, Nemotron Nano 2 12B: CacheBack 8× compression reaches 48% accuracy near 133 seconds, while fixed-budget settings show a different tradeoff; smaller text senders can be faster."><figcaption>Nemotron: relative 8× compression beats 12B text on both axes; 4B text remains faster but less accurate.</figcaption></figure>
</div>
</div>
<p class="curve-caption">Lines join the best accuracy-time tradeoffs within each channel; faded points are dominated settings. × means the receiver gets only the question. 50 concurrent tasks on 8 H100s; accuracy averaged over three receiver draws. Completion includes queueing. <a href="https://arxiv.org/pdf/2609.32046">Full setup and results in the paper.</a></p>
</div>
</section>
<section id="booking-demo" class="section wrap" aria-labelledby="booking-title">
<div class="booking-heading"><div><p class="eyebrow">Interactive demo · Unstructured documents</p><h2 id="booking-title">Multi-hop document Q&A</h2></div><a class="text-link" href="demo/index.html" target="_blank" rel="noopener noreferrer">Open full window <span aria-hidden="true">↗</span></a></div>
<p class="demo-scenario">In multi-agent question answering, a large document collection can be split across agents, each with its own context. Answering a question requires combining information across those contexts.</p>
<p class="demo-scenario">With text communication, agents write messages summarising what they read. CacheBack lets them pass latent thoughts together with a subset of internal state selected for what the next agent needs. This example shows selected-state handoffs: three agents read separate documents in sequence, and a fourth agent produces the final answer.</p>
<p class="demo-scenario">The top row shows each agent’s input document: a booking confirmation, a location-change notice, then an ID policy. Agents 2 and 3 also receive the previous agent’s handoff; Agent 4 answers from the final handoff alone. The bottom row shows the messages generated by the text agents.</p>
<div class="question-brief"><span>Question</span><h3 id="booking-question">Where should Maya collect her pass on Friday, and what ID should she bring?</h3></div>
<p class="visual-note">Press <strong>Run</strong> to watch the handoffs.</p>
<div class="recorded-demo">
<div class="demo-window"><iframe id="flagship-demo" src="demo/index.html?v=selection-1" title="Multi-hop document question answering: selected internal state versus generated text messages" loading="lazy" sandbox="allow-scripts allow-same-origin" referrerpolicy="no-referrer"></iframe></div>
</div>
</section>
<section class="code-section" id="code" aria-labelledby="code-title"><div class="wrap code-layout"><div><h2 id="code-title">Use CacheBack</h2><p>Install <code>rclc</code>, bind your agents, and pass the receiver’s request.</p><a class="button primary" href="https://github.com/agentcacheback/cacheback">Code and documentation</a></div><div class="code-card"><div class="code-bar"><span>Python</span><button class="copy-button" data-copy="quickstart" type="button" aria-label="Copy code"><svg class="copy-icon" aria-hidden="true" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.7"><rect x="8" y="8" width="12" height="13" rx="2"/><path d="M16 8V5a2 2 0 0 0-2-2H5a2 2 0 0 0-2 2v9a2 2 0 0 0 2 2h3"/></svg><span class="copy-feedback" aria-hidden="true">Copied</span></button></div><pre><code id="quickstart"><span class="code-keyword">import</span> rclc
sender = rclc.<span class="code-function">bind</span>(
model, tokenizer,
messages=sender_history, backend=<span class="code-string">"hf"</span>,
)
receiver = rclc.<span class="code-function">bind</span>(
model, tokenizer,
messages=receiver_history, backend=<span class="code-string">"hf"</span>,
)
<span class="code-focus"><span class="code-keyword">await</span> rclc.<span class="code-function">transfer</span>(
sender, receiver,
<span class="code-string">"Who owns the Cedar booking?"</span>,
)</span>
inputs = receiver.<span class="code-function">pop</span>()</code></pre><p class="snippet-note">Model, tokenizer, and histories loaded; run in an async function or notebook. <a href="https://github.com/agentcacheback/cacheback/blob/main/examples/quickstart.py">Runnable quickstart</a></p></div></div></section>
<section id="paper" class="section wrap paper-section" aria-labelledby="paper-title"><div><p class="eyebrow">THE PAPER</p><h2 id="paper-title">Receiver-Conditioned Latent Communication gives 94% CacheBack</h2><p>Maximillian Rossi, Prajwal Raghunath, Haoqing Xuan, Yusen Zhang, and Eugene Wu.</p><div class="hero-actions"><a class="button primary" href="https://arxiv.org/pdf/2609.32046">Read the PDF</a><a class="text-link" href="https://github.com/agentcacheback/cacheback/tree/paper">Reproduce the paper</a></div></div><div class="citation" id="citation"><h3>Cite this work</h3><button class="copy-button" data-copy="bibtex" type="button" aria-label="Copy citation"><svg class="copy-icon" aria-hidden="true" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.7"><rect x="8" y="8" width="12" height="13" rx="2"/><path d="M16 8V5a2 2 0 0 0-2-2H5a2 2 0 0 0-2 2v9a2 2 0 0 0 2 2h3"/></svg><span class="copy-feedback" aria-hidden="true">Copied</span></button><pre><code id="bibtex">@misc{rossi2026cacheback,
title = {Receiver-Conditioned Latent Communication
gives 94\% CacheBack},
author = {Rossi, Maximillian and Raghunath, Prajwal
and Xuan, Haoqing and Zhang, Yusen
and Wu, Eugene},
year = {2026},
eprint = {2609.32046},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.32046}
}</code></pre></div></section>
</main>
<footer class="site-footer wrap"><a class="brand" href="#">CacheBack</a><span>Receiver-conditioned communication · 2026</span><a href="https://github.com/agentcacheback/cacheback">Code & reproducibility</a></footer>
<div id="copy-status" class="sr-only" role="status"></div>
</body>
</html>
|