Research
Receiver-Conditioned Latent Communication gives 94% CacheBack
Overview Research area: Multi-agent LLM systems, latent (KV-cache) communication between agents, and query-based KV-cache compression for long-context inference. Technical level: Advanced. The paper a

- arXiv
- 2609.32046
- Published
- 2026-09-25
- Authors
- Maximillian Rossi, Prajwal Raghunath, Haoqing Xuan, Yusen Zhang, Eugene Wu
AI summary
Overview
- Research area: Multi-agent LLM systems, latent (KV-cache) communication between agents, and query-based KV-cache compression for long-context inference.
- Technical level: Advanced. The paper assumes familiarity with transformer attention, KV caches, prefill/decode, and agent topologies.
- Scope: The paper introduces "receiver-conditioned communication," in which a receiving agent tells senders what it needs, and instantiates it as CacheBack, a training-free method that selects which sender KV positions to transmit.
What This Paper Is About
Multi-agent systems split large contexts across agents that must then exchange information. Sending full KV caches avoids expensive text decoding, but a full cache grows with both the sender's context and the number of agents, so the receiver's memory and attention costs grow linearly and often exceed available GPU memory or the context window. This paper asks whether a sender should send only what the receiver actually needs, and shows that a simple, training-free selector based on the receiver's query can cut the transmitted state by 75% (at 4x compression) up to 94% (at 16x compression) while improving both accuracy and latency over same-size text communication.
Key Contributions
- Receiver-conditioned communication. The paper formalizes a protocol in which the receiver passes the sender a query describing its information needs, and the sender uses that query to decide which information to send. The authors argue this is a property of message construction, not of whether the message is text or latent.
- CacheBack. A simple, robust, training-free selector for latent communication that uses the receiver's query attention to score sender positions, with a correction term
S_t^CB = Q_t · C_t²that upweights positions whose attention is concentrated in a few query tokens rather than spread across all of them. - End-to-end evidence that receiver conditioning beats same-size text. Experiments across four model families (dense Transformers, Mamba-attention hybrids, sliding-window attention) and two topologies (parallel fan-in and sequential chain) shift the accuracy–latency Pareto frontier toward faster, more accurate task completion.
- Diagnostics on selection behavior. A held-out diagnostic set of 13 FanOutQA tasks shows the receiver query identifies needed evidence better than receiver-independent selectors, and that this improves with model scale.
Main Findings
- FanOutQA accuracy and latency gains. Against same-size text, CacheBack improves strict accuracy by 7.3–20.7 percentage points and reduces median TTEOA by 1.3x–8.0x. For Qwen 3 8B at 4x compression: 55.3% vs 40.7% accuracy (+14.7 pp), 192 s vs 612 s median TTEOA (3.2x). For Ministral 3 14B at 32x: 59.3% vs 38.7% (+20.7 pp), 193 s vs 1,537 s (8.0x). For Gemma 4 12B at 4x: 59.3% vs 52.0% (+7.3 pp), 201 s vs 603 s (3.0x). For Nemotron Nano 2 12B at 8x: 50.0% vs 38.7% (+11.3 pp), 102 s vs 131 s (1.3x).
- Message recall. At the selected operating points, CacheBack achieves 91.5%–95.8% message recall versus 84.8%–91.1% for same-size text.
- LongBench v2 gains. Relative-budget CacheBack at 4x improves Qwen 3 8B accuracy by 14.7 pp (54.0% vs 39.3%) with 3.7x lower median TTEOA (237 s vs 884 s); Fixed 64K improves Qwen by 7.3 pp with 2.5x speedup. Nemotron Nano 2 12B Relative at 8x improves accuracy by 9.3 pp with 2.5x speedup; Fixed 32K improves by 6.0 pp with 1.9x speedup.
- The receiver query matters. Replacing the query with a fixed unrelated query ("Why was Le Chaton Fat denied a small-business loan?") lowers message recall at every compression ratio. At 16x, recall falls from 91% to 65% and strict accuracy falls from 46% to 23%.
- CacheBack beats receiver-independent selectors. In the 13-task diagnostic, at 8x compression the strongest receiver-independent selector retains 68% of evidence versus 93% for CacheBack. At 128x, CacheBack reaches 25% evidence recall versus 1–3% for the other selectors.
- Scale helps selection. Mean query attention ranks gold tokens higher as models grow: median gold-token percentile rises from 0.61 (Qwen 3 1.7B) to 0.76 (Qwen 3 32B), where 0.75 corresponds to the top 25% of positions. CacheBack evidence recall at 8x rises from 68% for 1.7B to 93% for 32B. CacheBack's largest gain over mean query attention shifts toward higher compression ratios as model size increases.
- Averaging hides query-specific evidence. Mean query attention can rank a broadly-but-weakly attended position (B) above a position strongly attended by only one query token (A); CacheBack's
C_t²correction addresses this. - Over-compression hurts. For Qwen on LongBench v2, raising compression from 16x to 128x reduces median TTEOA from 165 s to 143 s but drops accuracy from 44.7% to 32.7%.
- Encoding sizes. Full-cache communication costs 144 KiB per position for Qwen 3 8B in bfloat16. Continuous-row encoding is about 18x smaller per position than full KV state; token-ID encoding reduces payload by a further 75x relative to continuous rows. On a 300 Mbit/s Wi-Fi link that is 664 ms versus 9 ms; on the 450 GB/s NVLink used in the evaluation it is 55 μs versus 0.7 μs. Reported experiments use continuous rows.
- Full-cache transfer can fail outright. With Qwen 3 8B in the FanOutQA settings, full-cache transfer leaves insufficient context for receiver generation on every task.
- Text is slow to generate. Generating three Qwen 3 8B text messages takes a median 73.2 seconds, measured over 50 questions on one NVIDIA H100 80GB node.
Methodology in Plain English
The researchers model an agent system as a directed acyclic graph where nodes send messages along edges. A sender first processes its input tokens, optionally generates G_i latent steps, and ends up with a sequence of T_i = H_i + G_i positions, with one KV entry per layer per position. Rather than sending all of that, the receiver appends a short query q (in these experiments, the task question) after the sender has finished its prefill, and the query's attention over the sender's existing cache is used to score positions.
The scoring rule adapts SnapKV: attention from the query tokens to earlier sender positions is averaged over query tokens and query heads, summed across eligible layers, and pooled over a window of seven sender positions to give a score Q_t. CacheBack then reweights this by C_t², a term that grows when a position's support comes from a concentrated subset of query tokens. The top-scoring positions under a message budget B_i are transmitted, with source positions sent as token IDs or continuous rows and latent steps sent as continuous vectors.
Experiments compare text and latent messages that use the same receiver query, so the comparison isolates how the message is constructed rather than what information need drives it. Two benchmarks are used: FanOutQA (fan-in, three agents each reading different Wikipedia pages, at least 40K source tokens per sender and 120,000 total per task) and LongBench v2 (sequential chain, 100K–246K-token documents split into four equal parts of 25K–61K tokens). Models span Qwen 3 (1.7B/4B/8B), Ministral 3 (3B/8B/14B), Gemma 4 (E2B/E4B/12B), and Nemotron Nano 2 (9B/12B) plus 3 Nano (4B); the receiver is always the largest model in the family. Each benchmark uses 50 tasks submitted as a batch, with p50 and p95 TTEOA reported, on an 8x 80 GB H100 GPU node.
Why This Matters
- It separates two questions that latent communication research had merged. Prior work varied how latent state is represented and transmitted; this paper shows that what is transmitted is a separate, high-leverage decision, and that a full cache is not the right default.
- It recovers the point of splitting contexts. Partitioning documents across agents only helps if the receiver does not have to re-absorb everything the agents read. Equation 2 (
H_R + Σ B_i + O_R ≤ L_R) makes the budget requirement explicit, and the paper's FanOutQA results show full-cache transfer exhausting the receiver window on every task. - It is training-free. CacheBack requires no learned communication module, no fine-tuning, and no changes to model weights, which makes it deployable wherever a KV cache already exists.
Real-world applications:
- Long-document question answering and research assistants, where a task needs evidence assembled from multiple documents or sections that no single agent can hold at once.
- Repository exploration and codebase agents, which the paper names as a case where several agents must return detailed findings before planning can continue.
- Sequential multi-agent pipelines, such as chained summarization or review workflows, where each agent must finish before the next begins and text decoding at every step compounds latency.
- Wide-area or distributed agent deployments, where message size matters: token-ID encoding cut a payload from 664 ms to 9 ms on a 300 Mbit/s Wi-Fi link.
Industry relevance: the paper targets the practical failure mode of multi-agent systems at scale, namely memory, context-window, and latency costs growing with the number of coordinating agents. Cloud providers, agent framework vendors, and teams running concurrent agent workloads on GPU nodes (the setup tested here uses an 8x H100 node) are the direct beneficiaries. Funding comes from NSF awards 2103794, 2312991, and 2551201 plus corporate support from Amazon, IntellectAI, Infosys, Tidalwave, Veris, Shopify, Microsoft, Thinking Machines, Dandy, Perplexity, and Daytona; code is at github.com/maxr0ssi/rclc.
Future Directions
- How should queries be written, and when should a receiver ask again? Because the query is applied after sender computation, one completed sender state can be scored for multiple queries. The paper does not evaluate how many queries a sender can serve before selection and transfer become the bottleneck.
- Adaptive querying and budgets. The authors fix the query to the task question, issue it once after sender computation, and fix the message budget before execution. They suggest a receiver could ask for evidence about an unresolved claim, request a specific file or dependency, or ask for more detail after an initial message, with the budget changing as the task changes.
- Better selectors. CacheBack's scoring rule is deliberately fixed and simple, and hybrid models expose only their global-attention layers to it. Learning the scoring rule or using more of the model state may improve selection; the paper does not establish the best selector.
- Broader workload characterization. The paper notes the advantage narrows when network transfer or receiver prefill costs more than the sender decoding avoided, and does not characterize which workload mix pushes each cost to dominate.
Target Audience
Researchers and engineers working on multi-agent LLM systems, latent/inter-agent communication, and long-context inference; practitioners deploying agent pipelines with memory or context-window constraints on multi-GPU serving infrastructure; and readers interested in KV-cache compression who want to see query-based selection transferred from single-model continuation to cross-agent messaging.
Authors’ abstract
Multi-agent systems distribute large contexts across agents that communicate to solve a task. Text messages are compact but require decoding and may omit evidence the receiving agent needs. Recent latent communication instead transfers KV caches. This avoids text generation and can improve accuracy and latency. However, a full KV cache grows linearly with both the context an individual agent processes, and the number of agents that coordinate together. This raises memory and context costs, often far exceeding available GPU resources and context window sizes. Our key observation is that agents need only send what the receiving agent requires for its local task -- which we call receiver-conditioned communication. The receiver agent passes the sender a small description of its information needs, which serves to filter and compress the sender agent's KV cache. CacheBack is a simple, robust, training-free instance of receiver conditioning based on the sender's attention weights. On FanOutQA, CacheBack with Qwen 3 removes 75% of the state the agent would otherwise receive, improving accuracy by 14.7 percentage points and reducing median task-completion latency by 3.2x relative to text communication. We show comparable improvements across model families that span dense Transformers, Mamba-attention hybrids, and sliding-window attention.