Skip to content
AI.info

Research

KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems

Overview Research area: Efficient inference and serving systems for large language models, specifically KV cache sharing across prompt-specialized agents in multi-agent systems (ML/systems). Technical

KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
arXiv
2609.34060
Published
2026-09-28
Authors
Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang, Jae-Joon Kim

AI summary

Overview

Research area: Efficient inference and serving systems for large language models, specifically KV cache sharing across prompt-specialized agents in multi-agent systems (ML/systems).

Technical level: Advanced. The paper assumes familiarity with transformer KV caches, prefill versus decode, multi-agent workflows, low-rank matrix factorization, and LLM serving metrics such as TTFT and QPS.

Scope: A single paper presenting KVCMAS, a training-free online KV cache correction framework that stores cross-agent cache deviations in compact low-rank form and chains corrections along the agent workflow.

What This Paper Is About

In prompt-specialized multi-agent systems, several agents share one model but prepend their own distinct role prompts ("prefixes"). Because those prefixes differ, each agent generates a different KV cache for the same shared context, so every agent must repeatedly prefill the growing shared text and keep its own memory-heavy cache. KVCMAS aims to reuse one agent's cache for the next agent while correcting the deviation introduced by the different prefix, doing so online for dynamically changing context, without a separate reference prefill, and without large memory overhead.

Key Contributions

  1. Low-rank anchor pools. The authors observe that cross-agent KV cache deviations live in a low-dimensional feature space (effective rank below 32 across workloads), and store both base cache representations and delta corrections as truncated-SVD factors, reducing anchor memory from O(VNLD) to O(VNr(L+D)).

  2. Chained correction along the workflow. Instead of correcting relative to a separately constructed context-free reference, KVCMAS corrects each shared segment relative to the source agent's already-produced cache, densely prefilling only the first agent. This preserves an exact first-agent cache and removes an out-of-workflow reference prefill.

  3. Online correction for dynamically changing shared context. KVCMAS maintains an anchor pool per shared placeholder slot with similarity-weighted combination of corrections and an entropy-based reliability gate, so it handles new user queries and agent outputs rather than only recurring context relations.

  4. Comprehensive evaluation across language and vision-language workloads. Results span MMLU, GSM8K, HumanEval, MathVista and Video-MME, comparing against NonShared, FullShared, DroidSpeak, CacheBlend, RelayCaching, GraphFlow and KVComm on accuracy, latency, TTFT, peak GPU memory and reuse ratio.

Main Findings

  • Accuracy. KVCMAS achieves the highest accuracy among KV cache sharing methods on three of the five benchmarks (MMLU 65.45%, MathVista 53.80%, Video-MME 52.36%) and ranks second to KVComm on GSM8K (79.53% vs 80.14%) and HumanEval (84.47% vs 85.09%), within 0.7 percentage points. NonShared reference accuracies are 67.54%, 85.60%, 86.96%, 55.90% and 53.00% respectively; FullShared drops sharply to 36.45%, 29.42%, 80.75%, 50.20% and 37.25%.

  • Memory. KVComm's full-dimensional anchor pool uses up to 3.7x more peak GPU memory than KVCMAS. In the reported table, KVComm peak memory is 52.40 GB on MMLU, 59.18 GB on GSM8K, 59.19 GB on MathVista and 53.27 GB on Video-MME, versus 16.06 GB, 16.10 GB, 16.77 GB and 16.61 GB for KVCMAS. A motivating example: Llama-3.1-8B in BF16 with a 4K-token segment pool, V=20 and corrections for N=4 agents requires approximately 50 GiB for full-dimensional anchor states alone.

  • Serving efficiency. KVCMAS provides a 2.0x TTFT speedup over NonShared at 32K shared tokens and 8 QPS under the rho=0.8 setting, and 27% lower TTFT than GraphFlow, described as the next-fastest KV cache sharing work. KVComm runs out of memory at 8K and 32K shared context in these experiments.

  • Chained versus non-chained. At 32K shared tokens and 8 QPS, chained correction reduces median TTFT by 34% by removing the reference prefill that would otherwise compete with concurrent requests for batching and compute resources.

  • Low-rank structure. Delta corrections show effective rank below 32 across workloads, while PF-induced deviations are more compact, with effective rank around 4-6 and rank 8 retaining over 90% of the singular-value energy. Base caches exhibit a similar low-rank structure with slightly higher effective ranks.

  • Limits of selective recomputation. In the reconstruction-coverage analysis, at least half of the KV cache reuse error remains outside the recomputed subset across all workloads. On MathVista and Video-MME, deviation and attention selection remove only about 20% of the error and remain close to random selection. Coverage values range from 0.03 (layer selection, MMLU) to 0.50 (layer selection, MathVista).

  • Latency. KVCMAS provides the lowest end-to-end latency and TTFT among the delta correction methods across all workloads and lower end-to-end latency than the selective recomputation methods except on GSM8K, despite using lower reuse ratios.

  • Ablations. Rank r=32 and V=10 anchors capture most of the accuracy gains; larger settings give only marginal improvements at higher memory cost. KVCMAS retains the lowest TTFT and highest throughput as the number of agents and interaction rounds increases.

Methodology in Plain English

The starting point is a simple observation: agents in a workflow alternate between their own private prefix segments (system prompts and "glue" prompts placed before another agent's output) and shared segments (the user query, task observations, and other agents' outputs). Only the shared segments are candidates for reuse, but the private prefixes shift the cache values for those shared segments.

Rather than recomputing part of the cache through the model (selective recomputation) or building a separate prefix-free reference cache (delta correction), KVCMAS does three things. First, it densely prefills the very first agent so that the workflow begins from an exact cache. Second, for each subsequent shared segment, it compares the source agent's cache against a pool of stored "anchors" for that slot and estimates the needed correction as a weighted combination of the anchors' stored corrections, with weights set by cache similarity. Third, it treats the corrected cache as the reference for the next edge in the workflow, so corrections are chained rather than always measured against an external reference.

To keep memory small, the stored base caches and corrections are compressed with truncated SVD into low-rank factors, and the correction is applied directly from those factors to produce a full-dimensional cache for attention. A confidence gate based on normalized entropy of the matching scores decides whether to accept the correction or fall back to dense prefill, with a threshold tau controlling the reuse ratio rho.

Experiments use the KVComm multi-agent framework (built on AgentPrune and GPTSwarm), with Llama-3.1-8B-Instruct for MMLU and GSM8K, Qwen2.5-Coder-7B-Instruct for HumanEval, and LLaVA-OneVision-7B for MathVista and Video-MME. Each accuracy workload uses three task-specific agents followed by a reflection agent. Selective recomputation is evaluated at rho=0.9 and 0.8; delta correction at approximately rho=0.8 and 0.6 for LLM and VLM workloads. Accuracy is averaged over three runs. Latency, TTFT and peak memory are measured on an NVIDIA A100 80 GB GPU at 1 QPS with TTFT averaged across each trajectory, and single-stream memory/throughput measurements use an NVIDIA A6000 48 GB GPU. Serving efficiency uses controlled four-agent traces that fix the schedule and incorporate 128-token agent outputs into subsequent inputs.

Why This Matters

Impact on research. The paper reframes cross-agent cache correction as a low-rank online estimation problem and shows that the choice of reference cache (workflow-relative versus context-free) changes aggregate approximation error. It also documents the reconstruction coverage limits of selective recomputation, which helps explain why layer and token selection methods degrade on vision-language workloads.

Real-world applications:

  • Enterprise multi-agent assistants where a planner, a tool executor, a critic and an orchestrator share one base model but use different role prompts.
  • Vision-language agent pipelines that process long video or document contexts, where the paper reports the largest accuracy gaps for token-level selective recomputation.
  • Coding assistants with separate generation, testing and review agents operating over the same large repository context.
  • High-concurrency serving deployments where TTFT and GPU memory, not raw throughput, dominate customer experience and cost.

Industry relevance. Memory reduction up to 3.7x relative to KVComm and a 2.0x TTFT speedup over no cache sharing at 32K shared tokens and 8 QPS translate directly into serving capacity and latency improvements on fixed GPU hardware. The approach requires no additional training or calibration, and the authors note that system-level cache optimizations are complementary.

Future Directions

  • Extending the anchor pool regime. The paper evaluates up to V=10 anchors per shared slot and reports that larger pools give only marginal accuracy gains; how the method scales to much larger pools, longer shared contexts, or many more agents and interaction rounds remains an open question beyond the reported appendices.
  • Applying low-rank correction to recurring prefixes. The paper shows PF-induced deviations have effective rank around 4-6 and suggests profiling them once, but focuses its main evaluation on the harder dynamically changing PH case; combining both regimes systematically is a natural next step.
  • Interaction with other serving optimizations. The authors describe system-level cache optimizations as complementary to cross-agent correction, leaving open how KVCMAS composes with other cache management and scheduling techniques.
  • Generalization beyond the evaluated model and workload set. The results cover four models and five benchmarks within the KVComm/AgentPrune/GPTSwarm frameworks; behavior under other workflow topologies, model families, and precision settings is not reported.

Target Audience

Researchers and engineers working on LLM inference systems, KV cache management, and multi-agent serving infrastructure. The paper is most useful to readers who already understand transformer prefill/decode behavior and serving metrics like TTFT, and who want a memory-efficient, training-free alternative to selective recomputation and existing delta correction methods. Practitioners building production multi-agent deployments with shared base models will find the memory and latency comparisons particularly relevant.

Authors’ abstract

Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and construct a separate cache with high computation and memory overhead. Selective recomputation reduces this redundancy but still retains substantial model execution, while existing delta correction methods either support only recurring context relations or maintain memory-intensive online correction states for dynamically changing context. For first seen shared context, these methods also construct a reference cache outside the agent workflow, and an approximate correction at the first agent affects the outputs passed to subsequent agents. We present KVCMAS, an online KV cache correction framework that represents cross-agent cache deviations using compact low-rank states and seamlessly chains corrections along the agent workflow without an additional reference prefill. This design supports dynamically changing shared context while preserving an exact first-agent cache. Across multiple language and vision-language workloads, KVCMAS matches or improves the accuracy of prior KV cache sharing methods while achieving the lowest TTFT under highly concurrent serving. Under controlled serving traces, it provides a 2.0x TTFT speedup over inference without KV cache sharing and reduces peak GPU memory by up to 3.7x relative to a prior KV cache correction method. These results establish KVCMAS as an accurate and scalable KV cache sharing approach for prompt-specialized multi-agent serving.

Read the original paper