Skip to content
AI.info

Research

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

Overview Research area: Efficient inference and key-value (KV) cache reuse for multi-agent large language model systems, specifically communication between agents built from different model families.

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
arXiv
2609.32259
Published
2026-09-26
Authors
Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram, Sai Praneeth Karimireddy

AI summary

Overview

Research area: Efficient inference and key-value (KV) cache reuse for multi-agent large language model systems, specifically communication between agents built from different model families.

Technical level: Advanced. The paper assumes familiarity with transformer attention, KV caches, tokenizers, rotary position embeddings, and linear algebra (whitening, Procrustes alignment, covariance matching).

Scope in one sentence: The paper introduces HeteroFold, a method that transfers a sender model's KV cache directly into a receiver model from a different family so the receiver can decode without re-processing ("prefilling") the shared text, validated on six transfer directions among Llama-3.1-8B-Instruct, Qwen3-4B, and Ministral-3-14B-Instruct.

What This Paper Is About

In multi-agent LLM systems, agents often combine models from different families for specialized roles, but they exchange shared context as text. That forces every receiving agent to prefill the same context again, rebuilding its own KV cache at added cost and latency as context grows. HeteroFold's goal is to hand the sender's already-computed KV cache straight to the receiver—across different tokenizers, model depths, and KV layouts—while keeping both models frozen and still preserving the receiver's native behavior.

Key Contributions

  1. Tokenizer mismatch identified and solved via Token Alignment (TA). The authors identify tokenizer mismatch as a key obstacle to prefill-free cross-family KV reuse and introduce TA, which establishes sender–receiver token correspondence through shared character-end boundaries (with forward filling for unmatched receiver boundaries and no averaging in many-to-one cases).
  2. A cross-family K/V mapping pipeline. HeteroFold combines layer alignment (proportional depth matching with three sender layers at offsets −4, 0, +4 around a matched layer), cross-head mixing via flattened features, and Recolor—a moment-matched affine initialization built from whitening, Procrustes alignment, and restoration of receiver means and covariances.
  3. Receiver-aware calibration demonstrating that low reconstruction error is not enough. The paper shows KV Ridge achieves lower cache reconstruction error than HeteroFold yet distorts receiver attention, and introduces output-aware calibration (separate key and value objectives) whose learned rank-16 corrections fold into fixed affine maps.
  4. Empirical validation across six cross-family transfer directions. HeteroFold reports the best cache-transfer performance on all four long-context benchmarks and most short-context settings, matches text-based communication on the HiddenBench multi-agent benchmark, and reduces receiver-side latency.

Main Findings

  • Long-context gains. Across all six transfer directions, HeteroFold reports the highest cache-transfer scores on the four long-context benchmarks: Qasper and HotpotQA (token F1), LoCoMo (token F1), and QuALITY (accuracy). For Qwen3-4B → Llama-3.1-8B it reaches 29.12 Qasper F1, 38.48 HotpotQA F1, 29.66 LoCoMo F1, and 62.85 QuALITY accuracy, versus 16.06 / 34.04 / 13.56 / 54.60 for the strongest KV Ridge + TA baseline in that direction.
  • Most short-context settings. HeteroFold leads most short-context benchmarks, but not every one. For example, on Ministral-3-14B → Qwen3-4B HellaSwag, KV Ridge + TA reports 45.12 against HeteroFold's 41.46, and WinoGrande remains a harder setting for cache transfer generally (HeteroFold reports 57.30 on Qwen3-4B → Llama-3.1-8B versus 65.82 for TextMas).
  • Cache error does not equal behavioral fidelity. KV Ridge achieves lower K/V reconstruction error than HeteroFold on Qasper yet distorts receiver attention, which the authors use as the central motivation for calibrating against native receiver attention and outputs rather than reconstruction loss alone.
  • Token Alignment is the single most important component. In the Ministral-3-14B → Llama-3.1-8B ablation, replacing TA with same-index token pairing collapses performance (GSM8K 7.13, Qasper 3.19, HotpotQA 0.72, LoCoMo 1.47, QuALITY 24.59), far below TA + Recolor (62.77, 26.52, 40.71, 40.81, 79.34 respectively).
  • Recolor beats Ridge-style fitting; calibration adds further gains. TA + Ridge reports 4.55 GSM8K and 6.82 HotpotQA F1; TA + Recolor raises these to 62.77 and 40.71; adding both key and value calibration (full HeteroFold) yields 63.46 GSM8K and 50.29 HotpotQA F1, with key correction contributing most of the long-context improvement.
  • Multi-agent decisions match text communication. On all 65 HiddenBench tasks, initial accuracy is 11.5% for every method; after 15 communication rounds HeteroFold reaches 30.6% average accuracy and 29.2% majority accuracy with a 3.2% invalid rate, compared with TextMas at 29.2%, 27.7%, and 4.0%, and baselines Dense Latent + TA (18.5%, 7.7%, 3.2%) and KV Ridge + TA (21.8%, 16.9%, 20.9%).
  • Latency. At batch size 2 with two NVIDIA H100 80GB GPUs connected by NVLink and FlashAttention-2, Llama-3.1-8B → Ministral-3-14B at 32,768 tokens takes 481.3 ms (10.74× faster than Native Prefill's 5167.4 ms), versus 705.7 ms for Dense Latent + TA and 569.9 ms for KV Ridge + TA. Speedup over Native Prefill grows from about 3.5–3.8× at 4,096 tokens to about 10.7× at 32,768 tokens.
  • Alignment quality. On 48 HotpotQA documents at 4K, 16K, and 32K character lengths, mean character-position mismatch between matched sender and receiver tokens grows with context length under same-index pairing but stays low under TA.
  • Behavioral preservation. On QuALITY, HeteroFold more closely preserves the native receiver's next-token distributions, attention weights, and answer choices than both baselines across all six transfer directions.

Methodology in Plain English

HeteroFold works in three stages, and both the sender and receiver models stay frozen throughout—only small mapping parameters are learned.

  1. Line up the tokens. Different tokenizers split the same text differently. Instead of pairing token 1 with token 1 (which drifts badly as documents get longer), TA pairs tokens that end at the same character in the original text. If the receiver has a boundary the sender lacks, it reuses the sender state at the latest preceding shared boundary; many-to-one cases are resolved by picking the sender token at the matched boundary rather than averaging.
  2. Line up the layers and map the features. Because models differ in depth, each receiver layer is matched proportionally to a sender layer, and features from three neighboring sender layers (±4) are concatenated around that match. Keys are mapped before key normalization and RoPE; values after projection; the receiver then applies its own native normalization and RoPE. Mapping is done head-flattened so head counts do not need to agree.
  3. Match the receiver's statistics, then calibrate against the receiver's behavior. "Recolor" replaces plain least-squares fitting with a moment-matching procedure: both feature spaces are whitened, aligned with a Procrustes solution, then rescaled and shifted to reproduce receiver means and covariances. Finally, output-aware calibration adds a small rank-16 correction to each layer's K and V. The key correction tries to match native attention weights and their effect through the output projection; the value correction optimizes the mapped values while holding the attention pattern fixed. These corrections are algebraically folded back into a single fixed affine map per layer and role, so inference needs no extra module, no receiver prefill, and no gold answers at training time—calibration uses only the question span of the prompt.

Calibration uses 1,600 training prompts (800 Open-R1 and 800 HotpotQA) and 400 held-out prompts (200 from each), with AdamW at learning rate 3×10⁻⁴, weight decay 10⁻⁴, gradient clipping at 1, and four epochs, selecting the checkpoint with the lowest combined held-out loss. Baselines Dense Latent and KV Ridge had no public implementations, so the authors reimplemented them, including TA-enabled variants, and evaluated text-based communication (TextMas) with native receiver prefill as the lossless reference.

Why This Matters

Research impact. The paper reframes cross-family cache transfer as a behavior-matching problem rather than a reconstruction problem: it shows empirically that minimizing K/V reconstruction error can make receiver attention worse, and that aligning tokens across tokenizers matters more than any other single design choice. It also demonstrates that KV reuse can cross model-family boundaries, not just same-family ones, which prior prefill-free work (Dense Latent, KV Ridge, LatentMAS) had not addressed.

Real-world applications.

  • Heterogeneous agent pipelines where a solver, evaluator, or aggregator run on different model families and must share the same long document or conversation history.
  • Long-context question answering and document analysis, where the four evaluated long-context benchmarks (Qasper, HotpotQA, LoCoMo, QuALITY) mirror realistic "read this and answer" workloads.
  • Multi-agent decision systems resembling HiddenBench, where several agents with private information deliberate over rounds and converge on a group answer.
  • Serving systems that keep a cheap sender model and an expensive receiver model on separate GPUs, using cache transfer to avoid re-processing shared context.

Industry relevance. The latency numbers translate directly into cost: at 32K tokens the transfer avoids roughly 90% of the receiver prefill time reported for Native Prefill, and the mappers are single fixed affine maps rather than multi-layer networks, which suits deployment on existing inference stacks. Because both models stay frozen, organizations can add cross-family cache sharing without retraining or fine-tuning their production models.

Future Directions

  • Extending beyond the three evaluated model families. The paper covers six directions among Llama-3.1-8B-Instruct, Qwen3-4B, and Ministral-3-14B-Instruct; whether TA, Recolor, and calibration transfer to other tokenizers, depths, and architectures (for example mixture-of-experts or multi-head latent attention designs) is untested here.
  • Closing the remaining short-context gap. Cache transfer still trails text-based communication on short-context tasks such as GSM8K and on settings like WinoGrande and one HellaSwag direction, and the question of when prefill-free transfer is preferable to simply sending text is left open.
  • Calibration cost and mapper construction. The authors include an appendix on mapper size and construction cost, but the practical trade-off between offline calibration effort and per-deployment savings across many model pairs remains an engineering question.
  • Breaking the requirement for aligned calibration prompts. Calibration needs native receiver queries, attention weights, values, and outputs on prompts whose question spans overlap the transfer; scaling this to arbitrary domains, languages, or prompts without held-out calibration data is unresolved.
  • Longer conversations and autoregressive stability. The method is evaluated with communication limited to 15 rounds and generation caps of 160 tokens per message and 384 per vote; behavior over much longer multi-agent dialogues is not reported.

Target Audience

This paper is best suited to machine-learning systems researchers and inference engineers working on LLM serving, KV cache optimization, or multi-agent orchestration, particularly those who need agents on different model families to share context cheaply. It also benefits readers tracking the prefill-free and latent-communication literature (TextMas, C2C, LatentMAS, Dense Latent, KV Ridge), and graduate students comfortable with transformer internals who want a concrete example of aligning representations across models without fine-tuning either one.

Authors’ abstract

Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.

Read the original paper