Skip to content
AI.info

Research

Draft-KV: Learning Useful Latent Communication Between Language Models

Overview Research area: Natural language processing, specifically latent communication between language models (passing internal key–value states between frozen models instead of decoded text), with e

Draft-KV: Learning Useful Latent Communication Between Language Models
arXiv
2609.34754
Published
2026-09-28
Authors
Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang, Peng Zhong, Fengming Zhu, Xi Peng, Linqi Song, Jacky Keung, Jingyu Zhang

AI summary

Overview

Research area: Natural language processing, specifically latent communication between language models (passing internal key–value states between frozen models instead of decoded text), with evaluation methodology for model collaboration.

Technical level: Intermediate. The paper assumes familiarity with transformer attention, key–value (KV) caches, and gated attention, but it explains each interface component and defines its evaluation contrasts explicitly.

Scope: The paper audits five existing latent-communication method–dataset pairs to show that their reported gains do not depend on the transmitted message, then proposes Draft-KV, a 1.05M-parameter gated interface that transmits the key–value states formed while a frozen sharer drafts an answer, and evaluates it across seven model pairs and seven benchmarks.

What This Paper Is About

Latent communication between language models — passing internal states instead of text — is usually judged by whether the receiver's accuracy goes up. This paper shows that higher accuracy does not prove the receiver used the message: replacing a message with one produced for an unrelated question changes accuracy by at most 0.60 points across five audited method–dataset pairs, even where communication adds 15.44 points over the receiver alone. The goal is to build and evaluate an interface whose gain depends on correctly paired content, so that a stronger sharer actually delivers more to the receiver.

Key Contributions

  1. A criterion. The paper defines the pairing gain P = A_M − A_D (Matched minus Deranged accuracy) alongside the system gain G = A_M − A_R (Matched minus Receiver-only), and uses it to expose an "information-use gap" in C2C, LatentMAS, and DLC (Section 3).
  2. A method. Draft-KV sends the key–value states formed at draft positions while the sharer answers the current question, projected into the frozen receiver's layout and read through a gated attention branch. The interface trains 1.05M parameters, 348× fewer than C2C and 36× smaller than DLC (Section 4).
  3. Evidence that gains are content-dependent. Gains depend on paired content in all 35 Public settings; with a Qwen3-8B sharer, MMLU-Redux reaches 78.04% Matched, 36.40% Deranged, and 37.45% Receiver-only.
  4. Evidence that stronger sharers transfer. Scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%, and communication transfers to held-out tasks and to a Private protocol where the two models hold different evidence.

Main Findings

  • Existing interfaces show near-zero pairing gain. Across the five audited method–dataset pairs, pairing gain never exceeds 0.60 points while system gain reaches 15.44 points. For C2C, MMLU-Redux shows a 12.29 pp system gain with a −0.47 pp pairing gain; OpenBookQA shows 13.40 pp with 0.00 pp; ARC-Challenge shows 15.44 pp with 0.42 pp. LatentMAS keeps a positive pairing gain but only 0.60 pp of its 4.86 pp system gain on ARC-Challenge depends on pairing.
  • The retained gain is receiver-conditioned. An adapter-only control — averaging fused-cache outputs under deranged sharer messages — retains 96.06% of Matched accuracy on MMLU-Redux, 95.93% on ARC-Challenge, and 93.16% on OpenBookQA for the Qwen2.5-0.5B-Instruct to Qwen3-0.6B C2C pair. Matched exceeds adapter-only by only 1.69, 2.22, and 3.60 pp respectively.
  • Stronger sharers barely help existing systems. Across C2C's 0.6B–8B sharers, standalone sharer accuracy rises by 29.85 pp while system accuracy rises by 3.54 pp. With the Qwen3-4B receiver used by Chen et al. (2026), moving from an 8B to a 14B sharer raises standalone sharer accuracy by 4.48 pp but changes system accuracy by only 0.77 pp.
  • Draft-KV headline result. With a Qwen3-8B sharer and a frozen Qwen2.5-0.5B-Instruct receiver, Draft-KV reaches 78.04% on MMLU-Redux versus 37.45% for the receiver alone and 36.40% with reassigned messages.
  • Consistent improvement in the Public protocol. Draft-KV beats the receiver alone in all 35 settings, is the most accurate method in 30, and exceeds Text-to-Text in 33, with an average gain of 5.85 points. System gain and pairing gain are positive in all 35 settings, with method means of 24.24 and 25.65 pp.
  • Mismatch harm is usually small, but not always. The harm A_R − A_D has a median of 1.19 pp and stays within 2.5 pp in 33 of the 35 settings, but reaches 11.65 pp on MMLU-Redux with the Qwen3-4B sharer. C2C's pairing gain averages 1.31 pp, stays within ±7 pp everywhere, and its system gain is negative in 29 of the 35 settings.
  • The receiver does more than relay the draft. On ARC-Challenge, Draft-KV recovers 24.1% of the questions the sharer answers incorrectly while losing only 1.1% of those it answers correctly.
  • Split evidence helps. In the Private protocol, Draft-KV beats Text-to-Text in all 14 model–benchmark settings by 6.41 points on average, ranks first in seven and second in seven, and pairing gain is positive in 13 of 14. A deranged packet falls 10.98 pp below Receiver-only, against 1.41 pp under the Public protocol.
  • Scaling with the sharer. With a fixed Qwen2.5-0.5B-Instruct receiver, scaling Qwen3 sharers from 0.6B to 8B raises accuracy from 46.11% to 78.04%. Mean system gain rises by 32.6 pp from 0.6B to 8B, against 30.2 pp for Text-to-Text and 1.4 pp for C2C, and Draft-KV beats the sharer-only reference in 19 of 20 settings.
  • Draft-position states make pairing matter. Replacing draft-position KV with prompt-position KV reduces Matched accuracy by 33.70 pp on ARC-Challenge and 21.48 pp on MMLU-Redux, and collapses pairing gain to 0.17 and 0.09 pp. Removing reconstruction pretraining gives 76.71 Matched / 38.74 pairing gain on ARC-Challenge and 52.40 / 16.58 on MMLU-Redux; removing answer alignment gives 80.03 / 41.72 and 55.40 / 19.25.
  • Adapter-only behavior differs between methods. C2C retains 93.16–116.42% of Matched accuracy across its 12 settings, whereas Draft-KV with the 1.7B and 4B sharers retains 42.28–51.05% and falls below Receiver-only in all six; restoring the correct donor is worth 26.96 and 29.49 pp.
  • Reproduction of the C2C baseline. Three training checkpoints of the Qwen2.5-0.5B-Instruct to Qwen3-0.6B direction and one Llama-3.2-3B-Instruct to Qwen2.5-0.5B-Instruct fuser show 95% per-question paired bootstrap intervals on the Matched–Deranged difference containing zero in 13 of 20 cases across the five Public benchmarks, with the seven exceptions staying within a few points of zero.

Methodology in Plain English

Both language models stay frozen; only a small interface is trained.

The sharer first greedily drafts an answer to the question without seeing the dataset answer. The sharer is then run once over the question plus that draft, and its key and value tensors are extracted at selected layers — after the sharer's native key normalization and before rotary position encoding, so positional information is removed from the transmitted tensors. A visibility mask exposes only draft positions, so the packet is a memory of what the sharer computed while answering.

Two bias-free linear maps per communication layer flatten the sharer's KV heads and reshape them into the receiver's KV-head layout. The receiver reads the packet through an attention branch that reuses its own frozen query projection, normalization, and head layout, so the branch adds no parameters except one signed gate per KV head, initialized to zero. The packet stays a separate memory and is not concatenated into the receiver's own self-attention cache; it joins the residual stream alongside native self-attention. Because the branch borrows the receiver's other weights, interface size follows the two models' cache widths rather than their depth or parameter count.

Training proceeds in three stages that inherit the same parameters. Stage 1 trains message reconstruction: the sharer encodes the last assistant turn of an OpenHermes conversation plus a per-sample transmission key, and the receiver — seeing only a fixed decoding instruction — must reproduce that payload from the packet alone. Stage 2 aligns answers: the packet now comes from the sharer's own draft, and the target is the dataset reference answer, with reconstruction replayed every fifth update. Stage 3 trains on ARC-Easy and ARC-Challenge training data with a one-sided guard: on identical receiver inputs and answer prefixes, the loss compares a Matched packet, a Deranged packet from a fixed derangement, and no packet, and penalizes only when the mismatched packet makes the gold answer more than τ = 0.1 nats per token costlier than sending nothing, with λ_p = 0.1.

Evaluation uses two protocols: Public, where sharer and receiver see the same question across five multiple-choice benchmarks (MMLU-Redux, ARC-Easy, ARC-Challenge, OpenBookQA, C-EVAL), and Private, where the annotated gold evidence of HotpotQA and 2WikiMultihopQA is split at random between the two models. Comparisons are Sharer-only, Receiver-only, Text-to-Text, Cache-to-Cache (either the released fuser or one retrained under the authors' recipe), and Draft-KV, all under Matched, Deranged, and Receiver-only conditions. Baselines include Qwen3 models from 0.6B to 8B, Qwen2.5-0.5B-Instruct, and Llama-3.2-3B-Instruct; each model pair is trained separately.

Why This Matters

The paper changes what counts as evidence for latent communication. Reporting a receiver's accuracy gain is not enough, because the interface itself can supply that gain while making the sharer dispensable — an "information-use gap" the authors measure directly. Pairing gain, not accuracy alone, becomes the standard test of whether a channel carries content, and the paper shows that a small interface can in fact be paid for content, so that a stronger partner's advantage reaches the receiver instead of being stranded.

Real-world applications:

  • Heterogeneous model pipelines. Serving a small, cheap frozen model that is assisted by a larger frozen model without retraining either, changing only a 1.05M-parameter interface per pair.
  • Split-evidence agent systems. Settings such as multi-hop retrieval or distributed agents where each model holds information the other lacks; the Private protocol shows gains exceeding both models answering independently when they hold different evidence.
  • Draft-and-verify and speculative pipelines. Draft states already produced during ordinary decoding become a communication channel, with no extra inference pass beyond the sharer's forward pass over question plus draft.
  • Deployment where weights cannot be modified. Because both models stay frozen and only a tiny gated branch is trained, the approach fits licensed, proprietary, or otherwise immutable checkpoints.

Industry relevance follows from the parameter count: 1.05M trained parameters, 348× fewer than C2C and 36× smaller than DLC, means the trainable footprint is negligible relative to the frozen backbones, and per-pair training is the main cost.

Future Directions

  • Testing whether a single interface can transfer across sharer–receiver pairs, since the paper trains each model pair separately and does not report cross-pair transfer.
  • Extending the protocol beyond multiple-choice option selection and short-text answers to long-form open-ended generation, which the paper does not report.
  • Improving the guard so that mismatch harm stays bounded in the settings where it does not: the harm reaches 11.65 pp on MMLU-Redux with the Qwen3-4B sharer, far above the 1.19 pp median.
  • Extending evaluation to more than two communicating models or to chains of sharers, and understanding why adapter-only retained accuracy for Draft-KV with the 1.7B and 4B sharers falls to 42.28–51.05% and below Receiver-only.

Target Audience

Researchers and engineers working on LLM systems, model collaboration, and multi-agent communication; practitioners building heterogeneous inference pipelines where models are frozen and only small adapters can be trained; and evaluation-focused researchers who need a contrast-based protocol (Matched versus Deranged versus Receiver-only) for measuring whether a communication channel carries content rather than merely adapting the receiver.

Authors’ abstract

Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.

Read the original paper