Skip to content
AI.info

Research

No Mean Feat: Simple, Strong Baselines for Context Compression

Overview Research area: Natural Language Processing, specifically context compression for efficient Transformer inference. Technical level: Intermediate. Scope: The paper introduces a standardized eva

No Mean Feat: Simple, Strong Baselines for Context Compression
arXiv
2510.20797
Published
2025-10-23
Authors
Yair Feldman, Yoav Artzi

AI summary

Overview

Research area: Natural Language Processing, specifically context compression for efficient Transformer inference. Technical level: Intermediate. Scope: The paper introduces a standardized evaluation suite (BenchPress) for context compression and two simple soft-compression baselines — mean pooling and bidirectional compression tokens — that outperform the widely used causal compression-token approach across six models and three model families.

What This Paper Is About

Context compression replaces long inputs with shorter pre-computed representations so that a language model can reuse the same document cheaply across many queries, which matters for retrieval-augmented generation (RAG). Progress in this area is hard to measure because evaluations differ in datasets, metrics, context lengths, and model scales, and because the common baseline — causal compression tokens — sets a low bar. The paper builds a reproducible evaluation suite for the problem and shows that two very simple methods beat the standard approach.

Key Contributions

  1. BenchPress, a standard, easy-to-reproduce evaluation suite for context compression, applicable to any compression paradigm but instantiated here for soft compression. It spans short (fewer than 1K tokens) and mid-range (fewer than 8K tokens) contexts, uses an explicit in-domain/out-of-domain split, teacher-normalized scoring for fair cross-model comparison, and a standardized training mixture to separate methodological gains from data effects.

  2. Mean pooling, a compression operator that averages adjacent hidden states produced by a bidirectionally-encoding fine-tuned LLM using non-overlapping windows with window size and stride equal to the compression ratio. It adds no parameters beyond the encoder backbone and processes only the original L tokens, whereas compression tokens require an encoder input of size L + L/r.

  3. Bidirectional compression tokens, a modification of the standard causal compression-token approach in which compression tokens attend bidirectionally among themselves while context attention remains causal. This makes the model aware of its compression budget and removes the Matryoshka-style constraint that shorter compressed representations must be prefixes of longer ones. The paper states this modification has not been explored in prior work.

  4. A systematic empirical study across six models from three families and four scales — Llama3.2-1B, Gemma2-2B, and Qwen3-0.6/1.7/4/8B — covering single-ratio and multi-ratio training, short and mid-range contexts, in-domain versus out-of-domain behavior, and model-scale effects.

Main Findings

  • Mean pooling is the strongest method overall, especially at ratios up to 16x. At 128x compression, bidirectional tokens are competitive with or superior to mean pooling in several settings, and Llama3.2-1B consistently shows this pattern.

  • Bidirectional attention during encoding is critical for compression quality. Adding bidirectional attention among compression tokens improves over the causal variant in most settings, with gains ranging from modest to over 11 F1 points depending on model and ratio.

  • Multi-ratio training helps bidirectional tokens but is only a minor trade-off for mean pooling. Bidirectional compression tokens benefit markedly from multi-ratio training, while mean pooling shows a modest degradation. The paper's explanation is that bidirectional attention lets compression tokens "see" how many tokens are available, giving the model a signal about its compression budget, whereas mean pooling must produce a one-size-fits-all representation at each position.

  • Compression quality scales favorably with model size. Teacher-normalized F1 increases across the four Qwen3 scales from 0.6B to 8B, meaning larger models retain more information under compression.

  • Mean pooling's advantage is larger at longer contexts. On LongBench-E QA with contexts up to 8K tokens (Table 4), mean pooling remains superior, with margins over compression-token approaches even larger than in the short-context setting.

  • Causal compression tokens are a weak baseline against external systems. In short-context results (Table 3), the Qwen3-8B teacher scores 74.33 F1 with full context and 23.06 with no context; causal compression tokens reach 47.47 at 128x single-ratio, while mean pooling reaches 47.90. Baseline systems such as PCC Large (Llama3.1-8B) report 37.24 at 128x, and ICAE (Mistral-7B) reports 42.40 at 4x.

  • The in-domain/out-of-domain gap shrinks at higher compression ratios. Using teacher-normalized F1 for Qwen3-8B, the performance gap is higher at low ratios and lower at high ratios; the paper suggests that at high ratios fine-grained contextual detail is already lost to compression noise, which dominates over the domain gap.

  • Encoder capacity matters most in ablations. For mean pooling with Gemma2-2B (Table 8), freezing the encoder costs 13.3 points on average and removing the encoder costs 12.6 points, while freezing the decoder costs 3.6 points, dropping the linear layer costs 0.7 points, and using ratio sampling costs 1.4 points.

  • A persistent gap remains at 128x. All methods trail the teacher at 128x compression, which the paper reads as a sign that compressor architecture advances are needed alongside scaling.

Methodology in Plain English

The authors first define soft context compression formally: a document of length L is mapped to C dense vectors, where C = ceil(L/r) for a compression ratio r, and the goal is for a model using the compressed context to match the conditional distribution of a model using the full context.

They then build the evaluation suite. It covers six short-context reading comprehension datasets — SQuAD, NarrativeQA, HotpotQA, AdversarialQA, TriviaQA (verified subset), and ParaphraseRC — totaling 22,344 samples and 9,742 contexts, with an overall weighted average context length of 375 tokens. Training uses the train splits of SQuAD, NarrativeQA, and HotpotQA (in-domain), while AdversarialQA, TriviaQA, and ParaphraseRC are held out entirely (out-of-domain). A mid-range tier uses QA tasks from LongBench-E with contexts up to 8K tokens: QASPER, MultiFieldQA-en, HotpotQA, and 2WikiMultihopQA, totaling 578 samples and 492 contexts with a weighted average of 5,044 tokens.

Metrics are exact match (EM) and F1, plus teacher-normalized versions computed as (M_fc − M_T^∅) / (M_T − M_T^∅), where M_T is the teacher's full-context score and M_T^∅ is its no-context score. The authors deliberately avoid substring accuracy because it is easily exploitable, which excludes some baselines (such as PISCO) from primary comparisons.

Both baselines are trained by knowledge distillation from a teacher LLM that has full context access, with the loss being the summed KL divergence between the teacher's and student's step-wise token distributions. For multi-ratio training, compressed representations are generated for all ratios in {4x, 8x, 16x, 32x, 64x, 128x} and losses are summed before a single parameter update. Each teacher is first finetuned on the shared training mixture with LoRA; encoder and decoder are initialized from the same instruction-tuned LLM but trained with separate LoRA weights. Training used 48,000 steps, batch size 32, max context length 1024, and max answer tokens 256. Long-context experiments used Qwen3-1.7B (pretrained at 32K context) with a three-stage procedure, 4,800 steps, and max context length 8,192. All experiments except the baseline systems ran on Google Cloud preemptible TPUs using JAX and Flax NNX.

Why This Matters

Impact on research. The paper argues that without a shared evaluation framework it is unclear whether reported compression improvements are real advances or favorable experimental choices. BenchPress, the standardized training mixture, and the teacher-normalized metric give the field a way to compare compression paradigms on equal footing, and the paper demonstrates that even simple baselines can meaningfully advance the state of the art. The finding that causal attention is a poor inductive bias for compression points toward encoder-style or prefix-LM backbones as more natural starting points.

Real-world applications.

  • Retrieval-augmented generation, where the same evidence is repeatedly processed across queries and compression cuts both time and key-value cache costs from dependence on L to dependence on C.
  • KV cache management in long-context serving, since smaller compressed representations reduce memory pressure and caching overhead.
  • Storage and networking in RAG frameworks, where caching compressed representations avoids recomputation but where higher-capacity KV-style compressed representations raise storage and networking challenges.
  • Multi-ratio deployment, where a single compressor handling several ratios at once (as studied here) avoids training separate models per compression budget.

Industry relevance. Inference cost and latency are central concerns for deployed LLM systems handling long documents; the paper reports rough training costs on a v4-64 TPU with a 2B model at 48,000 steps and batch size 32: 4 hours for a teacher model, 23 hours for a multi-ratio compressor, and 10 hours for a single-ratio compressor. Because mean pooling adds no parameters beyond the encoder backbone and has negligible computational overhead, it is attractive for practical deployment.

Future Directions

  • Testing directly whether forward attention gives a model an explicit view of its compression budget, which the paper offers as its explanation for why bidirectional tokens benefit from multi-ratio training while mean pooling does not.

  • Developing new compressor architectures to close the persistent gap between all methods and the teacher at 128x compression.

  • Investigating scaling of individual components, since the experiments only consider compressors where the encoder and decoder are the same size.

  • Extending BenchPress to additional compression paradigms; the suite is designed to apply to soft compression, KV cache methods, and hard prompt compression, and the paper demonstrates this by evaluating LLMLingua2 within the same framework. Closing the evaluation gaps that excluded methods such as GMSA (whose code and models were not publicly available at the time of writing) is also implied.

Target Audience

Researchers and engineers working on efficient LLM inference, retrieval-augmented generation, and long-context modeling who need a reproducible way to compare context compression methods. It is also useful for practitioners choosing a compression operator for deployment, and for authors of new compression methods who need strong, standardized baselines before claiming improvements. The paper is accessible to readers with intermediate familiarity with Transformer attention, knowledge distillation, and LoRA-style finetuning.

Authors’ abstract

Context compression reduces Transformer inference costs by replacing lengthy inputs with shorter pre-computed representations. It carries significant benefits for retrieval-augmented generation (RAG) and has attracted growing research attention. However, progress remains difficult to measure due to inconsistent evaluations and baselines. We design a standard, easy-to-reproduce evaluation suite for context compression, BenchPress, along with simple, high-performance baselines for English reading comprehension. BenchPress supports benchmarking across model scales, datasets, compression ratios, and short ($<$1K tokens) to mid-range ($<$8K tokens) contexts. While the suite is applicable to any compression paradigm, our baselines target soft context compression. We establish two simple baselines that strongly outperform the widely used causal compression-token approach: mean pooling and a bidirectional compression-token variant. Our results show the benefit of bidirectional attention when computing compressed representations, and that simple pooling is an expressive compression operator.

Read the original paper