Research
Retrieval Capacity of Self-Attention Under Competition
Overview Research area: Natural Language Processing / mechanistic interpretability of transformer self-attention (context retrieval, attention sparsity, long-context behavior). Technical level: Interm

- arXiv
- 2609.37879
- Published
- 2026-09-29
- Authors
- Timur Mudarisov, Mikhail Burtsev, Radu State
AI summary
Overview
Research area: Natural Language Processing / mechanistic interpretability of transformer self-attention (context retrieval, attention sparsity, long-context behavior).
Technical level: Intermediate. The paper assumes familiarity with attention weights, value vectors, and negative log-likelihood, but its central experimental logic is simple to follow.
Scope: The paper introduces a method for measuring how many context tokens a pretrained language model must retain at each attention operation to preserve its average predictive loss, and it tests what makes that number grow — context length, competition from background text, and how retained weights are renormalized.
What This Paper Is About
The paper asks a deceptively simple question: how many tokens from its context does a language model actually use, and what determines that number? The authors answer it operationally by keeping only the top-ranked tokens at every attention head, layer, and query (unchanged weights, no retraining) and measuring how much the model's negative log-likelihood degrades as the retained set shrinks. This yields an "effective attention set size," which they then stress-test by extending context, adding background around a fixed supporting fact, and changing how retained weights are aggregated.
Key Contributions
-
A useful-token framework for attention selection. The authors separate an unknown "useful set" of tokens (whose size K(C) varies by sequence, layer, head, and query) from the observable top-N set produced by attention or contribution ranking, and derive recovery bounds linking ranking inversions to the number of tokens that must be retained.
-
Geometric and functional analysis of selected sets. They compare geometric separation of selected vs. random sets (Euclidean and cosine, using precision, recall, and their harmonic summary F_N) against actual loss preservation, and estimate effective attention set sizes across nine decoder-only checkpoints under attention ranking, contribution ranking, and random selection.
-
Evidence that set sizes are context-dependent. They measure the effect of context extension at fixed prediction targets (lengths 256, 512, 1024, 2048) and examine competition via BABILong qa1 experiments that add 0K–4K tokens of background around one fixed annotated supporting fact.
-
Aggregation controls and conditional explanations. They compare deletion against renormalization of retained weights and develop conditional theoretical models showing how ranking competition and preservation of attention mass can force larger set sizes even when the amount of distinct task-relevant information is unchanged.
Main Findings
-
Attention-based selection substantially outperforms random selection. Retaining the highest-attention tokens produces much lower NLL degradation than random sets of the same size, at 1k-token contexts on OpenWebText. Selecting by contribution magnitude (attention weight times value norm) shows the same qualitative pattern.
-
Relatively small selected sets can keep NLL close to the full-attention baseline, but the required size varies across models. Contribution ranking reduces the required size most clearly for Llama-2-7B; for most other models the attention and contribution rankings give similar estimates.
-
Geometric separation is not sufficient evidence of functional sufficiency. Score-based selection (attention and contribution) shows stronger Euclidean separation than random sets across all four models measured geometrically (Qwen-2.5-7B, Gemma-7B, Llama-3-8B, Mistral-7B-v0.3), with smaller advantages under cosine distance. The larger Euclidean differences indicate a substantial magnitude-related component, and geometric separation alone does not determine how large a set is sufficient. Geometry and loss instead move together: the paper reports positive Spearman rank correlations along the plotted F_N-versus-NLL trajectories.
-
Longer context increases the required set size while shrinking its fraction of the context. Using nested suffixes at L ∈ {256, 512, 1024, 2048} and scoring the same final 128 target tokens, the effective attention set size at a 5% relative NLL tolerance grows with length, while the selected fraction of context decreases. Full-model NLL also improves, meaning the added natural context carries useful predictive information. The required size still grows when the allowed absolute NLL increase is held fixed, ruling out the tightening relative criterion as a complete explanation.
-
Additional background displaces a fixed supporting fact. In BABILong qa1, with the same annotated supporting fact and background growing from 0K to 4K, the support tokens move down the attention ranking, receive less total attention mass, and are less often retained among the 64 highest-weight tokens (statistics use 86 examples per background after tokenizer-span validation). Mean candidate answer loss increases from 0K to 4K in all four models, with paired loss-change intervals above zero. The effective set size needed to stay within 0.10 nats of the full model increases substantially in several models, though Qwen's response is weaker and nonmonotonic.
-
Pairwise ranking improves while overall competition worsens. An individual non-support token becomes less likely to outrank a support token as background grows in every model, yet the rising number of competitors produces more tokens ranked above support overall. A conditional ranking model formalizes this: with K fixed useful tokens out of L tokens and competitor scores drawn independently, the expected size needed to retain all useful tokens is K + (L − K)·p_L.
-
How retained weights are combined matters as much as which tokens are kept. Renormalizing the retained attention weights (dividing each by the total retained mass, restoring the sum to one) can substantially reduce the required set size and its increase between background conditions. The local error identities show deletion error is (1 − M_S)·µ_T while renormalized error is (1 − M_S)·(µ_T − µ_S), so rescaling can reduce the perturbation when retained and discarded value means are close — though it can also increase local error.
-
The measurement is an average over the whole model, not a per-operation count. Localized interventions reveal substantial variation across layers and heads, so the common set size does not describe a typical individual attention operation. Set sizes must be read as conditional on the evaluated set sizes, and the experiments do not demonstrate an inference speedup because they compute dense attention before selection.
Methodology in Plain English
The authors take already-trained language models and interfere with attention at inference time. At each query position, head, and layer, they compute the normal attention weights, then keep only the top N tokens by some ranking — either attention weight, contribution magnitude (weight times value-vector length), or a random set of the same size used as a control. The retained weights are left untouched, so the attention output is just a partial sum, and the selection is recomputed at every operation during the intervened forward pass. No retraining or fine-tuning happens.
By sweeping N (evaluated at powers of two, with observed crossings refined to integer sizes rather than testing every smaller integer) and recording the average loss increase, they find where each curve crosses a chosen tolerance — the "effective attention set size." For language modeling this is relative NLL increase; for question answering it is candidate-normalized answer loss over six candidates scored by first-continuation-token logits.
They run three families of experiments. First, a geometric comparison: do the selected tokens form a distinct cluster in the space of weighted value vectors relative to their own aggregate, compared with random sets of the same size? Second, a context-length study using nested suffixes of the same 50 documents per model and corpus, always scoring the same final 128 target tokens so the prediction targets are fixed. Third, a controlled-retrieval study on BABILong qa1 where a single annotated supporting fact is held constant while background text grows, letting them track support rank, attention mass, and recall separately from the functional set-size requirement. Finally, the same pipeline is rerun with renormalized weights to isolate the effect of aggregation, and conditional theoretical models are built to show which mechanisms could produce growing set sizes.
Why This Matters
Impact on research. The paper gives a general, retraining-free procedure for converting attention rankings into a quantitative statement about context utilization, and it explicitly separates two explanations that are often conflated: needing more tokens because more information is present, versus needing more tokens because of competition in the ranking or because the retained sum must be preserved. It also provides a caution against reading geometric separation as evidence of relevance — a useful methodological result for the interpretability literature on attention sinks, active/dormant heads, and geometric token selection. The recovery bounds connect ranking inversions to required set size in a form that can be checked against measurements.
Real-world applications:
- KV-cache and context-window budgeting: deciding how aggressively a deployed model's per-query context can be truncated while staying within a loss budget.
- Long-document and retrieval-augmented pipelines: predicting when added context (more retrieved passages) will displace the passage that actually answers the question.
- Evaluation of attention-based pruning or sparsification heuristics before committing to a specific selection rule.
- Diagnosing "lost in the middle"-style failures by tracking support rank, attention mass, and recall as background grows.
Industry relevance. The experiments use open decoder-only checkpoints at deployment-relevant scales (1B to 24B, including 8-bit weights for the 24B model), and the reported finding that renormalization substantially shrinks the needed set size is directly actionable for anyone deciding between zeroing out and renormalizing attention during approximate inference. The authors are explicit that their intervention does not demonstrate an inference speedup, because dense attention is computed before selection — so the near-term value is in measurement and design guidance, not a drop-in speedup.
Future Directions
- Verify the theory's assumptions on real models. The conditional ranking and attention-mass-retention models are not verified against the tested models, and the paper states the experiments do not establish a universal scaling law. Testing whether the K + (L − K)·p_L form holds empirically is a natural next step.
- Develop a general predictor of effective set size. The authors note the observed geometry–loss trajectories do not establish a general predictor; a validated one would let practitioners estimate set size without running the sweep.
- Reconcile the whole-model average with per-layer/per-head heterogeneity. Localized interventions show strong variation across layers and heads that cannot be treated as independent components of the full-model requirement; finding a principled aggregation is open.
- Move from measurement to efficient inference. Since selection currently happens after dense attention, turning the measured effective sparsity into an actual compute or memory saving remains unresolved.
- Better relevance references. Annotated BABILong support is only a partial reference, and the useful set is never directly observed; methods that approximate ground-truth token relevance would sharpen the recovery analysis.
Target Audience
Researchers and engineers working on transformer interpretability, attention sparsity, long-context modeling, and KV-cache or context-budgeting methods. It will be most useful to readers who already understand how attention weights produce a weighted sum of value vectors, and who want a quantitative handle on context utilization rather than a purely qualitative account. Practitioners choosing between attention-pruning strategies will find the deletion-versus-renormalization comparison and the context-length results directly relevant; readers looking for an off-the-shelf speedup will not, since no inference speedup is demonstrated.
Authors’ abstract
How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size needed to stay within a chosen loss tolerance. Relatively small selected sets can keep NLL close to the full-attention baseline, although the required size varies across models. Attention-based selection substantially outperforms random selection. Selected sets exhibit geometric structure, although geometric separation alone does not establish that model loss is preserved. Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range. Experiments with a fixed supporting fact show that additional background pushes its tokens down the attention ranking and reduces their attention mass. Renormalizing the retained weights can substantially reduce the required set size, showing that it also depends on how selected representations are combined. Conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. These results provide a way to measure effective attention set size in language models and investigate its dependence on context, competition, and aggregation.