Research
More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
Overview Research area: Efficient LLM inference — sparse attention and attention-head architecture co-design (cs.CL, per the arXiv listing). Technical level: Advanced. The paper combines a systems-lev

- arXiv
- 2610.04753
- Published
- 2026-10-03
- Authors
- Noam Elata, Itay Lamprecht, Mikey Shechter, Daniel Ohayon, Itay Hubara, Daniel Soudry
AI summary
Overview
- Research area: Efficient LLM inference — sparse attention and attention-head architecture co-design (cs.CL, per the arXiv listing).
- Technical level: Advanced. The paper combines a systems-level latency analysis with a closed-form linear-algebra derivation (minimum-weight perfect matching on attention-logit reconstruction error).
- Scope in one sentence: The paper analyzes how sparse decoding shifts the attention bottleneck from the probability-value multiplication to the query-key multiplication, and exploits that shift with a sparse attention method (Atop-N) plus an architecture (SAGA) that decouples key-head and value-head counts.
What This Paper Is About
Autoregressive LLM decoding is bottlenecked by attention memory and bandwidth costs, and sparse attention methods try to cut those costs by computing only a subset of attention entries. The authors observe that when sparsity is applied, the probability-value multiplication becomes negligible and the dominant cost moves to the query-key step, so the number of key heads can be reduced while keeping many value heads. They build both a sparse attention implementation (Atop-N) and an asymmetric attention architecture (SAGA) around this observation, and add a recipe for converting existing pretrained GQA models into SAGA models.
Key Contributions
- An analysis of LLM inference trade-offs under accelerated sparse attention, identifying that sparsity shifts the decoding bottleneck from the probability-value multiplication to the query-key multiplication.
- Atop-N, a simple and portable approximate top-N sparse attention algorithm for decoding. It reuses exact top-N thresholds computed once every n tokens and identifies top-N entries via per-element comparison rather than sorting.
- SAGA (Sparse Asymmetric Group-Query Attention), an extension of GQA that decouples key heads
h_kfrom value headsh_v, reducingh_kindependently while keepingh_vlarge to preserve capacity at limited decoding cost. - A conversion procedure for pretrained GQA (and naturally MHA) models, using a closed-form merged key projection derived from minimum-weight perfect matching on attention-logit reconstruction error, followed by short distillation fine-tuning.
Main Findings
- Bottleneck shift: Under sparse attention, the probability-value multiplication accesses only
h_q × l_q × Nvalue vectors regardless of the number of key-value headsh_kv, soh_kvaffects cost only through the query-key multiplication. In autoregressive generation (l_q = 1), Atop-N reduces the probability-value cost fromO(h · l_kv · d)toO(h · N · d). - Atop-N standalone speed: For long sequences, Atop-N yields an overall speedup of nearly 2× for the full attention operation.
- End-to-end decoding speedups: Evaluating Llama 3.2-1B (
h_k: 8 → 4) on H200 and Qwen2.5-1.5B (h_k: 2 → 1) on B200, with sequence lengths from 8K to 128K, N ∈ {512, 1024, 2048}, mean per-token latency over 10 runs of 32 generated tokens each. At 128K with batch size 1, SAGA Atop-N reaches up to 2.2× on Qwen and 2.1× on Llama over baseline full attention; at batch size 16 it reaches 2.2× and 2.6× respectively. - Merging alone helps, and scales with batch size: On Llama at batch size 16, SAGA full attention is up to 1.47× faster than baseline full attention at 128K context (reduced query-key cost from fewer key heads). The benefit grows with batch size as queries compete for memory bandwidth.
- Short-context overhead: Combining merged heads with Atop-N introduces overhead at short sequences due to threshold calibration and sparse gather, with gains appearing at long contexts.
- Quality retention under sparsity: On Qwen2.5-1.5B, Atop-N retains 83–97% of full attention quality across 8K, 16K and 32K context lengths on RULER, with most tasks virtually unaffected; degradation is concentrated in multi-key retrieval tasks.
- Training from scratch on SmolLM2-360M (
h_q = 16, three seeds): SAGA(h_k, h_v) = (1, 8)closely tracks the(8, 8)high-capacity baseline despite using a little over half the KV cache, and outperforms both the key-budget-matched(1, 1)baseline and the roughly parameter-matched(4, 4)baseline. - Head count beats head dimension: A Sigma-style variant
(h_k, d_k) = (2, 32)matches SAGA's(1, 64)in KV-cache size and nearly in parameter count (both withh_v = 8,d_v = 64), but reaches 41.7% final validation accuracy versus 42.7% for SAGA. - Larger-scale from-scratch confirmation: A Qwen3-0.6B experiment trained on 50B tokens exhibits the same trend (details in Appendix D.2).
- Distillation results: Converting Qwen2.5-1.5B (
h_k: 2 → 1, a trivial single pair that reduces to averaging) and Llama 3.2-1B (8 key heads merged into 4) via about 23K iterations, the distilled student retains 96.3% of the teacher's aggregate RULER score across 8K, 16K and 32K contexts. - Comparison with other decode-acceleration methods: Sliding window with attention sinks (SW+S), H2O, and KIVI fall below 50% aggregate RULER accuracy at 8K context and degrade rapidly at longer contexts, despite a 4096-token budget for sliding window and H2O.
- General-capability benchmarks: Qwen2.5-1.5B HellaSwag accuracy drops from 0.678 (baseline) to 0.644 (SAGA) and WikiText word perplexity rises from 12.10 to 12.43. Llama 3.2-1B HellaSwag drops from 0.643 to 0.583 and WikiText perplexity rises from 11.98 to 13.09.
- Large-model Atop-N-only evaluation (no retraining needed): On Llama3.1-70B-Instruct with GPQA Main limited to 8K output tokens and GPQA Diamond limited to 16K, with average outputs over 1.2K tokens (Main) and 2.1K tokens (Diamond): baseline scores are 0.4241 (Main) and 0.4343 (Diamond). At N = 2048, 0.4129 and 0.3939; at N = 1024, 0.4330 and 0.3838; at N = 512, 0.4241 and 0.3586. GPQA Main accuracy is unharmed even at N = 512, while GPQA Diamond incurs some degradation.
- Flash attention note: The authors state that flash attention provides little benefit during autoregressive decode (attention is a matrix-vector product with no quadratic attention matrix to tile), so their full-attention baseline uses faster standard compiled attention.
Methodology in Plain English
The authors start by asking where time actually goes in decoding attention once sparsity is applied. They reason through the two attention multiplications and note that a sparse method touching only N entries per query head makes the value-side multiplication tiny, leaving the key-side multiplication as the bottleneck. From this they conclude that the number of key heads, rather than value heads, is the useful lever.
To test the idea without depending on a complicated sparsity kernel, they build a simple approximate top-N method. Instead of sorting all attention values every step, they compute exact thresholds once every n tokens, then reuse those thresholds so that finding top-N entries becomes a cheap parallel comparison. This keeps the implementation portable.
They then define SAGA, which lets h_k and h_v be set independently, and evaluate it in two ways. First, they train small models from scratch with fixed h_q and varying head configurations, comparing SAGA against GQA variants that match its key budget or roughly its parameter count, with three seeds each. Second, they convert existing pretrained GQA models: they find which original key heads to merge by minimizing the squared error in reconstructed attention logits. For a fixed pair of heads, the best merged key projection has a closed-form solution (a weighted combination using the Moore–Penrose pseudoinverse), and the best overall pairing is a minimum-weight perfect matching over a graph whose edge weights are those optimal pairwise errors. After initialization, they distill the student from the teacher using KL divergence plus hard-label cross-entropy over stages of increasing context length, with a data mix including SlimPajama, QA and long-context data plus synthetic needle-in-a-haystack examples. Speedups are measured end-to-end with compiled CUDA graphs.
Why This Matters
This work argues for co-designing attention architecture with the sparsity pattern used at inference time, rather than treating sparsity as a purely runtime add-on. It also shows that the long-standing GQA convention of tying key and value head counts may be suboptimal once decoding is sparse, and it provides a way to retrofit the change onto existing checkpoints instead of requiring expensive pretraining.
Real-world applications:
- Long-context document and code assistants that decode over contexts of tens of thousands of tokens, where per-token latency at 128K is the practical constraint.
- High-throughput serving at large batch sizes, where the paper reports the merging benefit growing as queries compete for memory bandwidth.
- Retrieval-augmented and agentic pipelines that generate long reasoning traces before answering, as explored in the GPQA experiments with long outputs.
- Memory-constrained deployment, since fewer key heads shrink the key projection and the KV cache footprint, enabling longer contexts or larger batches within the same memory budget.
Industry relevance: the approach is designed to be portable (no training required for Atop-N) and compatible with existing GQA/MHA checkpoints and with orthogonal memory-reduction methods such as cache quantization and eviction, which the authors explicitly note SAGA can be combined with.
Future Directions
- Replacing the simple Atop-N method with hardware-aware sparse attention kernels to gain further speedup.
- Compensating for the reduced key capacity in distilled models by reallocating parameters, for example by increasing value head count or dimension, instead of keeping the student identical to the teacher apart from reduced key projections.
- Validating SAGA at full-scale pretraining, since the paper's quality experiments are limited to models of up to 1.5B parameters and its from-scratch study is on SmolLM2-360M (with a Qwen3-0.6B run on 50B tokens).
- More broadly, treating inference-time sparsity patterns as a first-class consideration in LLM architecture design.
Target Audience
Researchers and engineers working on efficient LLM inference, KV-cache optimization, and attention architecture design. The paper is most useful to readers comfortable with attention mechanics, GQA, and decoding-time latency profiling, and to practitioners who want to reduce long-context decode latency on existing pretrained checkpoints without full retraining. Readers seeking large-scale (beyond 1.5B parameter) quality validation will find that the paper states this remains an open question rather than a reported result.
Authors’ abstract
Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while retaining more value heads preserves capacity with limited additional decoding cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple sparse attention method designed to study the interaction between sparsity and head-count asymmetry. We formalize the benefits of this asymmetry theoretically and validate them empirically through latency measurements and quality evaluations on models up to 1.5B parameters. Together, SAGA and Atop-N achieve end-to-end decoding speedups exceeding $2\times$ over our full-attention GQA baseline at long contexts. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants on the evaluated benchmarks. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture, enabling practitioners to benefit from our approach without costly retraining.