Research
Sliding-window beats linear attention
Overview Research area: Natural Language Processing, specifically efficient LLM attention, long-context inference, and alternatives to quadratic self-attention. Technical level: Intermediate. The pape

- arXiv
- 2608.28444
- Published
- 2026-08-28
- Authors
- Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais
AI summary
Overview
Research area: Natural Language Processing, specifically efficient LLM attention, long-context inference, and alternatives to quadratic self-attention.
Technical level: Intermediate. The paper uses attention terminology, benchmark names, and memory/speed measurements, but its central comparison is accessible.
Scope: The paper directly compares training-free Sliding Window Attention (SWA) with attention sinks against post-trained linear attention models across multiple LLMs, benchmark tasks, long-context reasoning tasks, and inference speed/memory measurements.
What This Paper Is About
Large Language Models scale quadratically with context length, so every new token increases memory and compute cost. Linear attention post-training has been proposed as a fix, but the paper argues it has not been properly compared to the simpler baseline of SWA with attention sinks. The goal is to show whether training-free SWA with sinks can match or beat post-trained linear attention on knowledge, reasoning, long-context, speed, and memory.
Key Contributions
- Provides a direct comparison between training-free Sliding Window Attention with 4 attention sinks and post-trained linear attention methods across multiple pretrained LLMs ranging from 1.3B to 70B.
- Shows that SWA with sinks matches or beats most linearized models on short-context general knowledge and reasoning benchmarks, while requiring 0 post-training tokens.
- Demonstrates large long-context advantages for SWA on Single Needle-in-a-Haystack (S-NIAH) and BABILong, where post-trained linear attention degrades sharply.
- Measures speed and memory, showing SWA is fastest and has similar or lower memory cost at window sizes smaller than 512, without needing specialized linear kernels or post-training.
Main Findings
- Short-context average performance: SWA obtains the best average downstream performance in 9 out of 11 cases. The exceptions are LoLCATs on Phi-1.5-1.3B, which scores barely higher than SWA (62.5 vs 62.4), and QRWKV6 on Qwen2.5-32B-Instruct, which performs as well as the baseline model (77.3) while SWA has a slight drop (76.6).
- MMLU performance: SWA is always the best-performing non-teacher model on MMLU except for Llama2.0-7B, where DiJiang is better (40.7 vs 39.8 for SWA).
- Summarized recovery: SWA recovers 93.2% of MMLU baseline performance and 99.0% of average baseline performance. QRWKV6 recovers 92.4% of MMLU and 99.1% of average baseline performance. SWA requires 0 tokens, while LoLCATs fine-tunes on 40M tokens to recover 83.2% of MMLU and 97.5% of average baseline performance.
- **Exact match
Authors’ abstract
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional token, the keys and values must be stored in memory indefinitely, which is unsustainable. Several alternatives have been proposed to fix the quadratic scaling problem, one of which is retrofitting LLMs to use Linear Attention. This idea has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, this line of research has not been properly compared to simpler baselines. In this work, we show that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models. We observe this across multiple LLMs on various downstream tasks. For long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no post-training, is extremely fast, and requires low memory; therefore, making it an extremely cheap and reliable solution. To reduce inference memory cost, we strongly recommend switching to SWA instead of post-training linear models. Linear attention models may have shown some promise, but they likely require to be trained from scratch or extensive post-training in order to even match SWA.