Research
V-CoLA: Vision Token Compression with Linear Attention
V-CoLA: Vision Token Compression with Linear Attention Overview Research area: Efficient vision-language models (VLMs) — specifically, compressing the vision tokens that dominate VLM input sequences,

- arXiv
- 2610.11251
- Published
- 2026-10-08
- Authors
- Hao Jiang, Yiru Mao, Tianpeng Bu, Hao Zhou, Hongtao Duan, Wang Jing, Bowen Xu, Xin Chen, Lulu Hu, Bin Yang, Yongliang Tao, Minying Zhang
AI summary
V-CoLA: Vision Token Compression with Linear AttentionOverview
- Research area: Efficient vision-language models (VLMs) — specifically, compressing the vision tokens that dominate VLM input sequences, in the setting of hybrid architectures that interleave linear attention with softmax attention (e.g., Qwen3.5, InfiniteVL).
- Technical level: Advanced. The paper combines an information-theoretic argument about bounded recurrent state capacity with a training-free algorithm whose correctness depends on the internals of Gated DeltaNet-style state recurrence and chunk-wise parallel kernels. The high-level idea is accessible, but the mechanism and the implementation section assume familiarity with linear attention.
- Scope: One paper proposing a training-free, linear-attention-aware vision token compression framework (uniqueness-aware importance scoring plus adaptive chunk-wise merging and early exiting), evaluated on Qwen3.5-9B, Qwen3.5-27B and InfiniteVL across eight multimodal benchmarks.
What This Paper Is About
Vision tokens take up far more of a VLM's input sequence than text tokens, so cutting them down is a natural way to save compute. Existing compression methods, however, were built for standard softmax attention: they rely either on attention scores or on feature similarity between tokens, and the paper shows that when these methods are moved to hybrid architectures with linear attention (Qwen3.5, InfiniteVL), they perform no better than randomly pruning tokens. The goal of V-CoLA is a compression method that reads importance signals from the linear-attention state recurrence itself, so it transfers to this new architecture family without any training.
Key Contributions
- A systematic diagnosis of why prior methods fail on hybrid linear-attention VLMs. The paper measures two representative categories — attention-based (FastV) and similarity-based (DART) — against a random pruning baseline on Qwen3.5-9B across MME, MMB, GQA, SQA, T-VQA, POPE, VizWiz and MMStar, and finds they show virtually no advantage over random. It further traces this to the bounded recurrent state of linear attention and its information ceiling.
- A uniqueness-aware token importance criterion derived from state recurrence. V-CoLA combines a long-term memory-reconstruction term (how well a token's key-value pair can be read back out of the final state) with a short-term uniqueness term (how much a token changes the state relative to its neighbors, probed with pseudo queries), and combines them with a weight λ.
- An adaptive chunk-wise token merging strategy. Rather than hard-pruning the top-k tokens, V-CoLA partitions the token sequence into chunks of roughly equal importance mass, so importance-dense regions get finer chunks and uniform regions get coarser chunks, then merges tokens within each chunk by an importance-weighted softmax.
- Implementation-level compatibility with chunk-wise parallelism of linear attention, plus a vision-token early-exit mechanism. The extended Gated DeltaRule supports injected queries for the uniqueness computation, and all vision tokens are discarded beyond layer-24. The paper reports that adaptive token merging adds only 0.86 ms at 8K tokens (about 0.1% of total model prefill cost), and that pruning overhead is only 1.6% of the forward-pass cost.
Main Findings
- Prior methods collapse to roughly random on hybrid architectures. On Qwen3.5-9B, both the attention-based FastV and the similarity-based DART fail to consistently outperform a random token pruning baseline across the eight evaluated benchmarks.
- Shallow-layer attention is not enough in hybrid models. Comparing Qwen3-VL-8B with Qwen3.5-9B shows that vision attention ratio declines rapidly in shallow layers for Qwen3-VL, whereas Qwen3.5 keeps attending to visual content into intermediate layers and only drops low in deeper layers — so shallow-layer attention alone misses important vision tokens.
- Similarity structure is disturbed. With fixed pivot tokens, the redundant tokens identified by DART on Qwen3.5 become markedly less reliable, meaning latent feature similarity is a fundamentally weaker signal in these models.
- The bounded state imposes an information ceiling. Treating the final state S_L as a lossy encoder gives I(M; S_L) ≤ H(S_L) = c·d² = O(d²) versus source entropy H(M) = L·h(k,v) = O(Ld), so the lost information has lower bound H(M | S_L) = Ω(Ld − d²), which grows linearly in L once L > d. The paper notes this argument depends only on the data-processing inequality and a finite, L-independent state, so it extends to bounded-state recurrent backbones generally.
- Linear attention's own retention encodes importance. Layer-wise measurement of visual information retention on Qwen3.5-9B shows retention is relatively high in shallow layers and declines with depth, dropping below 20% in deeper layers; retention also falls as sequence length grows (256, 1024, 4096), and the Pearson correlation of retention-rate rankings between adjacent layers stays high in shallow layers and declines in deeper ones.
- V-CoLA retains 99.5% of original average performance at 50.0% of vision tokens (440 of 880 tokens). Per benchmark at that rate it scores MME 2398.5, MMB 85.0, GQA 61.0, SQA 92.7, T-VQA 80.6, POPE 89.8, VizWiz 68.3, MMStar 49.6, against an upper bound of MME 2398.2, MMB 85.6, GQA 61.1, SQA 92.2, T-VQA 83.2, POPE 89.9, VizWiz 69.2, MMStar 49.3. Comparable baselines at that rate average: FastV 97.5, SparseVLM 97.5, DART 95.5, VisionZip 99.0, DTP 94.2.
- The advantage widens under aggressive compression. At 25.0% tokens (220) V-CoLA averages 94.8 versus VisionZip 94.5, SparseVLM 91.1, FastV 90.9, DART 88.4, DTP 86.3. At 12.5% tokens (110) V-CoLA averages 88.0 while the next best, VisionZip, reaches 86.4, and FastV falls to 63.9.
- Real prefill speedups, not just token counts. Prefill wall-clock falls from 269 ms to 145 ms (1.86x), 88 ms (3.06x) and 61 ms (4.41x) at 2K vision tokens; from 511 ms to 273 ms (1.87x), 153 ms (3.34x) and 94 ms (5.44x) at 4K; and from 1034 ms to 542 ms (1.91x), 289 ms (3.58x) and 168 ms (6.15x) at 8K.
- The extended operator is the source of the 4-query speedup. At 8K tokens with a single query, Extended Gated DeltaRule takes 34.2 ms versus 34.1 ms for the original Gated DeltaRule, but at 4 queries it takes 35.1 ms versus 135.1 ms — a 3.85x speedup on that component.
- Early exiting vision tokens is nearly free in accuracy terms. Perplexity drops sharply and plateaus after layer-24. Exiting at layer-24 yields MME 2401.6, MMB 85.2, SQA 92.3 versus the full-token baseline of 2398.2, 85.6, 92.2, while exiting at layer-12 gives 2169.5, 80.4, 90.6 and at layer-1 gives 967.3, 22.6, 83.7.
- Both main components contribute. At 75.0% token retention, removing adaptive merging drops MME from 2268.7 to 2197.4 and MMStar from 47.2 to 44.3; removing early exit drops MME to 2235.1 and MMStar to 45.6. At 87.5% retention the drops are larger: MME 2012.0 → 1913.4 (without merging) and 1970.3 (without early exit).
- λ = 0.1 is the chosen weighting. Ablating λ gives MME 2360.5 (λ=0.00, importance only), 2388.0 (0.05), 2398.5 (0.10), and 2333.1 (0.90).
- A forward window of 4 pseudo queries is best. MME is 2387.6 for (w₁, w₂) = (0,1), 2398.5 for (0,4), 2376.4 for (−4,0), and 2395.2 for (−4,4); the paper reports that going beyond this window brings no further gains.
- It generalizes across scale and family at a fixed 75.0% compression. On Qwen3.5-27B V-CoLA reaches MME 2414.5, SQA 96.1, POPE 88.8, MMStar 53.7 against the uncompressed 2517.4, 97.0, 90.5, 58.1, beating DART, VisionZip and DTP. On InfiniteVL it reaches 1893.1, 85.8, 86.3, 49.3 against the uncompressed 1998.9, 86.0, 87.9, 53.7, again leading the three baselines.
Methodology in Plain English
The authors start by asking why existing compression methods break on hybrid models. They run attention-based and similarity-based compressors side by side with a random-pruning baseline, then look at how attention is distributed across layers in Qwen3-VL-8B versus Qwen3.5-9B, and at how reliable the similarity signal is in each. Both diagnostics point to the same culprit: linear attention squeezes the whole history into a fixed-size state, so the signals these methods depend on are either absent or distorted.
The proposed fix uses that same squeeze as a feature. Every token writes its key-value pair into the recurrent state, and the state is continually trying to reconstruct those pairs. A token whose pair can still be read back accurately from the final state is treated as important; a token that barely moves the state when it is written — because the model has already seen something similar — is treated as redundant. To measure the second effect robustly, the authors feed a small window of nearby queries into the state before and after a given token is absorbed and compare the two readouts. The two scores are mixed with a weight λ into a single importance value.
Compression is then done by grouping rather than cutting. Instead of keeping the top-scoring tokens, the method slices the token sequence into chunks that each carry roughly the same total importance, then averages the tokens inside each chunk using a softmax over their importance scores with temperature τ (set to 1.0). Regions where importance is concentrated end up in small chunks and are preserved in detail; flat regions get merged aggressively. Finally, because vision information fades in the deep layers — confirmed by an experiment that removes all vision tokens layer by layer and tracks perplexity — the method simply drops all remaining vision tokens after layer-24. Important configuration details: compression is applied at the end of layer-1 (for attention-score baselines, at the first softmax-attention layer), ψ is cosine distance, (w₁, w₂) = (0, 4), λ = 0.1, and everything is evaluated on NVIDIA A100 GPUs with LMMs-Eval.
Why This Matters
- Research impact: The paper argues that the standard toolkit for token compression — softmax attention scores and hidden-feature similarity — does not transfer to hybrid linear-attention backbones, and it backs that up with an information-theoretic lower bound. It also opens a second thread: that the retention and forgetting behavior of a bounded recurrent state is itself a usable importance signal, which the authors note should extend to other recurrent state update rules of the form S_t = A_t S_{t-1} + B_t v_t k_tᵀ.
- Practical relevance to serving: The reported prefill reductions (1.86x to 6.15x across 2K, 4K and 8K vision tokens) and the 0.86 ms merging overhead at 8K tokens are the kind of numbers that matter for GPU serving costs, since prefill dominates latency for image-heavy requests.
Real-world applications (enabled by cheaper vision-token prefill):
- High-resolution image and screenshot understanding, where vision token counts are exactly the regime the paper targets (L ≫ d).
- Document, chart and OCR-style question answering, related to the TextVQA and VizWiz benchmarks used here.
- Assistive and accessibility use cases, related to the VizWiz benchmark, where a smaller, faster model could serve more requests.
- Multi-turn multimodal assistants and long-form dialogue, which the paper explicitly evaluates in its appendix.
Industry relevance: All authors are affiliated with Alibaba Cloud Computing, Alibaba Group, and the paper is explicit about deployment friendliness and about the method being training-free — meaning it can be layered onto an existing hybrid VLM without retraining or fine-tuning.
Future Directions
- Broader architecture coverage. The paper's own limitations section states that evaluation focuses on hybrid VLMs interleaving linear and softmax attention, and that purely Mamba-based or fully linear-attention VLMs are not covered; extending V-CoLA there would require revisiting how the uniqueness criterion behaves without periodic softmax-attention layers.
- Broader linear-attention formulations. The design is grounded in Gated DeltaNet. The authors call for a systematic study across other gated linear attention variants to establish how general the reconstruction-based criterion really is.
- Beyond the evaluated models. The two criteria are claimed to remain well defined for recurrent backbones with compatible matrix-state readouts, including S_t = A_t S_{t-1} + B_t v_t k_tᵀ, and evaluating on such backbones is flagged as a natural next step.
- Transfer of the default configuration. The paper reports that the same settings (λ = 0.1, (w₁, w₂) = (0,4), early exit at layer-24) work across Qwen3.5-9B, Qwen3.5-27B and InfiniteVL without per-model tuning; whether that holds for architectures outside this family is left open.
Target Audience
- VLM efficiency and inference-serving engineers who need concrete prefill latency numbers and a drop-in, training-free compression layer compatible with chunk-wise parallel kernels.
- Researchers working on attention alternatives — linear attention, Gated DeltaNet, Mamba-2, RWKV-style recurrent backbones — who want evidence about what a bounded state does and does not preserve.
- Practitioners benchmarking token compression who need a clear picture of when attention-score and similarity-based methods stop working, and an honest account of the failure modes.
- Graduate students and advanced readers looking for a worked example of using an information-theoretic bound to motivate a practical algorithm design.
Note: the paper does not report a human evaluation, a user study, or a deployment of any kind, and states that no user-facing deployment was conducted.
Authors’ abstract
Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating linear attention (\eg, Qwen3.5), prior methods designed for softmax attention struggle to generalize. Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime. To this end, we propose \textbf{V-CoLA}, an efficient training-free token compression framework specifically designed for linear attention. V-CoLA introduces a novel \textit{uniqueness-aware importance criterion} for identifying critical vision tokens, coupled with an \textit{adaptive token merging strategy} that performs compression. All components are optimized at the implementation level to remain compatible with the chunk-wise parallelism of linear attention, ensuring strong practical value. Extensive experiments across multiple benchmarks demonstrate the superiority of V-CoLA: it achieves 99.5\% of the original performance with only 50.0\% of vision tokens, and over 88.0\% with as few as 12.5\%, while delivering a 1.86$\times$ to 6.15$\times$ prefill speedup.