Skip to content
AI.info

Research

VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm

Overview Research area: Efficient multimodal AI — specifically, training-free token pruning for vision-language models (VLMs) in computer vision. Technical level: Advanced. The method is framed as a c

VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
arXiv
2512.02700
Published
2025-12-02
Authors
Zhenkai Wu, Xiaowen Ma, Zhenliang Ni, Dengming Zhang, Han Shu, Xin Jiang, Xinghao Chen

AI summary

Overview

Research area: Efficient multimodal AI — specifically, training-free token pruning for vision-language models (VLMs) in computer vision.

Technical level: Advanced. The method is framed as a constrained maximum diversity problem with a greedy (near-)submodular optimization scheme, spatial-distance modulation, and similarity-weighted feature aggregation.

Scope: The paper introduces VLM-Pruner, a training-free token pruning paradigm that balances token redundancy against spatial sparsity to preserve fine-grained object detail under aggressive pruning.

What This Paper Is About

VLMs encode images into large numbers of visual tokens, and the number of visual tokens can exceed textual tokens by hundreds or even thousands of times, which degrades efficiency because attention computation in the LLM decoder is quadratic in token count. Existing pruning methods split into importance-driven approaches (which retain similar local regions and therefore keep redundant tokens) and redundancy-reduction approaches (which spread selections too widely and miss fine-grained object detail). VLM-Pruner's goal is a training-free pruning scheme that deliberately trades off between those two failure modes.

Key Contributions

  1. A centrifugal token pruning paradigm. VLM-Pruner is a training-free framework that selects tokens in a near-to-far order — starting from pivot tokens, expanding into spatially adjacent tokens, and finally recovering the outermost information — to balance redundancy and local-detail completeness.

  2. The Buffering for Spatial Sparsity (BSS) criterion. BSS modulates candidate-to-selected token similarity by normalized nearest-selected spatial distance, deferring spatially distant tokens so that pruning proceeds in an ordered way rather than scattering selections across background and foreground.

  3. Similarity-Weighted Aggregation (SWA) for recovery. Discarded tokens are assigned to their most similar retained tokens and their hidden states are aggregated back into the retained set, with 70% of the aggregated hidden state incorporated (β = 0.3).

  4. Extensive empirical validation. Experiments on 13 benchmarks (9 image-language, 4 video-language) across 5 VLMs show consistent outperformance of strong baselines at an 88.9% pruning rate, with end-to-end inference speedups.

Main Findings

  • Consistent gains across pruning rates on LLaVA-1.5-7B: VLM-Pruner preserves 98.85% of the upper-bound average at 192 tokens (66.7% pruning), 98.07% at 128 tokens (77.8% pruning), and 95.61% at 64 tokens (88.9% pruning), reaching 95.61% versus 93.68% for DivPrune and 92.71% for DART at the most aggressive setting.

  • Gains grow as the budget tightens. The authors report that VLM-Pruner ranks first on seven benchmarks at 64 tokens on LLaVA-1.5-7B, with increasingly favorable performance-versus-sparsity tradeoffs over baselines.

  • Larger-model transfer: On LLaVA-1.5-13B at 64 tokens (88.9% pruning), VLM-Pruner attains an average of 92.68%, exceeding DivPrune by +2.48%, DART by +4.56%, and FastV by +7.93%, and holds first place on seven benchmarks.

  • Dynamic-resolution adaptability: On LLaVA-Next-7B with dynamic resolution (max pixels = 1344×336, max tokens 576×4 = 2304), VLM-Pruner reaches an average of 91.60% at 88.9% pruning and ranks first on eight benchmarks.

  • Architectural generality: On Qwen2-VL-7B-Instruct, the average improvement grows with sparsity, reaching +3.65% at 88.9% pruning. On OCRBench (sensitive to small text), VLM-Pruner scores 581 versus 481 for DART, an absolute gain of +12.56%.

  • Video-language results: On LLaVA-Video-7B-Qwen2, retaining 20 tokens per frame out of 13×14 = 182 (88.9% pruning), VLM-Pruner reaches an overall average of 90.55% versus 89.84% for the second best (+0.71%). NExTQA is 72.93/32.02 (+0.09 EM and +0.15 WUPS over second best), VideoMME P-score is 54.96 (+1.00), and EgoPlan accuracy is 29.58% (+0.68%).

  • Efficiency: On LLaVA-1.5-7B at 64 tokens, VLM-Pruner reports 1.39× speedup on POPE (23:59 versus 33:21) and 1.19× on OCRBench (6:04 versus 7:14), with FLOPs at 22.09%. DivPrune achieves 1.50×/1.26× with 16.89% FLOPs but at significantly poorer performance. On Qwen2-VL-7B, VLM-Pruner achieves the fastest OCRBench speed at 1.60× and the best performance in that table.

  • Structural ablation (LLaVA-1.5-7B, 64 tokens, full model at 95.30%): Replacing max-min diverse pivot initialization with Top-4 L1-distance keys drops the average by −1.27% (to 94.03%). Substituting Stage 2 with DART reduces it to 90.46% and with DivPrune to 46.47%. Removing the normalized nearest-distance term drops it by −1.11% (to 94.19%). Dropping Stage 3 lowers the average from 95.30% to 95.07%.

  • Hyperparameter sensitivity: Four pivots are most robust; variance-based channel selection is best around q = 256; τ(0) = 0.8 performs best; larger batch size B reduces latency while performance remains stable, with B = 16 offering a good trade-off.

  • Additional model check (Supplementary): On Qwen3-VL-4B retaining about 120 tokens (88.9% pruning), VLM-Pruner reaches 89.53% average, outperforming the second-best SAINT (87.38%) by 2.15%; BTP reaches 62.06% and CDPruner 84.31%.

Methodology in Plain English

VLM-Pruner operates on the second layer of the LLM decoder and keeps a fixed number of visual tokens, selected in three stages.

Stage 1 — Pivot initialization. A small set of "pivot" tokens (κ) is chosen to coarsely represent distinct subjects. The first pivot maximizes its L1 norm among key vectors; each subsequent pivot maximizes the minimum L2 distance to already-chosen pivots (a max-min rule), so pivots are widely separated across semantic regions.

Stage 2 — Greedy selection with BSS. To avoid both redundancy and scattered coverage, the method first screens channels by variance and keeps the top q = 256 channels, then builds a cosine-similarity matrix over the reduced hidden states. For each candidate token, the nearest spatial distance to the selected set is normalized by D_max (the diagonal of the token grid, e.g. 24√2 for H = W = 24, N = 576). That normalized distance scales the candidate's similarity with a strength λ = 0.5, so spatially far candidates appear more redundant and are deferred. Candidates are sorted by a non-duplication score r_i = 1 − max similarity, processed in parallel batches (B = 16), and accepted when their score falls below a threshold τ that starts at 0.8 and increases by 0.1 each loop — a strict-to-loose schedule that prevents early admission of isolated, far-away tokens.

Stage 3 — Recovery via SWA. Retained tokens lose the outermost information. Each discarded token is assigned to its most similar retained token, forming clusters; within a cluster, weights are similarity-normalized and the discarded hidden states are aggregated into the retained token. With β = 0.3, the final representation is 0.3 × retained hidden state + 0.7 × aggregated hidden state.

The paper notes that when λ = 0 the objective reduces to a standard monotone submodular function with the classical greedy approximation guarantee; the λ > 0 case retains the submodular structure because the normalized distance term decays monotonically as the selected set grows. Experiments used 8 GPUs with 32 GB VRAM each, batch size 1, and the lmms-evals package for Qwen2-VL and LLaVA-Video.

Why This Matters

Impact on research. The work reframes token pruning as an explicit trade-off between redundancy reduction and spatial coverage rather than treating importance scoring alone as the objective. It shows that the failure modes of the two dominant pruning families are complementary, and that a spatial "centrifugal" prior can bridge them. Because the method is training-free, it can be evaluated on top of existing VLMs without retraining, and the reported margins widen with the pruning rate, which is the regime where prior methods degrade most.

Real-world applications:

  • On-device or edge multimodal assistants, where the abstract identifies mobile deployment as the motivating constraint.
  • OCR-heavy document and scene-text workflows, where the paper reports the largest gains on OCRBench (581 versus 481 for DART on Qwen2-VL-7B).
  • High-resolution image understanding pipelines, since VLM-Pruner's advantage grows on the dynamic-resolution LLaVA-Next and on Qwen2-VL, which the authors describe as handling high-resolution inputs better.
  • Long-video question answering and planning, where retaining 20 tokens per frame out of 182 keeps accuracy within a small margin of the unpruned model.

Industry relevance. The work is co-authored by researchers at Huawei Technologies and Zhejiang University, and the speedups reported (1.19×–1.60× depending on model and benchmark) alongside FLOPs reductions to roughly 19–22% bear directly on the serving cost of multimodal inference.

Future Directions

  • The paper does not report results for CDPruner, BTP, and SAINT on the main image benchmarks in the truncated content beyond Qwen3-VL-4B; broader comparison across all five base VLMs is left open here.
  • The ablation shows Stage 3 (SWA) yields only a modest gain (95.30% to 95.07% without it); the authors do not report how far that gain could be pushed with other aggregation designs.
  • The reconstruction target is a fixed token budget; whether the pivot count, channel count, and threshold schedule can be made input-adaptive rather than tuned per model is not reported.
  • Applying BSS beyond the second decoder layer, or combining it with learned/training-required pruning, is not explored in the reported content.

Target Audience

Researchers and engineers working on multimodal model efficiency — particularly those optimizing VLM inference latency, memory, or FLOPs for deployment. It is also relevant to practitioners who need training-free compression methods they can apply to existing checkpoints, and to readers interested in submodular/greedy selection formulations for token or feature reduction. Readers should be comfortable with attention mechanics, cosine similarity, and approximate greedy optimization; the paper is not an introductory treatment of VLMs.

Authors’ abstract

Vision-language models (VLMs) excel at image understanding tasks, but the large number of visual tokens imposes significant computational costs, hindering deployment on mobile devices. Many pruning methods rely solely on token importance and thus overlook inter-token redundancy, retaining numerous duplicated tokens and wasting capacity. Although some redundancy-aware approaches have been proposed, they often ignore the spatial relationships among visual tokens. This can lead to overly sparse selections of retained tokens that fail to adequately cover the regions of target objects. To address these limitations, we propose VLM-Pruner, a training-free token pruning algorithm that explicitly balances redundancy and spatial sparsity. We introduce a centrifugal token pruning paradigm that enables near-to-far selection while prioritizing the preservation of fine-grained object details. Moreover, we design a Buffering for Spatial Sparsity (BSS) criterion that defers the selection of spatially distant tokens. We further adopt a parallel greedy strategy to conduct token selection efficiently. To mitigate information loss from pruning, we selectively fuse salient information from the discarded tokens into the retained ones. Comprehensive comparisons demonstrate that VLM-Pruner consistently outperforms strong baselines across five VLMs with an 88.9\% pruning rate, while delivering an end-to-end inference speedup. The code is available at https://github.com/Casey-bit/VLMPruner.

Read the original paper