Skip to content
AI.info

Research

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

Overview Research area: Computer vision and multimodal AI — specifically visual token pruning for vision-language models (VLMs) that reason about 3D scenes from multi-view images. Technical level: Int

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
arXiv
2609.08345
Published
2026-09-08
Authors
Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye, Dilin Wang, JQ Huang, Rakesh Ranjan, Aviral Chharia, Fernando De la Torre

AI summary

Overview

Research area: Computer vision and multimodal AI — specifically visual token pruning for vision-language models (VLMs) that reason about 3D scenes from multi-view images.

Technical level: Intermediate. The paper is readable without deep 3D-vision background, but it assumes familiarity with transformer-based VLMs, visual tokens, attention, and KV caching.

Scope in one sentence: This paper introduces CoVeR, a training-free, geometry-only token selector that keeps an exact per-scene token budget while maximizing spatial coverage of a multi-view 3D scene, outperforming both voxelization-based and learned-importance-based pruning methods across three 3D reasoning benchmarks and four VLMs.

What This Paper Is About

Feeding a 3D scene to a 2D vision-language model as multiple camera views lets the model reuse internet-scale 2D pre-training priors instead of relying on scarce 3D-language data — but it produces thousands of visual tokens whose count scales with the number of views (12 views yield 8,748 visual tokens in LLaVA-OneVision-7B). Existing token pruners either rank tokens by learned importance (attention or encoder features), which keeps near-duplicate tokens from a few prominent regions, or pool tokens into 3D voxels, which improves coverage but cannot hit an exact token budget and saturates as views overlap. CoVeR's goal is to select tokens that collectively cover every region of the scene, using only token coordinates, while guaranteeing exactly the requested number of tokens for each individual scene.

Key Contributions

  1. Problem analysis of both pruning families. The paper identifies why voxelization-based methods cannot enforce exact per-scene budgets and are capped by voxel saturation, and why learned-importance methods spend their budget on near-duplicates while leaving much of the scene unrepresented.

  2. A coverage-based pruning paradigm (CoVeR). A training-free, deterministic, geometry-only framework that optimizes scene coverage under a guaranteed exact per-scene budget, combining adaptive voxelization (coverage initialization) with farthest point sampling (coverage expansion).

  3. Analytical insights linking coverage to 3D reasoning. Statistical metrics (Nearest Neighbor Index, Nearest Neighbor Distance at the 95th and 100th percentiles) connect geometric coverage to 3D reasoning performance; directed distances (Token Recovery and Token Expansion) show CoVeR preserves the regions favored by learned-importance methods while also covering additional ones; stage-wise ablations show original encoder tokens beat merged features and that pure spatial distance outperforms learned signals.

  4. Extensive evaluation. State-of-the-art results on ScanQA (spatial scene understanding), SQA3D (situated reasoning), and OpenEQA (embodied question answering), with plug-and-play generalization tested across four VLMs, plus efficiency measurements in TFLOPs, KV cache, GPU memory, and latency.

Main Findings

  • Best aggregate performance at every token budget. Averaging across the three benchmarks, CoVeR retains 93.5% of full-token performance at 8% token retention, versus 89.6% for SeGPruner and 85.9% for VisPruner. This corresponds to surpassing SOTA by 3.9 percentage points on average across benchmarks.

  • Strong absolute numbers on ScanQA. At 23% retention CoVeR reaches 28.5 EM@1 and 85.5 CIDEr, improving over the full-token baseline (28.2 EM@1, 83.6 CIDEr). At 9% budget it achieves 27.1 EM@1, 81.4 CIDEr, and 41.5 ROUGE-L.

  • SQA3D and OpenEQA gains. CoVeR reaches 48.6 EM@1 on SQA3D at 8% retention and outperforms prior pruning methods on OpenEQA at aggressive budgets.

  • Exact budget is a real gap in prior work. At a fixed voxel size of 0.2 m and a budget of 1342 tokens, 56% of scenes fall below budget and 44% exceed it. Voxel-only pruning keeps about 69% of tokens on average and only 46–57% in the most redundant scenes, since roughly 31% of ScanQA and SQA3D tokens spatially overlap. The plateau appears near v_s = 0.02 m.

  • Coverage metrics confirm the mechanism. At the most aggressive ScanQA budget, SeGPruner has NNI = 0.458 versus CoVeR's NNI = 0.924 (lower means clustered). CoVeR also wins on NND₉₅ (0.980 vs. 0.967) and NND₁₀₀ (0.977 vs. 0.917), with the largest gap on the worst case.

  • CoVeR preserves salient regions while adding coverage. Against SeGPruner at 9% retention, Token Recovery is 0.009 and Token Expansion is 0.020 (of the scene diagonal). Every SeGPruner token lies within 3.1% of the scene diagonal of a CoVeR token, while roughly 20% of CoVeR tokens are farther than that from any SeGPruner token.

  • Large efficiency gains. Relative to LLaVA-OV-7B on ScanQA at 14% retention, CoVeR cuts LLM TFLOPs by 8.6× and KV cache by 7×, with 1.4× lower GPU memory and a 2.5× inference speedup at a 1.1% relative drop. At 9% retention it achieves 13.3× fewer TFLOPs, 10.7× smaller KV cache, and a 2.9× speedup for a 1.1-point drop.

  • Highest accuracy at lowest peak memory among pruners. At 9% retention on ScanQA with LLaVA-OV-7B, CoVeR scores 27.1 EM@1 / 81.4 CIDEr / 41.5 ROUGE-L at 17.2 GB peak memory, versus SeGPruner at 24.5 / 71.2 / 37.0 at 22.1 GB and VisPruner at 23.4 / 66.9 / 35.6 at 22.1 GB. CoVeR's selection cost is 0.034 s, versus 0.008 s for SeGPruner and 0.010 s for VisPruner.

  • Generalization without retraining. On Video-3D LLM with 16 views at 10% budget, CoVeR obtains 26.5 ScanQA EM@1 versus 26.0 for Geo3DPruner, and retains 93.5% of full performance versus 90.7%, even though Geo3DPruner adds a VGGT-1B encoder and fully retrains the backbone. CoVeR transfers to Qwen2.5-VL-7B and Qwen3-VL-8B, retaining over 96% of ScanQA performance at retention levels above 20% and over 95% on SQA3D until retention falls below 20%.

  • Keeping original tokens beats merging. With stage 1 only (α = 1), pruning improves EM@1 by 7.7 on ScanQA and 18.1 on SQA3D at the tightest budget, with the largest SQA3D gains on Can (+37) and Which (+29).

  • Pure spatial distance beats semantic distance. With stage 2 only (α = 0), pure 3D distance beats spatial+semantic FPS by 1.7 / 2.1 on ScanQA/SQA3D and semantic-only FPS by 2.9 / 4.3, with largest SQA3D gains on How (+6.8) and What (+6.2).

  • Both stages matter. At 9% retention, replacing iterative FPS with Top-K or random selection loses 0.7 and 1.7 average ScanQA points; removing the stage-1 seed loses 0.5 points at 9%. Voxel seeding also reduces pruning time by roughly 1.5× (0.034 s vs. 0.049 s at 9% retention).

  • Robust hyperparameter and view scaling. Varying α from 0.1 to 0.9 changes accuracy within a narrow band; the paper uses α = 0.4. CoVeR leads VisPruner and SeGPruner at every view count and gains more from added views, while those baselines saturate.

Methodology in Plain English

The method starts from a formal statement of the goal. Given a token budget B, the authors want B tokens whose 3D world coordinates represent the whole observed scene. They measure coverage with the directed Hausdorff distance: the distance from the worst-covered original token to its nearest retained token. Minimizing this is equivalent to the discrete Euclidean k-center problem. A small value certifies that no region of the scene is dropped outright — a guarantee learned-importance ranking does not provide.

The pipeline has two complementary stages, both using only 3D token coordinates and no attention, visual features, or auxiliary encoders:

  1. Coverage initialization. Because voxel occupancy depends on scene geometry, extent, and view overlap, CoVeR searches for a voxel size per scene rather than using one global value. It performs a heuristic binary search over the interval [0.02 m, 5.0 m] for up to T = 16 iterations, stopping when the occupied-voxel count falls within a tolerance τ = 0.05 of the stage-1 target B_init = max(1, ⌊αB⌋). From each occupied voxel it keeps the single original token closest to the mean of the other tokens in that voxel — a representative near the voxel's geometric center. This removes cross-view duplicates in parallel and produces a coarse scene-wide cover, but its count is capped by how many voxels are occupied.

  2. Coverage expansion. To get past the saturation ceiling, CoVeR applies farthest point sampling starting from the stage-1 selection. At each step it adds the unselected token whose squared Euclidean 3D distance to its nearest already-selected token is largest, then updates the minimum distances. It repeats this B_expan = B − |C_init| times, so the final selection always has exactly B tokens. Stage 2 is only needed when the voxel count falls short of B; if the voxel search overshoots, a safeguard keeps representatives of the B most populated voxels. The authors state that this safeguard is not activated at any evaluated scene or budget with α = 0.4.

When the safeguard is inactive, every unselected token shares a voxel of side v_s with a selected token, giving a bound of √3 · v_s on the directed Hausdorff distance for stage 1, and each FPS step is non-increasing in that distance.

Integration: CoVeR is a plug-in module. It picks token indices right after the visual encoder; all tokens pass through the projector, and only the selected features reach the LLM in their original sequence order.

Evaluation setup: All three forms of 3D reasoning are tested — ScanQA (4,306 questions, 71 scenes), SQA3D (3,519 questions, 67 scenes), and OpenEQA (1,636 questions, 152 scenes, split into 1,079 ScanNet questions / 89 scenes and 557 HM3D questions / 63 scenes). Metrics are EM@1, CIDEr, and ROUGE-L for ScanQA; EM@1 for SQA3D; and LLM-Match for OpenEQA. Four VLMs are used: LLaVA-OV-7B, Video-3D LLM, Qwen2.5-VL-7B, and Qwen3-VL-8B. Following prior work, 12 views are uniformly sampled, and comparisons are made against VTC, DTC, VisPruner, and SeGPruner under matched inputs; for Geo3DPruner the paper follows its protocol with Video-3D LLM using 16 views. Baseline importance ratios are set to 0.5. Everything runs on one NVIDIA H100 GPU.

Why This Matters

The paper reframes multi-view 3D token pruning as a coverage problem rather than an importance-ranking problem, and shows that this shift is measurable: better coverage metrics track with better downstream 3D answering accuracy. It also exposes a practical weakness of voxel-based methods that is easy to overlook — dataset-average retention is not the same as a per-scene memory or latency guarantee, and a single scene can blow past a fixed limit even when the average looks fine. CoVeR's deterministic, training-free, model-agnostic design means it can be bolted onto different VLMs without architectural changes or retraining, which lowers the cost of scaling 3D reasoning on top of existing 2D models.

Real-world applications:

  • Mixed-reality and smart-glasses assistants that need to answer questions about the physical space around a user from headset cameras, where memory and latency budgets are tight.
  • Embodied agents and service robots performing open-vocabulary question answering about an indoor environment, since OpenEQA is explicitly an embodied-QA benchmark.
  • 3D scene understanding pipelines for home and industrial inspection, where questions concern object recognition, attributes, counting, localization, and spatial relations across many views.
  • Accessibility tools that describe a scene to a user from multi-camera input, where the cost of running a large VLM per query must be reduced.

Industry relevance: the affiliations (Carnegie Mellon University and Meta Reality Labs) point directly at head-mounted and embodied AI hardware, where the paper's headline efficiency numbers — 13.3× fewer TFLOPs, 10.7× smaller KV cache, and 2.9× faster inference at 9% retention — map onto the constraints of on-device or latency-sensitive deployment.

Future Directions

  • Reducing dependence on geometry quality. The authors note in their limitations that CoVeR, like all prior methods, requires depth and camera information and is designed for indoor scenes, so its performance may depend on the quality of estimated geometry. They propose combining coverage with reliable depth and pose estimation.
  • Extending beyond indoor scenes. The same limitations section calls for hierarchical or streaming selection to handle outdoor scenes.
  • Characterizing when coverage and accuracy decouple. The paper reports that across all ablation variants higher coverage (NNI, NND₉₅, NND₁₀₀) accompanies higher accuracy, but the exact conditions under which this relationship breaks down are not established.
  • Broadening the benchmark and backbone matrix. The paper notes that official code for VTC/DTC and Geo3DPruner was not publicly available, so those comparisons followed reported protocols rather than same-environment reproduction; and category-level OpenEQA results are deferred to the Appendix, which is truncated in the provided content. Extending same-environment comparisons to more backbones and task categories is a natural next step.

Target Audience

This paper is most useful to researchers and engineers working on efficient multimodal inference, visual token reduction, and 3D or embodied vision-language systems — particularly those deploying VLMs on hardware with tight memory, latency, or KV-cache budgets. It is also relevant to practitioners building multi-view or MR/AR perception pipelines, and to students looking for a clear worked example of how replacing a learned-signal heuristic with a geometric objective can yield gains that are both measurable and interpretable.

Authors’ abstract

Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods improve spatial coverage but cannot enforce an exact token budget and saturate as multi-view observations overlap in 3D, capping retention well below the target. We show that spatial coverage is associated with 3D reasoning performance and introduce CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals. CoVeR selects tokens that collectively cover every region of the scene, and solves the limitations of both families: it enforces an exact per-scene budget, breaks the voxelization saturation plateau, and avoids the near-duplicate selections of learned importance. Extensive experiments show CoVeR outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs. Notably, with only $\approx$8% of visual tokens, it preserves 93.5% of full-token performance, surpassing SOTA by 3.9 percentage points on average across benchmarks.

Read the original paper