Skip to content
AI.info

Research

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

Overview Research area: Efficient long-video understanding with vision-language models (VLMs) — specifically the allocation of a fixed visual-token budget across frame count, per-frame resolution, and

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency
arXiv
2610.04318
Published
2026-10-03
Authors
Sixun Dong, Wei Li, Andong Deng, Qi Qian, Victor Zhu, Zhengping Ji, Chen Chen

AI summary

Overview

Research area: Efficient long-video understanding with vision-language models (VLMs) — specifically the allocation of a fixed visual-token budget across frame count, per-frame resolution, and front-end video decoding latency.

Technical level: Intermediate. The paper assumes familiarity with VLM input pipelines (video vs. image pathways, visual tokens, mRoPE positional encoding), video codec concepts (I-frames, P/B-frames, Groups of Pictures), and determinantal point processes (DPPs).

Scope: An empirical study of the frame–resolution trade-off across multiple VLMs and long-video benchmarks, plus a training-free two-stream framework (LoHi) that combines dense low-resolution video with sparse high-resolution images.

What This Paper Is About

Existing efficient long-video VLMs treat the problem as selecting informative tokens — which frames or which visual tokens to keep — while holding per-frame resolution at its native scale. This leaves two axes unused: resolution can itself be traded for more frames, and front-end decoding latency grows with video length rather than with the token budget. The paper asks how to jointly allocate an end-to-end budget across frame count, per-frame resolution, and front-end latency, and proposes a framework that answers this by pairing a dense low-resolution video stream with a sparse high-resolution image stream.

Key Contributions

  1. A re-framing of long-video VLM efficiency as a joint allocation problem over frame count, per-frame resolution, and front-end latency, supported by three empirical lessons distilled across multiple VLMs and long-video benchmarks.
  2. LoHi, a training-free framework that composes a dense low-resolution video stream (Lo-V) with a sparse high-resolution image stream (Hi-I), routed through the VLM's native video and image pathways and linked by an inline natural-language marker, with no fine-tuning or architectural changes.
  3. Plug-and-play Hi-I selectors offering different cost–accuracy trade-offs: LoHi-Uniform (zero extra overhead), LoHi-Anchor (codec I-frame metadata, no CLIP forward pass, no question), and LoHi-SemDiv (query-relevance plus visual-diversity DPP over CLIP features).
  4. An entropy-triggered adaptive variant that runs the Lo-V base first and invokes the Hi-I stream only when the answer-distribution entropy exceeds a threshold θ.

Main Findings

  • Lesson 1 — Dense low-resolution sampling is the new recipe. At a matched token budget (about 5,760 tokens), trading per-frame resolution for denser temporal coverage outperforms the sparse native-resolution recipe. For Qwen3-VL-4B, (256, 0.25) improves over the default (16, 1.0) baseline by +6.66% on VideoMME, +10.47% on MLVU, and +5.10% on LVBench, with the same trend on Qwen3-VL-8B and on VideoLLaMA3-7B for MLVU and LVBench.
  • Frame density gains are monotone at fixed low resolution. With r fixed at 0.25, VideoMME accuracy increases as N grows from 16 to 256 for every task category. The gains are largest when the default recipe critically under-samples the timeline and smaller when 16 frames already provide sufficient coverage.
  • Lesson 2 — Optimal frame–resolution allocation is task-dependent. At the same roughly 5,760-token budget, (64, 0.5) and (256, 0.25) swap dominance between resolution-sensitive tasks (OCR and Attribute Perception, which drop most from r = 1.0 to r = 0.25 at N = 16) and the remaining categories. No single (N, r) pair optimally serves all task types.
  • Lesson 3 — Front-end latency cannot be overlooked. Decode-then-select pipelines that build a dense candidate pool at FPS = 1 see latency scale linearly with duration: the gap reaches 7.4x on 60-minute videos and remains substantial on VideoMME itself. Profiled decode latency rises from 8.5 s (10 min, dense) to 25.3 s (30 min) to 53.3 s (60 min), while uniform sampling at N = 256 stays at 4.7 s, 6.5 s, and 7.2 s respectively.
  • Accuracy at matched budget. On Qwen3-VL-4B, LoHi-SemDiv reaches 66.81% / 74.38% / 45.90% on VideoMME / MLVU / LVBench, an average gain of +10.60% over the vanilla baseline (16F, 1.0) and +5.23% over the strongest prior baseline (BOLT) at identical token cost. The abstract states these gains as +10.6% and +5.2% respectively.
  • Lo-V alone is a strong baseline. The dense low-resolution stream (256F@0.25) already outperforms every prior keyframe-selection and token-pruning baseline on average, supporting Lesson 1 without any selection logic.
  • Semantic diversity helps. LoHi-SemDiv > LoHi-Anchor > LoHi-Uniform on every benchmark, with the largest semantic-diversity contribution on MLVU (+6.76% over LoHi-Uniform).
  • LoHi resolves the resolution-sensitivity trade-off. Under a single configuration, 64F@0.5 wins on R (72.73 vs. 71.89) while 256F@0.25 wins on ¬R (61.76 vs. 58.54). All three LoHi variants outperform both single-configuration baselines on both subsets simultaneously (SemDiv: 77.06% on R, 63.12% on ¬R, 66.81% overall).
  • Efficiency gains. LoHi decodes 128 frames versus 256 for every prior efficiency baseline, cutting front-end latency from 5.4 s to 3.3 s. It reduces ViT latency by approximately 12x (from 2,486 ms to 195 ms) and lowers TTFT from 5,963 ms to 3,906 ms, at a measured GPU memory footprint of 9.6 GB. Token pruning methods show a severe Vision Tower bottleneck (464.3 TFLOPs; 2,486 ms).
  • Token pruning falls short of naive resizing. At 25% of the default budget (about 1,440 tokens), resizing to 16F@0.5 (49.91% average) beats VisionZip (47.93%), MMTok (48.27%), and FlashVid (48.36%). At 50% (about 2,880 tokens), resizing to 128F@0.25 (56.20%) beats VisionZip (49.65%), MMTok (52.89%), and FlashVid (51.54%).
  • Lo-V and Hi-I are strictly complementary. At 5,760 tokens, 256 Lo-V frames alone give 64.44% and Hi-I K=16 (SemDiv) alone gives 63.74%, while combining 128 Lo-V frames with 8 Hi-I images gives 66.81%.
  • Robustness across selectors and budgets. All 12 configurations of K ∈ {2, 4, 8} across four training-free selectors achieve 65.22–66.81% on VideoMME, exceeding the 128-frame Lo-V-only baseline of 63.00%.
  • Adaptive triggering. The adaptive variant nearly matches full LoHi performance while invoking the high-resolution stream on only one-third of the queries.
  • Scaling to a larger model. On EgoLongQA with Qwen3.5-27B, LoHi matches the accuracy of a higher-resolution single-stream baseline using 29% fewer visual tokens.

Methodology in Plain English

The authors begin with a controlled empirical study rather than a new architecture. They define a fixed visual-token budget B and a uniform sampling configuration (N, r), where N is the frame count and r is the per-frame resolution scale relative to native, constrained by N · τ(r) ≤ B. They then sweep this two-dimensional space across VLMs (Qwen3-VL-4B/8B, VideoLLaMA3-7B) and benchmarks to see where accuracy actually comes from.

From that study they derive a decomposition: instead of one configuration, split the budget into a dense low-resolution video stream of N frames at scale rℓ (temporal coverage) and a sparse high-resolution stream of K ≪ N frames at scale rh > rℓ (spatial detail). The K high-resolution indices are drawn as a subset of the low-resolution grid, so each frame is decoded once at native resolution and reused at both scales — meaning the number of decoded frames depends only on N, not on video length.

For selecting which frames get the high-resolution treatment, they offer three options. LoHi-Uniform picks evenly spaced indices at zero overhead. LoHi-Anchor reads pre-computed codec I-frame indices (via decord.VideoReader.get_key_indices(), no temporal-distance threshold) and maps each back to the nearest grid index, falling back to uniform if keyframe retrieval fails. LoHi-SemDiv builds a quality–similarity DPP kernel L = diag(q) · E Eᵀ · diag(q) from L2-normalized CLIP frame embeddings and min-max-normalized, power-sharpened query–frame cosine similarities (α = 5), then greedily maximizes the regularized log-determinant log det(L_S + I), which is submodular and gives a (1 − 1/e) approximation in O(NK²) time.

The two streams are fed through the VLM's native video pathway (3D temporal-spatial mRoPE) and image pathway (independent 2D positional encoding), bridged by a text marker telling the model that K high-resolution images show selected key frames. An optional variant computes answer entropy from the Lo-V base and only triggers Hi-I when it exceeds θ.

Evaluation uses VideoMME (without subtitles), MLVU, and LVBench, with a primary backbone of Qwen3-VL-4B and generalization checks on Qwen3-VL-8B, Qwen3.5-4B, VideoLLaMA3-7B, VideoChat3-4B, and Qwen2.5-VL-7B. Baselines include keyframe selection (AKS, BOLT, CLIP-Topk) and token pruning (FlashVid, VisionZip) at a pruning ratio of 93.75%. LoHi uses N = 128 at rℓ = 0.25 with K = 8 Hi-I frames at rh = 1.0, giving N_dec = 128 — half the baselines' 256-frame pool at the same token budget.

Why This Matters

Impact on research. The paper challenges the default assumption that efficiency in long-video VLMs is a token-selection problem at fixed native resolution. It shows that resolution itself is a free axis, that decode-then-select pipelines pay a latency cost that scales with video length, and that token pruning — which requires costly native-resolution encoding before discarding tokens — can be beaten by simply resizing frames. It also argues that evaluation should report decode budget alongside visual-token budget.

Real-world applications.

  • Hour-scale video triage and search over surveillance or body-camera footage, where front-end decode latency rather than model inference dominates wall-clock time.
  • Assistive or agentic systems that need to answer many queries over the same long recording and can use entropy-triggered early exit to skip high-resolution processing on easy questions.
  • OCR-heavy and attribute-perception workflows (document screens, dashboards, scene signage) where moderate downscaling fails and a small number of high-resolution frames is needed.
  • Deployment on memory-constrained accelerators, since LoHi reports a 9.6 GB GPU footprint with a 12x ViT latency reduction.

Industry relevance. The framework requires no fine-tuning or architectural changes — it plugs into existing unified video-image VLMs such as the Qwen-VL series — and its selectors range from zero-overhead (uniform, codec metadata) to one CLIP forward pass over low-resolution frames, making the cost–accuracy trade-off explicitly tunable. Affiliations include Meta Reality Labs and Axon, suggesting direct relevance to AR/VR and public-safety video workloads.

Future Directions

  1. Richer Hi-I selectors and adaptive allocation, which the conclusion names explicitly as future work beyond the uniform, anchor, and SemDiv options presented.
  2. Whether the joint-allocation recipe generalizes across all task types — the paper shows no single (N, r) is optimal for both resolution-sensitive and insensitive subsets, so principled per-query allocation remains open.
  3. Extending the decomposition beyond the tested backbones and benchmarks, since primary results use Qwen3-VL-4B with generalization runs on six other models, and EgoLongQA is evaluated only with Qwen3.5-27B.
  4. Refining the adaptive trigger, as the entropy-based early exit invokes Hi-I on one-third of queries; the threshold θ and its calibration across tasks are not characterized in detail in the provided content.

Target Audience

Researchers and engineers working on efficient video-language models, long-video understanding, or multimodal inference serving. It is most useful to readers who already understand VLM token budgets and video decoding pipelines, and who want a concrete, training-free recipe with measured latency and memory numbers. Practitioners deploying long-video VLMs on latency- or memory-constrained hardware will find the efficiency profiling in Table 5 directly actionable.

Authors’ abstract

Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding latency depends on the size of the candidate pool rather than the final token budget. An empirical study across multiple VLMs and long-video benchmarks yields three findings: dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; resolution-sensitive tasks benefit from selected high-resolution frames; and front-end decoding dominates wall time for hour-long videos. Motivated by these findings, we introduce LoHi, a training-free, single-pass framework that combines a dense low-resolution video stream with sparse high-resolution image frames through the VLM's native video and image pathways. LoHi-Anchor selects high-resolution frames using codec I-frame metadata, while LoHi-SemDiv uses query relevance and visual diversity over CLIP features. Across three long-video benchmarks, LoHi improves average accuracy by 10.6 percentage points over the native-resolution baseline at a matched token budget and by 5.2 percentage points over the strongest prior efficiency method. It also reduces front-end decoding latency by up to 7x on hour-long videos. Project page: https://sixundong.com/projects/lohi

Read the original paper