Skip to content
AI.info

Research

Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers

Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers Overview Research area: Cross-modal video retrieval, specifically list-wise reranking and multi-retriever rank fusion using frozen Vision-Langu

Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers
arXiv
2511.01617
Published
2025-11-03
Authors
Mohamed Eltahir, Ali Habibullah, Lama Ayash, Tanveer Hussain, Naeemullah Khan

AI summary

Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers

Overview

Research area: Cross-modal video retrieval, specifically list-wise reranking and multi-retriever rank fusion using frozen Vision-Language Models.

Technical level: Advanced. The paper assumes familiarity with two-stage retrieval pipelines, dual-encoder retrievers, rank/score fusion methods, and large multimodal model prompting.

Scope: The paper introduces Vote-in-Context (ViC), a training-free framework that serializes candidate content and retriever metadata into a VLM prompt so a frozen model can act as both a single-list reranker and an ensemble fuser, demonstrated on text-to-video and video-to-text retrieval benchmarks (arXiv:2511.01617v1, 03 Nov 2025).

What This Paper Is About

Video retrieval systems typically retrieve a broad candidate pool with a fast dual-encoder, then refine it with a second-stage reranker. When several retrievers are available, their lists are usually merged using fixed formulas such as Reciprocal Rank Fusion (RRF) or CombSUM/CombMNZ, which look only at ranks or scores and ignore what the candidates actually contain. The paper's goal is to replace these content-blind fusion formulas with a zero-shot VLM that sees both the candidates' content and the retrievers' rank/consensus signals, and reasons over them jointly.

Key Contributions

  1. Vote-in-Context (ViC): A generalized, training-free framework that turns a frozen VLM into a list-wise reranker and fuser by serializing both content evidence (images, text) and retriever metadata (per-list ranks, cross-list multiplicity) directly into the prompt.
  2. S-Grid: A compact video representation that composites uniformly sampled frames into a single image grid, optionally paired with subtitles or ASR text, enabling VLM reasoning over full video candidate lists without costly sequence processing.
  3. Dual-mode evaluation: ViC is evaluated as a single-list reranker (M=1), where it improves every backbone tested, and as an ensemble fuser (M>1), where it consistently outperforms RRF, CombSUM, and CombMNZ.
  4. Released framework and protocols: Code and evaluation protocols are publicly released, including analysis of scaling properties, sensitivity to context size, grid size, VLM type, and different assembly strategies.

Main Findings

  • Single-list reranking delivers large gains. On MSR-VTT (t2v), ViC lifts the weakest backbone, CLIP4Clip, by 29.8 points (34.4 to 64.2) and the strongest, InternVideo2, by 21.4 points (54.5 to 75.9). On ActivityNet (t2v), it adds 31.6 R@1 to InternVideo2 (58.2 to 89.8). On VATEX (v2t), it boosts VAST by 22.0 points (77.6 to 99.6).

  • Subtitles help. Comparing "Grid" (visuals only) against "S-Grid" (visuals plus subtitles) shows consistent improvements when textual evidence is included, indicating the VLM uses all available modalities. For example, VAST on MSR-VTT t2v goes from 67.3 (Grid) to 68.7 (S-Grid), and on VATEX v2t from 99.4 to 99.6.

  • Fusion beats traditional baselines. ViC achieves R@1 of 87.1 (t2v) / 88.1 (v2t) on MSR-VTT, 87.4 (t2v) / 84.3 (v2t) on DiDeMo, 96.0 (t2v) / 96.2 (v2t) on ActivityNet, and 97.5 (t2v) on VATEX. On MSR-VTT t2v it surpasses the best baseline (CombMNZ, 85.3) by +1.8 points; on DiDeMo t2v it exceeds the next-best baseline (CombSUM, 80.4) by +7.0 points.

  • Headline zero-shot numbers. The abstract reports Recall@1 of 87.1% (t2v) / 89.0% (v2t) on MSR-VTT and 99.6% (v2t) on VATEX, described as gains of up to +40 Recall@1 over previous state-of-the-art baselines. Table 2 lists the ViC fusion v2t result on MSR-VTT as 88.1.

  • Multiplicity carries real signal. Removing duplicates from the candidate sequence before prompting the VLM drops MSR-VTT t2v from 87.1 to 84.2 (the ViC "No Duplicates" variant), confirming the model uses cross-list consensus as a relevance cue.

  • Fusion and reranking are complementary. Against the best single-backbone reranking result (75.9 on MSR-VTT), single-list ViC reranking adds +21.4 points over the un-reranked InternVideo2 baseline (54.5), and adding fusion contributes a further +11.2 points.

  • Grid size has a sweet spot. Testing 1x1 through 4x4 grids, 2x2 and 3x3 perform best: 1x1 undercovers the video, while 4x4 compresses frames too aggressively and can add redundant visual tokens. The evaluated datasets are short (MSR-VTT clips 10-30 s, DiDeMo about 25-30 s, VATEX around 10 s), while ActivityNet Captions has longer untrimmed videos averaging minutes.

  • Bigger rerankers help, but 8B is the floor. Scaling from 8B to 38B parameters at a fixed 3x3 grid steadily improves R@1 until saturating. The 8B model already performs strongly, whereas models smaller than 8B fail to produce consistent permutations.

  • Context size has diminishing returns. For t2v, moving from K=10 to K=14 raises R@1, but K=30 lowers it while barely moving R@10. For v2t, K=20 is the most effective operating point. Qwen3-VL performs well on t2v but degrades substantially on v2t as context grows; Gemma-3 is an exception, staying stable at K=30 and achieving the highest R@10 in both directions. R@30 was confirmed to be effectively saturated near 100% across benchmarks, meaning the correct item is almost always in the top 30.

  • A new efficiency frontier. ViC pushes average R@1 from roughly 57% to roughly 90%, establishing a dominant Pareto frontier for the time-per-query versus performance trade-off, though at higher latency than RRF or CombSUM. Latency was measured on a single NVIDIA A100 80GB GPU, averaged over 50 queries for a 1k video retrieval task.

Methodology in Plain English

The pipeline has two stages. In stage one, ordinary frozen retrievers (CLIP4Clip, VAST, GRAM, InternVideo2-6B) fetch a shortlist of candidates for each query. In stage two, ViC hands that shortlist to a frozen VLM (mainly InternVL 3.5 38B) and asks it to output a reordered list.

What makes ViC distinctive is what goes into the prompt. For each video, the method builds an "S-Grid": a single image made by tiling uniformly sampled frames in a grid (3x3 by default), optionally with the subtitle or ASR text appended. Because each video becomes one image rather than a frame sequence, the cost per query is proportional to the number of candidates K, not to the raw length of the videos: O(K · C_VLM).

For fusion across M retrievers, each list is truncated to a per-list depth of ceil(K/M), then the lists are interleaved in round-robin order, and duplicates are kept. That ordering and duplication pattern implicitly encodes retriever metadata: position in the list signals each retriever's rank, and an item appearing multiple times signals cross-list consensus. The VLM then weighs retriever agreement against the visual and textual content on a per-query basis, with no fixed weighting formula and no fine-tuning. The single-retriever case (M=1) is just the standard top-K list, so ViC acts as an ordinary list-wise reranker.

The final output is a permutation of the candidate sequence; only the highest-ranked instance of each duplicate candidate counts at evaluation. Default candidate counts are K=14 for t2v and K=20 for v2t.

Why This Matters

Impact on research. The paper reframes rank fusion from a fixed arithmetic formula into a zero-shot reasoning task, showing a frozen VLM can serve as a universal, modality-aware fuser. It also shows that an 8B model already does useful list-wise reranking, which lowers the practical entry barrier for this line of work and points toward lightweight fine-tuned rerankers.

Real-world applications:

  • Video search engines and media archives that need precise top-ranked results under fixed latency budgets.
  • Retrieval-augmented generation over video libraries, where the reranker quality directly affects downstream generated answers.
  • Video-to-text retrieval for accessibility workflows, such as matching a clip to its most relevant caption or description.
  • Content discovery and recommendation systems that aggregate results from several heterogeneous retrieval models.

Industry relevance. Many production search systems already run multiple retrievers and merge their outputs, since ensembling typically gives significant gains. ViC offers an alternative to hand-tuned fusion weights and hyperparameters, at the cost of a full VLM forward pass per query. The paper's efficiency-versus-performance analysis, plotted on a single A100 80GB GPU, gives practitioners a concrete sense of where that trade-off lands.

Future Directions

  1. Prompt engineering and lightweight fine-tuning to enable smaller and more efficient VLMs to perform robustly, since models below 8B failed to produce consistent permutations and the 8B model already performs well untrained.
  2. Query-aware or adaptive keyframe selection to build more representative S-Grids within a fixed token budget, addressing the fact that uniform sampling is lossy and can miss short but semantically important events in long untrimmed videos.
  3. Mitigating context-window and reliability limits, since performance degrades as K grows and fidelity depends on the VLM's instruction-following, with positional bias and output parsing failures observed.
  4. Improving the cost-performance balance through lightweight, fine-tuned rerankers, which the authors suggest could outperform the current Pareto frontier.

Target Audience

Researchers and engineers working on retrieval, reranking, and multimodal systems; practitioners building multi-retriever search pipelines who want a training-free alternative to RRF and CombSUM; and anyone studying how far frozen VLMs can be pushed as zero-shot relevance judges. Readers should be comfortable with standard retrieval metrics and two-stage retrieval terminology.

Available resources: Code and resources are publicly available at https://github.com/mohammad2012191/ViC, and the final VATEX video list used in the re-indexed 1,252-video evaluation is released to facilitate reproducibility.

Authors’ abstract

In the retrieval domain, candidates' fusion from heterogeneous retrievers is a long-standing challenge, particularly for complex, multi-modal data such as videos. While typical fusion techniques are training-free, they rely solely on rank or score signals, disregarding candidates' representations. This work introduces Vote-in-Context (ViC), a generalized, training-free framework that re-thinks list-wise reranking and fusion as a zero-shot reasoning task for a Vision-Language Model (VLM). The core insight is to serialize both content evidence and retriever metadata directly within the VLM's prompt, allowing the model to adaptively weigh retriever consensus against visual-linguistic content. We demonstrate the generality of this framework by applying it to the challenging domain of cross-modal video retrieval. To this end, we introduce the S-Grid, a compact serialization map that represents each video as an image grid, optionally paired with subtitles to enable list-wise reasoning over video candidates. ViC is evaluated both as a single-list reranker, where it dramatically improves the precision of individual retrievers, and as an ensemble fuser, where it consistently outperforms strong baselines like CombSUM. Across video retrieval benchmarks including ActivityNet and VATEX, the framework establishes new state-of-the-art zero-shot retrieval performance, demonstrating its effectiveness in handling complex visual and temporal signals alongside text. In zero-shot settings, ViC achieves Recall@1 scores of 87.1% (t2v) / 89.0% (v2t) on MSR-VTT and 99.6% (v2t) on VATEX, representing massive gains of up to +40 Recall@1 over previous state-of-the-art baselines. We present ViC as a simple, reproducible, and highly effective recipe for turning modern VLMs into powerful zero-shot rerankers and fusers. Code and resources are publicly available at: https://github.com/mohammad2012191/ViC

Read the original paper