Skip to content
AI.info

Research

GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval

Overview Research area: Computer vision and multimodal video-language understanding, specifically zero-shot video moment retrieval (ZVMR) — finding the temporal segment in a video that matches a natur

GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval
arXiv
2601.00584
Published
2026-01-02
Authors
Mingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeoung Kim

AI summary

Overview

  • Research area: Computer vision and multimodal video-language understanding, specifically zero-shot video moment retrieval (ZVMR) — finding the temporal segment in a video that matches a natural language query without task-specific training data.
  • Technical level: Intermediate. The framework itself is training-free and conceptually simple, but it assumes familiarity with vision-language models (VLMs), large language models (LLMs), sentence embeddings, and standard temporal grounding metrics.
  • Scope: The paper diagnoses a "granularity mismatch" between language queries and video captions, proposes a dual-path, training-free alignment framework (GranAlign) to fix it, and evaluates it on QVHighlights, Charades-STA, and ActivityNet-Captions, plus a transfer test on Video Highlight Detection.

What This Paper Is About

Video moment retrieval systems try to pinpoint the exact stretch of video a sentence describes, but in the zero-shot setting they rely on pre-trained models whose language and visual representations are not matched in how coarse or fine they are. A vague query ("a cute dog") covers broadly but is imprecise, while a very specific one ("a baby Golden Retriever is walking around") is precise but often misses the right moment — a trade-off the paper calls the granularity mismatch. GranAlign's goal is to resolve this trade-off by aligning queries and video captions at two complementary levels of semantic granularity, with no training required.

Key Contributions

  1. A diagnosis of the granularity mismatch in ZVMR. The paper argues that prior methods treat queries monolithically, and shows via a breakdown by query type (Error, Simple, Detail, Else) on QVHighlights validation that performance varies substantially across these categories.
  2. Granularity-based query rewriting. Using LLaMA-3 (LLaMA3-8B), the input query is rewritten into a simplified version (rare words replaced with common alternatives, core entities and actions kept, errors corrected) and a detailed version (fine-grained expressions, temporal context, and specific lexical choices preserved), using multiple manually crafted instruction pairs rather than a single rigid prompt.
  3. Query-aware caption generation. Alongside general query-agnostic captions for all frames, the framework uses Qwen2.5-VL-7B to produce focused query-aware captions only for the top-K% of frames most similar to the query, guided by entities and actions extracted from it — a hybrid strategy for computational feasibility.
  4. A training-free, granular moment scoring and proposal pipeline. Simplified queries are paired with query-agnostic captions and detailed queries with query-aware captions; the two similarity scores are averaged into a composite moment score that feeds a Moment Proposal Generator, span scoring with a length regularizer, and NMS-based post-processing.

Main Findings

  • New state of the art on QVHighlights. On the validation split, GranAlign improves on the previous zero-shot state of the art (Moment-GPT) by +3.04% to +3.93%, with the largest gain in mAP@0.5; on the hidden test set it leads by +1.6% to +3.84%. The abstract reports a 3.23% mAP@avg improvement on QVHighlights.
  • Concrete QVHighlights numbers. Validation: 61.94 R1@0.5, 41.81 R1@0.7, 59.63 mAP@0.5, 39.12 mAP@avg. Test: 59.92 R1@0.5, 39.3 R1@0.7, 58.94 mAP@0.5, 38.23 mAP@avg. The previous zero-shot SOTA (Moment-GPT) reports 58.9/38.6/55.7/35.9 on val and 58.3/37.7/55.1/35.0 on test.
  • Transfer to other benchmarks. On Charades-STA, GranAlign exceeds the previous SOTA by +1.2% (R1@0.5) and +1.5% (mIoU), reaching 59.1 R1@0.3, 39.6 R1@0.5, 22.7 R1@0.7, and 38.0 mIoU. On ActivityNet-Captions it gains +2.9% (R1@0.5) and +2.3% (mIoU), reaching 50.3/34.0/16.5 and 33.1 mIoU.
  • Granularity-aligned pairing beats mismatched pairing. In the ablation on QVHighlights validation, mismatched combinations (for example simplified query with query-aware caption) decline, while aligned pairs do better. The full model combining both the (simplified, query-agnostic) and (detailed, query-aware) pairs gives the best results across all metrics (61.94 R1@0.5, 41.81 R1@0.7, 59.63 mAP@0.5, 39.12 mAP@avg).
  • Complementary single paths. The simplified–agnostic pair yields high recall but low precision; the detailed–aware pair achieves the best single-pair performance but is prone to hallucination or misalignment from over-relying on the query.
  • Robustness across caption backbones. Rerunning the ablation with BLIP-2 captions shows the mixed-granularity configuration (simplified query with query-agnostic captions, detailed query with query-aware captions) again achieving the highest average mAP (37.80), with 56.00 mIoU and 59.32 R1@0.5.
  • Transfer to Video Highlight Detection. GranAlign reaches 39.35% mAP and 66.34% HIT@1, outperforming the fully supervised QD-DETR (39.04% mAP, 62.87% HIT@1) and zero-shot Moment-GPT (36.7% mAP, 62.7% HIT@1). A variant using only the detailed-aware path (GranAlign†) reaches 37.41% mAP and 64.98% HIT@1.
  • Efficiency. GranAlign has zero training cost, 6.2s inference time, and 22G GPU memory on QVHighlights, compared with 16.1s and 32G for Moment-GPT (which reports 58.9 R1@0.5 and 35.9 mAP). It also reports the shortest inference time among the compared methods.
  • Prompt robustness. Across five instruction pairs, R1@0.5 scores cluster within a 0.7% range, from a minimum of 61.26% to a peak of 61.94%, indicating the method generalizes the concepts of simplification and detail preservation rather than overfitting to one prompt.
  • Hyperparameter behavior. Performance peaks with 3 rewritten queries (m=3) and λ=0.3, and remains stable across nearby values. Caption selection top-K of 10%, span merging gap τ of 6 frames, a bottom-20% exclusion threshold, and an NMS IoU threshold of 0.9 were the best settings in the reported ablations.
  • Qualitative example. For the query "An Asian girl wearing a white face mask with a heart on it walking on the street," the simple path over-extends the prediction from 78s, the detail path starts more precisely but under-segments after 138s, and GranAlign recovers the ground-truth segment of 90–142s.

Methodology in Plain English

The framework keeps the underlying vision-language models frozen and adds a layer of logic on top of them, so nothing is trained.

First, an LLM rewrites the user's query two ways: a simplified version that strips away rare or incidental wording while keeping the main entities and actions, and a detailed version that keeps every nuance. This is done with several hand-written instruction pairs, and multiple rewritten query pairs are generated for each input.

Second, the video is captioned twice. General, query-agnostic captions are generated for all frames offline. Then the system finds the frames whose CLIP similarity to the query is in the top 10% and asks a VLM to produce query-aware captions for only those frames, using semantic elements (entities and actions) lifted from the query as guidance. This avoids captioning every frame with the query in mind, which would be too slow.

Third, the two aligned pairs — simplified query with the general caption, detailed query with the focused caption — are compared using cosine similarity between sentence embeddings, and the two similarities are averaged into one frame-level score. This average is the key design choice: it lets the broad path contribute recall and the focused path contribute precision, and it prevents either path's failure mode from dominating.

Finally, frames are grouped into candidate spans. Adjacent high-scoring frames within 6 frames of each other are merged into one span, and spans whose average similarity falls in the bottom 20% are discarded. Each candidate span is scored by a weighted combination of its average semantic similarity and its normalized length (λ = 0.3), and overlapping spans are pruned with non-maximum suppression at an IoU threshold of 0.9. The system runs on four NVIDIA A6000 48GB GPUs in the reported setup.

Why This Matters

  • Impact on research: The paper reframes zero-shot VMR as a granularity-alignment problem rather than a representation-quality problem. It suggests that even high-quality pre-trained features can fail if the coarse-to-fine levels of the two modalities do not match, and it provides a training-free recipe that other multimodal alignment tasks could adopt.
  • Real-world applications:
    • Searching long video archives (surveillance, sports footage, lecture recordings) with natural language to jump directly to a described event.
    • Automatic highlight or clip generation for media and streaming platforms from a text description.
    • Video editing and content production workflows, where editors locate a described shot without scrubbing timelines.
    • Assistive and accessibility tools that let users retrieve specific moments from video using free-form spoken or typed descriptions.
  • Industry relevance: The method requires no labeled training data and no fine-tuning, and it reports a shorter inference time and smaller memory footprint than a comparable zero-shot baseline, which makes it attractive where annotation budgets and GPU budgets are both constrained. The fact that it can be swapped to a different captioning backbone (BLIP-2) without losing the granularity benefit supports practical deployment flexibility.

Future Directions

  • Fact-checking generated captions. The paper notes that query-aware captioning can hallucinate visual content that is not present in the video, and proposes a verification mechanism that checks generated details against visual evidence.
  • Semantic verification of rewrites. A verification step could ensure that LLM-based query rewriting preserves the original intent of the user's query rather than drifting away from it.
  • Broader downstream generalization. The authors demonstrate transfer to Video Highlight Detection; extending the granularity-aware design to other related temporal grounding and multimodal understanding tasks is an open direction.
  • Open questions raised by the results. How granularity alignment behaves with captioning or rewriting backbones other than the ones tested, and how the approach scales to longer or more densely annotated videos, are not resolved in the reported experiments.

Target Audience

Researchers and practitioners working on video-language understanding, temporal grounding, and zero-shot retrieval, especially those building systems on frozen VLMs and LLMs. It is also useful for applied engineers who need moment retrieval without labeled training data, and for readers interested in how prompt-level design, rather than model training, can resolve a semantic alignment problem. Readers without background in multimodal embeddings or temporal IoU-based metrics will need to consult the appendix definitions covered in the paper (R1@n, mAP@m, mAP@avg, mIoU, HIT@1).

Authors’ abstract

Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this setting lies in the mismatch in semantic granularity between textual queries and visual content. Previous studies in ZVMR have attempted to achieve alignment by leveraging high-quality pre-trained knowledge that represents video and language in a joint space. However, these approaches failed to balance the semantic granularity between the pre-trained knowledge provided by each modality for a given scene. As a result, despite the high quality of each modality's representations, the mismatch in granularity led to inaccurate retrieval. In this paper, we propose a training-free framework, called Granularity-Aware Alignment (GranAlign), that bridges this gap between coarse and fine semantic representations. Our approach introduces two complementary techniques: granularity-based query rewriting to generate varied semantic granularities, and query-aware caption generation to embed query intent into video content. By pairing multi-level queries with both query-agnostic and query-aware captions, we effectively resolve semantic mismatches. As a result, our method sets a new state-of-the-art across all three major benchmarks (QVHighlights, Charades-STA, ActivityNet-Captions), with a notable 3.23% mAP@avg improvement on the challenging QVHighlights dataset.

Read the original paper