Research
V-Agent: An Interactive Video Search System Using Vision-Language Models
Overview Research area: Computer vision and information retrieval — specifically multimodal (vision + audio + text) video search systems built on vision-language models (VLMs) and multi-agent orchestr

- arXiv
- 2512.16925
- Published
- 2025-11-04
- Authors
- SunYoung Park, Jong-Hyeon Lee, Youngjune Kim, Daegyu Sung, Younghyun Yu, Young-rok Cha, Jeongho Ju
AI summary
Overview
Research area: Computer vision and information retrieval — specifically multimodal (vision + audio + text) video search systems built on vision-language models (VLMs) and multi-agent orchestration.
Technical level: Intermediate. The paper combines several moving parts (VLM fine-tuning, weight arithmetic, agent routing, hybrid score fusion), but each is described in a self-contained way.
One-sentence scope: The paper presents V-Agent, a three-agent platform that turns a fine-tuned VLM into a video-text retrieval model and wraps it in an interactive search and video question-answering pipeline, evaluated on MSR-VTT and the MultiVENT 2.0 benchmark.
What This Paper Is About
Conventional video search — including on YouTube — relies on metadata such as titles, tags, and descriptions rather than analyzing what is actually shown or said in a video, which limits how well it can answer complex or visually specific queries. The authors build a system that instead embeds video frames and ASR-transcribed audio into a shared multimodal space so that visual and spoken content can both be searched, and they add conversational agents so users can ask follow-up questions about retrieved videos. The goal is a working, interactive video search platform rather than a benchmark-only model.
Key Contributions
- A VLM-based video-text retrieval model. Qwen2-VL-7B-Instruct is fully fine-tuned for two epochs on the ShareGPTVideo 17k video preference dataset using InfoNCE loss with in-batch negatives plus one hard negative, then augmented with a "retrieval vector."
- A retrieval-vector technique for limited video data. The vector τ is computed as the weight difference between GME (an image-text retrieval model) and the original Qwen2-VL-7B-Instruct, and added to the fine-tuned model's weights — inspired by the Chat Vector method — to improve vision-text alignment despite scarce video training data.
- A three-agent pipeline. A routing agent (gpt-4.1-mini) decides whether retrieval is needed; a search agent (gpt-4o) retrieves top-k candidates with the retrieval model and reorders them with an LLM reranker (gpt-4o-mini); a chat agent (gpt-4o) grounds answers in retrieved or user-selected videos.
- A multimodal indexing and fusion scheme. Whisper performs ASR, video descriptions are concatenated and translated to English via gpt-4o-mini when needed, 48 frames are extracted per video at uniform intervals, and frame and transcription embeddings are fused with a weighted average (α = 0.5) over inner-product scores, indexed in pgvector with HNSW settings m=16 and ef_construction=200.
Main Findings
- State-of-the-art zero-shot results on MultiVENT 2.0. V-Agent reaches nDCG@10 of 0.680 and R@10 of 0.676, ahead of the prior best MMMORRF (0.586 nDCG@10, 0.611 R@10), CLIP (0.304 / 0.333), LanguageBind (0.324 / 0.355), SigLIP (0.375 / 0.409), VAST (0.116 / 0.118), and InternVideo2-6B (0.005 / 0.004).
- Prompting alone is not enough for retrieval. Qwen2-VL-7B-Instruct records R@1 of 0.002, R@5 of 0.006, and R@10 of 0.010 on MSR-VTT, indicating that an instruction-tuned VLM without retrieval-specific training does not produce useful embeddings.
- The retrieval vector helps on MSR-VTT. The fine-tuned model M_F scores R@1 0.413, R@5 0.661, R@10 0.750; adding τ to form M_R raises these to 0.476, 0.720, and 0.798 — above GME-7B with mean pooling (0.411 / 0.655 / 0.764) and LamRA (0.447 / 0.686 / 0.786), but below InternVideo2-6B (0.559 / 0.783 / 0.851).
- GME inference strategy matters. Mean pooling over frame-level embeddings is competitive (R@1 0.411) while multi-image inference performs poorly (R@1 0.201); the authors note frame-wise extraction costs more latency but consider the trade-off favorable.
- LLM re-ranking is the largest single gain on MultiVENT 2.0. Under the same frame settings, re-ranking improves nDCG@10 by roughly 6 percentage points (0.680 versus 0.614 at 48 frames).
- More frames give only modest returns. Moving from 16 to 32 to 48 frames without re-ranking changes nDCG@10 from 0.607 to 0.611 to 0.614, while R@10 holds at 0.671 for all three.
- Descriptions matter substantially. Removing descriptions at 16 frames drops nDCG@10 from 0.607 to 0.518 and R@10 from 0.671 to 0.587; removing both descriptions and the retrieval vector yields 0.509 / 0.573.
- InternVideo2's ranking flips between benchmarks. It is the strongest tested baseline on MSR-VTT but weakest on MultiVENT 2.0; the authors cite Samuel et al. (2025) suggesting dependence on low-quality captions and query complexity, and argue multilingual VLM training helps query understanding.
Methodology in Plain English
The authors start from an off-the-shelf vision-language model and turn it into something that can judge whether a video matches a text query. They do this by feeding it prompts paired with videos, along with a good answer and a rejected answer, and training it to score the good answer higher. Because video training data is scarce, they borrow alignment knowledge from an image-text retrieval model: they subtract the weights of the original base model from the specialized image-retrieval model to get a "retrieval vector," then add that vector to their fine-tuned model. The arithmetic is simple — τ = θ_GME − θ_Qwen, and θ_M_R = θ_M_F + τ — and it transfers image-level retrieval skill into the video model.
For the actual search pipeline, each video is prepared once: its audio is transcribed with Whisper, any existing description is concatenated to the transcript, non-English text is translated to English, and 48 frames are sampled at even intervals. Both the frames and the transcript are embedded by the same model and stored in a vector database. At query time, the query is embedded the same way, frame scores and transcript scores are combined with a weighted average, and the top 10 videos are reordered by a language model given a reranking prompt.
Around this sits the agent layer. A routing agent inspects the user's message and decides whether it needs video retrieval; if yes, the search agent runs retrieval and reranking and hands results to the chat agent; if no, the chat agent answers directly. Once results are shown, the user can pick one or more videos and ask questions, and the chat agent answers using those videos plus the question as context. The user-facing interface is built in Flutter, and the agents are orchestrated through the OpenAI Agents SDK.
Why This Matters
The paper argues that most video-search research optimizes models and benchmark scores without embedding them into a usable system, and that most deployed search (YouTube included) still ignores actual video content. V-Agent closes that loop by connecting a retrieval model to indexing, scoring fusion, LLM reranking, and a conversational layer in one pipeline, and by demonstrating that a single fine-tuned VLM embedding both modalities into one shared space can beat a system (MMMORRF) that uses separate SigLIP and multilingual ColBERT-X embeddings.
Real-world applications implied by the system:
- Searching large video archives by visual content rather than titles and tags.
- Multilingual event-centric retrieval, where queries and footage span Arabic, Chinese, English, Korean, Russian, and Spanish.
- Question answering across multiple videos at once, such as comparing coverage of the same event.
- Video browsing with automatically generated summaries of retrieved clips.
Industry relevance: The system is built almost entirely from existing components and APIs — Qwen2-VL-7B-Instruct, GME, Whisper, gpt-4o/gpt-4o-mini/gpt-4.1-mini, pgvector, Flutter — and the fine-tuning reportedly takes only a few hours on 8 A100 GPUs with batch size 8, which makes it a plausible template for production video search rather than a research-only artifact. The retrieval model and demo videos are released at huggingface.co/NCSOFT/multimodal-embedding.
Future Directions
- Visual-aware reranking. The authors state that visual information is not sufficiently incorporated during reranking and hypothesize that adding visual cues would produce more effective reranking.
- Latency reduction. The agent-based flow introduces higher latency than simple retrieval; the authors propose streaming LLM responses to improve the user experience and aim to balance retrieval effectiveness against latency.
- Reducing dependence on video descriptions. Ablations show a large drop when descriptions are removed, raising the question of how to recover that performance when descriptions are unavailable.
- Extending beyond the current evaluation. MSR-VTT and MultiVENT 2.0 are the only benchmarks reported; broader evaluation across other video retrieval settings and languages is not covered.
Target Audience
Researchers and practitioners in multimodal retrieval, video search, and LLM agent systems who want a concrete example of combining VLM fine-tuning, weight-space transfer, and agent orchestration into a deployed-style application. It is also relevant to engineers building conversational search products, since the paper documents specific component choices (frame counts, fusion weights, database settings, API versions) rather than only reporting scores.
Authors’ abstract
We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enhancing it with a retrieval vector from an image-text retrieval model, we overcome the limitations of traditional text-based retrieval systems in multimodal scenarios. The VLM-based retrieval model independently embeds video frames and audio transcriptions from an automatic speech recognition (ASR) module into a shared multimodal representation space, enabling V-Agent to interpret both visual and spoken content for context-aware video search. This system consists of three agents-a routing agent, a search agent, and a chat agent-that work collaboratively to address user intents by refining search outputs and communicating with users. The search agent utilizes the VLM-based retrieval model together with an additional re-ranking module to further enhance video retrieval quality. Our proposed framework demonstrates state-of-the-art zero-shot performance on the MultiVENT 2.0 benchmark, highlighting its potential for both academic research and real-world applications. The retrieval model and demo videos are available at https://huggingface.co/NCSOFT/multimodal-embedding.