Skip to content
AI.info

Research

OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video

Overview Research area: Multimodal deep research agents — vision-language models (VLMs) that actively inspect images or video, call external search tools, and compose evidence into answers. Technical

arXiv
2610.12419
Published
2026-10-08
Authors
Hongyu Li, Manyuan Zhang, Kaituo Feng, Shu Chen, Dian Zheng, Hao Li, Hao Yu, Zhangquan Chen, Zoey Guo, Ray Zhang, Shaofei Huang, Tianrui Hui, Linjiang Huang, Si Liu

AI summary

Overview

Research area: Multimodal deep research agents — vision-language models (VLMs) that actively inspect images or video, call external search tools, and compose evidence into answers.

Technical level: Advanced. The paper combines agentic tool-use training, supervised fine-tuning, reinforcement learning with GRPO, and a structured evidence-graph data pipeline.

Scope: This paper introduces OneSearch-VL, a single 8B-parameter policy trained to perform deep research over single images, multi-image collections, and videos, together with a graph-based data engine, a process reward, and two new operation-oriented benchmarks.

What This Paper Is About

Multimodal deep research — locating visual clues, retrieving external knowledge from the web, and combining the two into a verified answer — has been studied mostly for single images or, separately, for video. These settings need different visual operations (region cropping versus temporal localization), yet they share the same underlying workflow: ground a visual anchor, retrieve facts about the entity behind it, and compose those facts into an answer. The paper asks whether one policy can learn this workflow jointly across single images, multi-image sets, and videos, and how to keep track of which visual anchor supports which retrieved fact so that both training data and evaluation can be checked at the level of individual research operations rather than only final answers.

Key Contributions

  1. VGEG-centered data engine. The authors propose the Visually Grounded Evidence Graph (VGEG), a task-level structure that links localized visual anchors to real-world entities, records multi-hop relational paths and source-supported facts, and specifies the operations that turn those facts into an answer. A data engine built on VGEG constructs and verifies multi-image and video questions and filters expert tool-use trajectories.
  2. A unified agent with evidence-level supervision. OneSearch-VL is trained for joint single-image, multi-image, and video deep research using OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K. Its RL stage adds the Evidence-aware Visual-Grounded Rubric reward (EVGR), which scores evidence traceability and visual grounding of complete trajectories.
  3. Operation-oriented benchmarks. OneSearch-MI-Bench and OneSearch-Video-Bench organize questions by the six principal research operations encoded in their VGEGs, enabling fine-grained analysis of fact retrieval and composition rather than only aggregate accuracy.

Main Findings

  • Single-image results: OneSearch-VL-8B reaches an average of 58.3 across seven single-image deep-research benchmarks (SimpleVQA, VDR, MMSearch, LiveVQA, BrowseComp-VL, FVQA, InfoSeek), outperforming the same-scale Qwen3-VL-8B Agent by 16.3 points and OpenSearch-VL-8B by 1.7 points. It improves over OpenSearch-VL on all seven benchmarks, with gains of 2.8, 2.1, and 2.0 points on VDR, InfoSeek, and MMSearch, and is the best listed agentic method on six of the seven.
  • Multi-image and video results: OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 points on OneSearch-MI-Bench, 17.6 points on OneSearch-Video-Bench, and 27.0 points on VideoDR. It improves all six research operations on both new benchmarks.
  • Where the operation-level gains come from: On multi-image tasks, knowledge-conditioned counting, multi-anchor arithmetic, and multi-anchor comparison improve by 29.8, 26.3, and 22.5 points. On video tasks, multi-hop retrieval, multi-anchor arithmetic, and multi-anchor joining improve by 25.5, 20.0, and 19.2 points — categories that require filtering, connecting, or computing over facts tied to multiple visual anchors.
  • SFT data mixture ablation: Training on any single trajectory type raises the six-benchmark average from 40.7 to between 51.7 and 55.1. The video-only model achieves the strongest single-type average of 55.1 and also improves all three single-image benchmarks. Joint training on all three types reaches 55.8 and gives the best result in this ablation on SimpleVQA, OneSearch-MI-Bench, and OneSearch-Video-Bench.
  • RL reward ablation: Starting from the joint SFT model at 55.8, accuracy-only RL reaches 56.3, adding query quality reaches 57.3, adding Trace or Ground alone reaches 59.0 and 59.3, and the full reward reaches 61.1. The full reward improves by 3.8 points over answer-plus-query rewards and by 2.1 and 1.8 points over the two single-dimension variants, performing best on five benchmarks and tying for best on the remaining one.
  • Shared action space: All three visual input types use the same dispatcher, observation format, and trajectory representation; only the tool set differs, with temporal tools (SelectTimespan, SelectFrame) added for video.

Methodology in Plain English

The authors build a data pipeline that turns raw web video into verified research tasks. It starts from 2.5M candidate YouTube video records and filters them down to 70,781 through upload-date filtering (on or after January 1, 2023), rule-based metadata scoring, category balancing, a GPT-4.1 snippet assessment, and duration balancing. Each retained video is sampled at 2 FPS, split into clips of ten frames, captioned, reduced to key frames with a similarity-based deduplication step (image similarity at least 0.9 or text similarity at least 0.8), aggregated into events, and annotated with localized searchable objects.

Those visual anchors are then connected to real-world entities and web facts through image search, OCR, and text retrieval, producing an input-level evidence graph. A task-conditioned projection selects the anchors, facts, and answer-producing operations needed for a particular question, yielding a VGEG. From that, the generator produces a question, a reference answer, and the graph itself, while the pipeline rewrites explicit entity mentions into visually grounded references, remaps frame references to image indices for multi-image tasks, and filters out answer leakage, ambiguous references, and questions answerable without the visual input. An expert model then interacts with the tool environment to produce trajectories, and only trajectories passing answer-correctness and process-quality checks become the 110K SFT set. A separate 10K set drives RL.

Training proceeds in two stages. SFT learns the tool-interaction policy from expert trajectories, with loss applied only to the model's own reasoning, tool-command, and answer tokens rather than to tool observations. RL then uses GRPO with a composite reward combining answer correctness, query quality, and EVGR, gated by format validity. EVGR uses two judge calls derived from the VGEG rubric: one checking evidence traceability (whether tool observations actually establish the required entities and fact hops) and one checking visual grounding (whether the right objects, regions, or frames are identified and used).

Why This Matters

Most multimodal agents are evaluated only on final answer accuracy, which hides whether a model found the right clue, traced the right relation, or simply guessed. This work shows that binding visual anchors to retrieved facts in an explicit graph, and then using that same structure for data construction, trajectory filtering, reward design, and evaluation, produces measurable gains on compositional questions that require combining evidence across images or frames. The finding that trajectories from one visual input type improve performance on others is a practical result for anyone building multi-task agent training pipelines.

Real-world applications:

  • Fact-checking and investigative research over photo sets or video footage, where each claim must be traceable to a specific frame or region and a specific web source.
  • Media and newsroom workflows that need to identify entities appearing across multiple images or video segments and assemble supporting context.
  • E-commerce and catalog work that must match visual product or object evidence to external specifications and prices.
  • Video archive and sports analysis pipelines that localize events temporally, then retrieve external knowledge about the entities involved.

Industry relevance: the recipe — a unified tool interface, expert trajectory SFT, and process rewards derived from structured evidence — is directly transferable to production research assistants that combine vision models with web search, and the two new benchmarks give teams a way to diagnose failures at the level of individual operations instead of overall accuracy.

Future Directions

  • Tool and web dependence: OneSearch-VL relies on external tools such as TextSearch and ImageSearch and on changing webpages, which affects evidence availability and exact reproducibility. The authors suggest preserving retrieval snapshots and evaluating robustness to tool failures.
  • Annotation and judging reliability: Automated VGEG construction and model-based rewards may introduce annotation errors or judging biases; human calibration and open multimodal process judges could improve trajectory-level assessment.
  • Cost control: Multi-turn interaction incurs additional inference and tool costs, motivating adaptive budgets and cost-aware training.
  • Open questions the paper raises but does not answer: how well the unified policy transfers to visual input types beyond the three studied, and how the operation-level taxonomy generalizes to domains whose answers require operations not covered by the six categories.

Target Audience

Researchers and engineers working on multimodal agents, tool-augmented VLMs, and vision-language RL, especially those interested in process supervision and evidence-grounded evaluation. It is also useful for practitioners building image- or video-based research assistants who need benchmarks that isolate retrieval and composition failures, and for readers interested in how structured intermediate representations can serve simultaneously as data scaffolding, reward signal, and evaluation schema.

Authors’ abstract

Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preserve the dependencies linking localized visual anchors, entity relations, source-supported facts, and answer-producing operations. We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation. Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. Using these data, we assemble OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K for SFT and RL, respectively. We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL. For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs. Experiments show that OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, while also achieving substantial gains across 7 image benchmarks and VideoDR. Project repository: https://github.com/appletea233/OneSearch-VL

Read the original paper