Skip to content
AI.info

Research

OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

Overview Research area: Multimodal / omni-modal foundation models — specifically agentic audio-visual reasoning, tool use, and reinforcement learning for large language models that process text, audio

OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
arXiv
2610.02181
Published
2026-10-01
Authors
Haibo Wang, Jiteng Mu, Jialu Li, Jingru Yi, Yuanjun Xiong, Jianming Zhang, Lifu Huang, Mingze Xu

AI summary

Overview

Research area: Multimodal / omni-modal foundation models — specifically agentic audio-visual reasoning, tool use, and reinforcement learning for large language models that process text, audio, and video jointly.

Technical level: Advanced. The paper assumes familiarity with Chain-of-Thought reasoning, multimodal LLM architectures, policy-gradient reinforcement learning (GSPO), attention masking, and multimodal benchmark suites.

Scope: The paper proposes OmniSeek, an agentic framework, dataset, and three-phase RL training recipe that lets a 30B Omni-LLM actively retrieve raw audio and video clips across multiple turns instead of passively encoding an entire audio-visual stream in one pass.

What This Paper Is About

Current Omni-LLMs ingest a whole audio-visual sequence in a single forward pass, which dilutes brief sounds and fine visual details as context grows, causing models to fall back on language priors or single-modality shortcuts. The paper's goal is to make evidence acquisition part of the reasoning process itself: the model decides whether to look or listen, over which temporal window, and when it has gathered enough evidence. It does so by training an Omni-LLM to alternate between thinking, calling audio/video retrieval tools, and observing the returned raw segments until it can answer a multi-hop question.

Key Contributions

  1. An agentic framework (OmniSeek) that enables Omni-LLMs to retrieve, interleave, and reason over decoupled audio and video evidence across multiple turns, using an iterative <think> → <tool_call> → <observe> loop with two tools: get_audio_clip(start, end) and get_video_clip(start, end, fps, resolution).

  2. OmniTraj-170K, a dataset of multi-turn Chain-of-Thought trajectories with interleaved audio-visual evidence — 169,725 trajectories over 39,797 videos spanning 19 cross-modal question types — built by a three-stage data engine.

  3. An Audio-Visual Necessity objective for RL training that rewards trajectories whose reasoning genuinely depends on both modalities, measured through modality-specific attention masking, which discourages single-modality shortcuts without requiring additional rollouts.

  4. A three-phase training strategy (cold-start SFT, broad-exploration RL, hard-example refinement) plus extensive experiments showing adaptive cross-modal evidence seeking and state-of-the-art or competitive results across ten omni benchmarks and four general video benchmarks.

Main Findings

  • Leading open-source results across omni benchmarks: OmniSeek (30B) reaches Daily-Omni 80.0, AVUT 78.8, WorldSense 62.4, FutureOmni 58.3, OmniVideoTest 69.5, VideoHolmes 74.6, JointAV 72.8, OmniVideoBench 47.7, MMOU 70.4, and LVOmni 44.2, against base Qwen3-Omni-Instruct scores of 71.9, 76.5, 55.1, 53.6, 54.5, 59.1, 63.6, 43.6, 54.1, and 35.8 respectively.

  • Largest gains on long-form and complex benchmarks: On MMOU and LVOmni, OmniSeek achieves 70.4% and 44.2%, exceeding the base Qwen3-Omni-Instruct by +16.3% and +8.4%.

  • Advantage over text-only CoT: Against OmniVideo-R1, OmniSeek scores 74.6% vs. 62.9% on VideoHolmes and 47.7% vs. 44.8% on OmniVideoBench. The paper reports a dedicated ablation where text-only CoT (single-turn text, no tools) gains only 71.9%→73.0% on Daily-Omni and actually drops 55.1%→54.3% on WorldSense, while OmniSeek's multi-turn multimodal paradigm delivers margins of +13.5% on OmniVideoTest, +8.1% on WorldSense, and +7.0% on Daily-Omni over that text-only baseline.

  • Competitive with closed-source models: The paper states OmniSeek is competitive against closed-source models, outperforming Gemini-2.5-Pro on FutureOmni, VideoHolmes, and JointAV (as named in the text).

  • General video understanding is preserved and improved: OmniSeek reaches Video-MME 78.5, LongVideoBench 66.4, MLVU 77.1, and LVBench 51.4, versus base Qwen3-Omni-Instruct at 76.8, not reported, 75.2, and 50.2. It surpasses much larger models such as Qwen2.5-VL (72B) and LLaVA-OneVision (72B) on these tasks, and holds an advantage over text-only reasoning approaches (78.5% vs. 73.6% on Video-MME for OmniVideo-R1).

  • An "alignment tax" appears during cold-start SFT, then is recovered by RL: Phase 1 SFT drops the base model from 71.9% to 69.1% on Daily-Omni and 55.1% to 50.3% on WorldSense. Phase 2 RL recovers and exceeds the base, pushing LVOmni from 35.0% to 41.6% and VideoHolmes from 55.9% to 69.6%.

  • Phase 3 rewards and rollout scaling help: Increasing rollout size from G=8 to G=16 raises Daily-Omni from 75.9% to 78.2% and FutureOmni from 56.4% to 58.1%. Adding the Audio-Visual Necessity reward contributes a further +1.8% on Daily-Omni, +3.2% on WorldSense, +2.2% on OmniVideoTest, and +2.9% on VideoHolmes over the same configuration without it.

  • Retrieval depth adapts to task complexity: Accuracy on JointAVBench peaks at exactly 3 tool calls, rising from 68.1% (2 calls) to 80.8%; Daily-Omni reaches 87.1% at four tool calls; long-form benchmarks VideoHolmes and MMOU peak at 75.9% and 73.3% when executing 5 or more tool calls.

  • The trajectory corpus is useful on its own: Under a controlled single-turn SFT paradigm, training on a 100K sample of OmniTraj-170K (7:3 open-ended to multiple-choice mix) yields gains over the zero-shot base of +5.8% on LVOmni and +3.9% on Daily-Omni, and leads OmniVideo-100K on Daily-Omni (+2.1%) and OmniVideoBench (+1.7%). OmniVideo-100K scores higher on OmniVideoTest (61.7% vs. 59.2%), which the authors attribute to OmniVideo-100K sharing origin, distribution, and stylistic design with that benchmark.

  • Dataset composition: 169,725 trajectories over 39,797 videos; 19 cross-modal question types; source videos from FineVideo spanning 122 categories; evidence spans mostly lasting between 3 and 10 seconds; 76.1% of trajectories require two tool calls and 23.9% require three or more; evidence chains contain an ordered list of 2 to 7 spans, each with at least one audio span and one video span.

Methodology in Plain English

The authors start by building training data rather than relying on the model's instincts. Their data engine works in three stages. First, it cuts videos into coherent scenes using PySceneDetect, merging adjacent shots until they reach roughly 15-second windows, then has Qwen3.5-397B-A17B write dense captions that are pinned to timestamps for both the visual track and (via original ASR transcripts or tools like Qwen3-Omni-Captioner or Qwen3-ASR) the audio track. Second, it synthesizes 19 kinds of audio-visual questions that cannot be answered from either modality alone, and requires each question to come with an ordered evidence chain of 2 to 7 spans whose text and timestamps are copied directly from the Stage 1 annotations. Third, it assembles full multi-turn trajectories: the tool calls and their returned observations are constructed deterministically from the evidence chain, and the model only writes the internal <think> nodes that plan, reflect, and finally synthesize the answer.

Training proceeds in three phases on top of Qwen3-Omni-30B-A3B-Instruct. Phase 1 supervises the model on OmniTraj-170K, but mixes only 10% tool-calling trajectories with 90% standard single-turn QA to avoid damaging general ability. Phase 2 switches to reinforcement learning with GSPO on roughly 31K multiple-choice questions (30K from OmniVideo100K plus 1K from the VideoHolmes training split), using three rule-based rewards: correctness of the answer, format compliance, and a tool-use reward granted only when the model both calls a tool and answers correctly. Phase 3 mines 8K cases the Phase 2 model failed on, doubles the rollout size to G=16, and adds the Audio-Visual Necessity reward.

That necessity reward is the paper's most distinctive technical idea. Given a completed trajectory, the authors rerun the same token sequence twice with teacher forcing, each time masking out one modality's keys in the attention mask. If removing audio from the attention mask sharply lowers the probability of the model's own response tokens, then those tokens genuinely depended on audio — and similarly for video. Negative drops are clipped to zero and the drops are averaged per modality. Because the goal is joint dependence rather than strong reliance on one stream, the final reward uses the minimum of the two necessity scores, gated by whether the answer was correct. This costs only two lightweight no-gradient forward passes, since the full-context likelihood is already available from the policy-gradient pass.

Why This Matters

Impact on research: The paper reframes omni-modal reasoning from a passive encoding problem into an active, agentic evidence-seeking problem, and shows that interleaving raw retrieved audio and video back into the context beats scaling up textual chains of thought. It also introduces a low-cost, token-level way to measure whether a model actually uses both modalities, which is relevant to a broader class of multimodal RL work where outcome-only rewards can silently reward shortcuts.

Real-world applications:

  • Long-form media and archive search, where a specific sound or a brief on-screen detail must be located inside hours of footage.
  • Accessibility tooling, such as question answering or description generation for audio-visual content that depends on both what is seen and what is heard.
  • Video editing and production assistants that need to point at exact time windows for audio and visual assets.
  • Content moderation, journalism, and sports or news analysis over the domains represented in the source corpus (education, science, news, sports among the 122 categories).

Industry relevance: The work is authored by Adobe Research together with University of California, Davis, and the authors report it was done during an internship at Adobe Research — pointing at direct relevance to creative, video-centric products. Because OmniSeek is built on a 30B open base model and improves on benchmarks while preserving general video understanding, it is a practical recipe for teams that want agentic capability without training a frontier-scale model. The result that video, audio, and text can be routed independently, and that the agent decides when to stop, also fits product settings where retrieval cost and latency must be balanced against answer quality.

Future Directions

  • Extending the tool space beyond the two retrieval operations. The paper equips the agent only with get_audio_clip and get_video_clip; adding tools such as code execution or web search — which the related-work section notes DeepEyesV2 explored for images — is an open extension.

  • Scaling context and retrieval depth further. The model already handles clips up to 2049-second benchmark videos (LVOmni) and 4038-second videos (LVBench) with a 65536 maximum sequence length and up to 8 interaction turns; whether longer horizons or more turns continue to help, and for which tasks, remains open.

  • Generalizing the Audio-Visual Necessity objective. The reward is defined for exactly two modalities via attention masking. How it would extend to more than two modalities, or to other agentic settings where outcome rewards hide shortcut behavior, is not addressed.

  • Broadening the data engine's domain coverage. OmniTraj-170K draws from FineVideo across 122 categories and 19 question types; whether the same pipeline transfers cleanly to domains such as medical, industrial, or surveillance audio-visual data is untested here, and the paper does not report results outside its evaluated benchmark suite.

Target Audience

Researchers and engineers working on multimodal and omni-modal foundation models, agentic tool-use and reinforcement learning for LLMs, and long-form video/audio understanding. It also suits practitioners who need a concrete, reproducible recipe — dataset construction, reward design, and hyperparameters — for turning a base Omni-LLM into an interactive reasoning agent, and readers interested in evaluation methods that detect when a model only pretends to use one of its modalities.

Authors’ abstract

We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.

Read the original paper