Research
CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
Overview Research area: Computer vision and multimodal video understanding — specifically agentic vision-language models (VLMs) applied to ultra-long video temporal grounding and long-video question a

- arXiv
- 2609.40048
- Published
- 2026-09-30
- Authors
- Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen
AI summary
Overview
Research area: Computer vision and multimodal video understanding — specifically agentic vision-language models (VLMs) applied to ultra-long video temporal grounding and long-video question answering.
Technical level: Advanced. The paper assumes familiarity with vision-language models, tool-using agents, prompt/skill evolution, and temporal grounding metrics (IoU, mIoU, Recall@0.5, Precision/Recall AUC).
Scope: The paper proposes CoEvoWhen, a framework that jointly evolves an agent's high-level tool-use policies and its executable media tools from a frozen VLM's own execution trajectories, evaluating the result on three ultra-long video temporal grounding benchmarks and two long-video QA benchmarks.
What This Paper Is About
Ultra-long videos (tens of minutes to hours) make temporal grounding hard: the target event may occupy only a tiny fraction of the timeline, so an agent must search a vast temporal range while still resolving fine-grained details and boundaries under a limited visual budget. Existing agentic systems handle this with multi-step search and tool use, but their observation strategies and media tools are largely predefined by humans, and evidence needs vary across videos and queries. CoEvoWhen's goal is to let task experience automatically improve both what the agent can observe (the tools) and how it decides to observe (the policies), without updating any model parameters.
Key Contributions
-
Policy–tool coevolution framework. CoEvoWhen jointly evolves high-level policies and executable media tools from a VLM's execution trajectories, distilling ultra-long video grounding experience into a reusable external skill while keeping model parameters frozen.
-
Evolved orchestration of complementary observations. The framework evolves coordination between image-based observations (compact coverage of long temporal ranges, comparison of distant candidates) and video-based observations (local temporal continuity for motion, event order, state transitions, boundaries), enabling the VLM to acquire evidence autonomously without hand-designed coordination strategies or a stronger external planner.
-
Systematic evaluation across five benchmarks and three VLMs, complemented by policy–tool and image–video ablations, substantiating effectiveness, cross-VLM generalizability, and transferability of the evolved skill to general long-video understanding.
-
A reusable external skill consisting of policy documents and executable tools, which the frozen VLM invokes at inference without a separate, stronger planning model.
Main Findings
-
Grounding accuracy improves across all three ultra-long benchmarks. With Qwen3.5-27B, the evolved skill beats the base skill on VUE-LVTR (IoU AUC 0.4107 to 0.5137, a relative gain of 25.1%), ExtremeWhenBench (mIoU +74.9%, Recall@0.5 +88.4%), and CoMET-Bench (mIoU +0.0280, a relative gain of 20.7%; Rejection F1 +10.93, a relative gain of 16.5%). The evolved skill achieves the best results on all reported metrics, surpassing seven VLM-based and agent-based baselines.
-
Accuracy gains come with lower visual token cost. Average visual tokens per query drop from 202.58k to 141.89k on VUE-LVTR (29.96% fewer), from 227.10k to 201.17k on ExtremeWhenBench (11.42% fewer), and from 119.66k to 97.10k on CoMET-Bench (18.85% fewer). Model calls per query also decrease on all three benchmarks.
-
The visual budget is reallocated, not simply shrunk. On ExtremeWhenBench, image-based observation tokens rise from 179.26k to 197.49k while video-based observation tokens fall from 47.84k to 3.68k, reflecting a shift toward image-based coverage and selective video verification.
-
Gains hold across different VLM backbones. On ExtremeWhenBench, coevolution improves mIoU by 99.2% on Qwen3.5-9B and 24.7% on Qwen3.6-27B, supporting cross-VLM generalizability.
-
The evolved skill transfers to long-video QA without task-specific evolution. Qwen3.5-27B's overall accuracy improves by 8.65 points on LVBench and 7.98 points on LSDBench when the grounding-evolved skill is applied directly.
-
Joint evolution beats evolving either component alone. Joint evolution exceeds policy-only by 0.0243 and tool-only by 0.0735 in IoU AUC while reducing visual token cost by 40.6% and 25.5%, respectively. Freezing policies during evolution causes a larger accuracy loss than freezing tools, suggesting tools matter most when strategically orchestrated.
-
Image and video observations are complementary. The evolved image+video skill outperforms both single-modality evolved skills on all grounding metrics while using fewer visual tokens. In the video-only setting, evolution cuts tokens from 693.52k to 221.98k while raising IoU AUC from 0.2457 to 0.4772; the image-only setting yields a smaller IoU AUC rise from 0.4173 to 0.4374.
-
Trajectory feedback drives specific, interpretable adjustments. Candidate omissions and boundary errors motivate image-based tools with candidate-retention and local-refinement policies, while video-based observation targets action verification, order discrimination, and continuity assessment.
Methodology in Plain English
The researchers start from a deliberately minimal "base skill" for a frozen VLM. The base policy states only the basic grounding protocol and elementary descriptions of observation modes, and the base tool set has six tools: probe_media (read video metadata), extract_frames_at (pull frames at given timestamps), extract_video_clip (extract an audio-free clip), inspect_images and inspect_video (place images or video into the VLM's visual context), and final_answer (submit the prediction).
Skill evolution proceeds over K rounds. In each round, the frozen VLM uses the current skill to ground a batch of queries, producing full execution trajectories. These trajectories are paired with ground-truth temporal intervals to form task feedback. An external skill updater — Codex (GPT-5.5, xhigh) — then analyzes the feedback, revising policy documents and using its coding ability to modify, consolidate, or retire existing tools, or write new ones. Notably, tools change in their description, interface, and source code, so improvements reach beyond prompts and invocation sequences into actual evidence sampling and presentation. A feedback batch is formed after every four completed queries, and evolution makes a single pass over 100 queries with VLM weights frozen throughout.
The evolution set comprises 100 queries from 100 distinct videos, drawn from purely visual temporal retrieval instances of VUE-TR-V2. Video durations range from 30.35 to 118.61 minutes (mean 52.90 minutes). The set includes 80 sparse-target queries, 88 boundary-sensitive queries (at least one target interval no longer than 10 seconds), and 30 video-edge queries; queries split into 18 keyword, 34 phrase, and 48 sentence types, and 56 single-interval versus 44 multi-interval cases.
At inference, the evolved skill is fixed. The VLM autonomously selects a tool each round based on the query, accumulated interaction history, and the evolved policy, with at most 24 inference rounds per query, thinking enabled, and greedy decoding. No separate, stronger planning model is used. Evaluation covers VUE-LVTR (video durations 30.1–105.5 minutes, mean 50.9), ExtremeWhenBench (45.0–542.5 minutes, mean 75.8), CoMET-Bench (30.0–123.7 minutes, mean 50.5), plus LVBench and LSDBench for QA transfer. Efficiency is measured as cumulative visual tokens summed across rounds and averaged per query, with separate image and video costs, alongside model calls, local tool calls, and visual observation counts.
Why This Matters
Impact on research. The paper argues that observation mechanisms and media processing capabilities in video agents have remained largely predefined, and that policy refinement alone is limited by existing tool capabilities while tool expansion alone does not guarantee effective use. By coupling the two within a single evolution loop driven by the same execution feedback, it offers a concrete alternative to hand-designed agentic pipelines and to approaches that evolve only skills or only tools (for example, the paper notes SkillSmith restricts tool updates to predefined operations on an existing library, and META leaves the task-level orchestrator outside the evolution loop). It also demonstrates that grounding-evolved skills transfer to unrelated long-video QA, a cross-task reusability result.
Real-world applications.
- Footage retrieval and clip extraction for media production, where users describe an event in natural language and need the exact segment on the timeline.
- Video editing assistance that must locate events in hours-long raw footage before cutting or assembling.
- Surveillance and compliance review, where a rare event occupies seconds within a long recording and both the event and its boundaries matter.
- Long-video question answering and summarization assistants that need to locate evidence efficiently under compute or context limits.
Industry relevance. The framework requires no parameter updates to the underlying VLM and no stronger external planner at deployment, so organizations can adapt an off-the-shelf model to a new long-video domain by evolving a portable skill artifact. The simultaneous reduction in visual token cost (for example, 29.96% fewer visual tokens per query on VUE-LVTR with Qwen3.5-27B) and model calls directly affects inference expense.
Future Directions
-
Scope of evolution data. Evolution used 100 queries from 100 distinct videos with a single pass; how far the skill generalizes as query distributions, video lengths, or languages shift is not established by the reported experiments.
-
Dependence on the external updater. The skill updater is Codex (GPT-5.5, xhigh), and the paper does not report how sensitive policy and tool quality is to the updater's capabilities or to the batch size of four queries.
-
Boundaries of cross-task transfer. Transfer was demonstrated on LVBench and LSDBench for grounding-evolved skills; whether skills evolved for other video tasks similarly transfer in the reverse direction is not reported.
-
Tool set restructuring over longer horizons. Tools can be created, consolidated, or retired across rounds, but the paper reports one evolution pass; the dynamics of tool proliferation or degradation over many more rounds remain open. The paper's listed appendix section on Limitations and Future Work is not included in the provided content.
Target Audience
Researchers and advanced practitioners in multimodal agents, long-video understanding, and tool-using LLM/VLM systems, particularly those working on temporal grounding, video question answering, or self-improving agent skills. It is also relevant to engineers building production video retrieval and editing pipelines who need accurate localization at controlled inference cost. Readers should be comfortable with agentic inference loops, temporal grounding metrics, and the distinction between parameter-level and non-parameter-level adaptation.
Authors’ abstract
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.