Research
ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing
Overview Research area: Instruction-guided video editing, multimodal large language models (MLLMs), and diffusion-based video generation. Technical level: Advanced. The paper assumes familiarity with

- arXiv
- 2609.38541
- Published
- 2026-09-29
- Authors
- Donghao Zhou, Haoyang He, Fan Zhang, Hao Yang, Guisheng Liu, Xin Gao, Zhongwei Wan, Xingyuan Bu, Jie Wang, Qiangpeng Yang, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng
AI summary
Overview
Research area: Instruction-guided video editing, multimodal large language models (MLLMs), and diffusion-based video generation.
Technical level: Advanced. The paper assumes familiarity with diffusion transformers (DiTs), MLLMs, cross-attention conditioning, VAE latents, and curriculum learning.
Scope: The paper proposes ThinkV2V, a reasoning-driven framework that makes an MLLM explicitly "think" before a DiT generates an edited video, and backs it with a new 150K training dataset and a 308-pair benchmark.
What This Paper Is About
Most instruction-guided video editors use MLLMs only as semantic encoders, so they handle explicit prompts but fail when a user's instruction is indirect and requires causal or semantic reasoning to decode. ThinkV2V instead activates the MLLM as an explicit thinker: it generates chain-of-thought reasoning plus a refined prompt before any pixels are generated, and turns those reasoning features into conditioning signals for the video generator. The paper also supplies the missing training data and evaluation benchmark for this reasoning-centric editing setting.
Key Contributions
-
A reasoning-driven MLLM-to-DiT architecture. ThinkV2V pairs Qwen3-VL-Thinking-8B with Wan-2.1-5B through a learnable-query connector, so explicit MLLM thinking becomes a fixed-length conditioning signal for editing, while the original instruction embeddings and source-video VAE features are retained to avoid an information bottleneck.
-
A dedicated training-and-inference recipe. Progressive Curriculum Training moves the model from a 480p basic-alignment stage up to 720p high-resolution and then reasoning-intensive tuning, while Inference-Time Thinking Scaling iteratively refines candidate prompts at test time and lets the MLLM select the best one.
-
A new dataset and benchmark for reasoning-driven editing. ThinkV2V-150K provides 150K video pairs with complex, implicit, reasoning-oriented instructions, and ThinkV2V-Bench provides 308 video-instruction pairs for evaluating whether complex language understanding translates into executable edits.
-
State-of-the-art results at smaller scale. A 5B-scale DiT model outperforms several 10B-scale and larger baselines on both the complex ThinkV2V-Bench and the standard OpenVE-Bench, under two separate judge models.
Main Findings
-
Best overall scores on the reasoning-focused benchmark. On ThinkV2V-Bench, ThinkV2V (5B) reaches an overall score of 2.72 under the Seed-1.6-VL judge and 2.61 under Gemini-2.5-Pro, the highest overall among all compared methods. It leads on Global Style (3.13 under Seed-1.6-VL) and is competitive on Background Change, Local Change, Local Remove, and Local Add.
-
Gains transfer to standard editing. On OpenVE-Bench, ThinkV2V again has the best overall score under both judges: 2.93 (Seed-1.6-VL) and 2.72 (Gemini-2.5-Pro), with especially strong results on Local Change and Local Remove.
-
Smaller model beats larger baselines. ThinkV2V uses a 5B-parameter video editor and still surpasses DITTO (14B), VACE (14B), ICVE (13B), and other baselines. Some baselines show localized strengths — DITTO on Global Style, ReCo on Local Remove, OpenVE-Edit on Background Change — but these are concentrated rather than consistent across the full task set.
-
Explicit thinking helps, and answer-token features help most. In ablation under Seed-1.6-VL judgment, treating the MLLM as a static encoder ("w/o Thinking") scores 2.51, enabling thinking with all features scores 2.63, and using only the hidden states of the final answer tokens ("w/ Thinking + Answer Features") scores 2.68. The authors attribute the last gain to reduced redundancy from the thinking process.
-
Simple-to-complex curriculum is the winning order. Among five training schedules, "Simple-to-Complex" scores 2.68, compared with "Simple Only" 2.52, "Mixed Simple+Complex" 2.47, "Complex-to-Simple" 2.41, and "Complex Only" 2.10. Training on complex instructions alone performs worst.
-
Iterative refinement needs selection to pay off. Without scaling, the overall score is 2.68; adding only Serial Refinement gives 2.67 (no improvement); adding Selection after serial refinement raises it to 2.72, the best in that ablation group.
-
Qualitative failures of baselines are intent-related. In the shown cases, prior methods produce incomplete style transformations, incorrect local modifications, misplaced additions, or removal of the wrong target, while ThinkV2V better preserves unedited content and temporal consistency.
Methodology in Plain English
The system has three pieces. First, an MLLM (Qwen3-VL-Thinking-8B) reads the source video and the user's instruction along with a system prompt, then writes out reasoning followed by a refined, more precise prompt. Only the hidden states of the tokens generated after the </think> tag are passed downstream, treated as high-level semantic features.
Second, a connector with learnable query tokens uses cross-attention to compress those hidden states into a fixed-length feature sequence. The query length is set to 512, and the connector's final layer is zero-initialized to stabilize early training.
Third, the DiT (Wan-2.1-5B, initialized from Lucy-Edit) generates the edited video under multiple conditions: the connector's fixed-length features are concatenated along the token dimension with the original instruction's text embeddings and injected via cross-attention, while the source video's VAE features are fused with the noisy latents by channel concatenation.
Training proceeds in three stages on a 32-GPU cluster. Stage 1 trains at 480p for 2 epochs at a learning rate of 1e-5; Stage 2 trains at 720p for 0.5 epochs at 1e-6; Stage 3 trains at 720p for 1.5 epochs at 1e-6. Stages 1 and 2 use OpenVE-HQ-1M (simple instructions); Stage 3 uses ThinkV2V-150K (complex instructions). The MLLM stays frozen throughout; only the connector and DiT are optimized.
At inference, the model feeds its refined prompt back into the MLLM over multiple rounds to produce a sequence of candidate prompts, then applies best-of-N selection with N=8, letting the MLLM pick the candidate that best matches the original intent and source video. The cached hidden states of that selected candidate drive the connector and DiT.
The data was built from OpenVE-3M: five categories were kept (Global Style, Background Change, Local Remove, Local Add, Local Change) and three excluded (Camera Edit, Subtitle Edit, Creative Edit). Two million samples were rescored with Gemini-2.5-Flash and filtered by average inter-frame CLIP similarity (CLIP-F) and Temporal Flickering (TF), keeping the top 50% to form OpenVE-HQ-1M. From there, 20K to 40K samples per category suitable for reasoning-instruction synthesis were selected, and Gemini-2.5-Pro rewrote their direct instructions into complex, reasoning-oriented ones, yielding 150K video pairs. The benchmark, ThinkV2V-Bench, was created by using Gemini-3.1-Pro to rewrite instructions in OpenVE-Bench, producing 308 video-instruction pairs across the same five categories.
Why This Matters
Impact on research. The paper argues that the bottleneck in instruction-guided video editing is not generation quality but the absence of an explicit "think-before-edit" step. It supplies both a mechanism (reasoning features as conditioning) and the evaluation infrastructure (a 150K dataset and a 308-pair reasoning benchmark) that the area previously lacked, and shows a 5B model can beat 10B-plus baselines.
Real-world applications:
- Consumer and creator video tools where users type loose, everyday-language requests such as "remove the thing blocking the view" rather than naming an object.
- Advertising and social media pipelines that need style or background changes described by mood or reference rather than by explicit label.
- Post-production cleanup, where removing or adding objects requires understanding object relations and scene dynamics across frames.
- Automated content moderation or compliance editing, where the target of an edit is defined by function or causal role.
Industry relevance. The work is a collaboration spanning ByteDance, The Chinese University of Hong Kong, Zhejiang University, and The Ohio State University. Its demonstration that a 5B editor can outperform 14B baselines matters for deployment cost, and the data construction pipeline (rescore, filter, rewrite) offers a repeatable pattern for building reasoning-focused datasets in adjacent visual generation tasks.
Future Directions
- More efficient thinking-involved frameworks. The authors state this explicitly as their first future goal, motivated by the cost of multi-round refinement and best-of-N selection at inference.
- Integrating reasoning editors into broader content creation systems. The conclusion calls for empowering integrated visual content creation with stronger instruction understanding, planning, and execution.
- Understanding when iterative refinement drifts. Serial refinement alone scored lower than no scaling at all (2.67 versus 2.68), so characterizing and preventing reasoning drift remains open.
- Benchmark and dataset limitations. The paper restricts itself to five edit categories and notes current limitations are discussed in Appendix C, leaving open how reasoning-driven editing should be evaluated on the excluded categories such as Camera Edit and Creative Edit.
Target Audience
Researchers and engineers working on controllable video generation, diffusion transformers, and multimodal LLMs; practitioners building instruction-following editing products who need to know the cost-quality tradeoff of adding explicit reasoning; and dataset or benchmark builders interested in how to convert direct editing instructions into reasoning-oriented ones through model-assisted rewriting.
Authors’ abstract
Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.