Research
VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling
Overview Research area: Computer vision — diffusion-based video editing, in-context generative modeling, and transfer of image editing priors to video. Technical level: Advanced. The paper assumes fam
- arXiv
- 2610.12104
- Published
- 2026-10-08
- Authors
- Leigang Qu, Feng Cheng, Ziyan Yang, Bangbang Yang, Zhaoyang Huang, Wei Chow, Yicong Li, Wenjie Wang, Tat-Seng Chua, Yan Zeng
AI summary
Overview
Research area: Computer vision — diffusion-based video editing, in-context generative modeling, and transfer of image editing priors to video.
Technical level: Advanced. The paper assumes familiarity with diffusion transformers (DiT), flow matching, rotary position embeddings (RoPE), VAE latent tokenization, and classifier-free guidance.
Scope: The paper introduces VINCIE-NExT, a framework that decomposes video editing into a Video → Image → Image → Video chain and drives it with an in-context image editing pair, so that video editing can be trained largely from image editing data rather than expensive paired video data.
What This Paper Is About
Building a video editor is much harder than building a video generator because editing supervision requires (source video, instruction, edited video) triplets, which are costly to annotate and hard to synthesize at scale, while image editing already has millions of (source image, instruction, edited image) pairs. VINCIE-NExT addresses this by routing editing intent through the image domain: it factorizes video editing into Video → Image → Image → Video sub-tasks and feeds an image editing pair as an in-context visual demonstration that acts as a spatial appearance blueprint for every generated frame. The goal is a single diffusion backbone that learns video editing from heterogeneous image and video corpora without large-scale paired video editing data.
Key Contributions
- An in-context video editing paradigm. The paper proposes transferring image editing capability to video through in-context visual demonstrations, alleviating the need for large-scale paired video editing data.
- Joint in-context image–video modeling with sub-task decomposition. A unified framework factorizes editing into Video → Image → Image → Video, enabling scalable training from heterogeneous image and video corpora under a single unified diffusion (flow-matching) objective.
- TDF3D-RoPE and Chain-of-Editing. A Time Dual-Frequency 3D Rotary Position Encoding enables pixel-faithful propagation of edits from image demonstrations to video frames, and a two-stage Chain-of-Editing inference strategy builds on it to provide test-time scaling without retraining.
- State-of-the-art results on OpenVE-Bench. Comprehensive experiments report state-of-the-art performance across diverse editing categories, with ablations validating each component.
Main Findings
- Overall benchmark score. On OpenVE-Bench (431 video clips, eight editing categories, evaluated by Gemini 2.5 Pro on a 1–5 scale), VINCIE-NExT scores 3.08 overall at 640×480, surpassing all open-source methods by a substantial margin, with a 24% relative gain over OpenVE-Edit (2.49), while closing the gap to the proprietary Runway Aleph (3.65). Results are reported as consistent under Seed1.6-VL as an independent judge.
- Category-level strengths. Global Style is the standout category at 4.17, versus Runway Aleph's 3.72. Local Remove improves dramatically from 1.85 to 3.24 over OpenVE-Edit. Local Change (3.48), Subtitle Edit (3.45), and Creative Edit (3.41) also benefit. The main limitation is Camera Edit (2.06) versus Runway Aleph (4.53), because it requires novel-viewpoint synthesis, a geometric operation outside the scope of appearance editing; Local Add (2.27) and Background Change (2.55) also show room for improvement since both require synthesizing content absent from the source video.
- Data composition matters, and contributions are additive. I2I-only training fails to generalize to video (1.14). The V↔I↔I data source (video-based in-context image editing) transfers image editing priors without direct video supervision (1.34, or 1.80 with CoE). V2V* provides the strongest spatially-aligned signal (2.49, or 2.57 with CoE). All three combined without CoE reach 2.53, and with CoE reach 2.65, the best overall score.
- Paired video data alone is not the driver. A Pure V2V model (same backbone and V2V data, no decomposition or CoE) scores only 2.32, below OpenVE-Edit (2.49). In contrast, full data + CoE reaches 2.65, a +14.2% gain, and +74.1% on Creative Edit, a category with no direct V2V supervision, while adding only about 6% visual training tokens.
- Human study confirms gains are perceptible. A blinded pairwise human study on all 431 clips with 10 raters (Fleiss' κ = 0.721) prefers full data + CoE over V2V*-only + CoE, most clearly on Prompt Following (38% vs. 24%) and Consistency (28% vs. 9%), indicating the MLLM-judged gap between 2.57 and 2.65 is human-perceptible.
- Chain-of-Editing is contingent on chain-structured training. Applied to an I2I-only model, CoE degrades performance because the two-stage format is out of distribution. The V↔I↔I model benefits most, since its training already includes an intermediate edited frame that CoE exploits as an appearance anchor. CoE also improves the V2V*-only model (2.49 → 2.57), showing it generalizes as a training-free boost.
- Intermediate image quality propagates through the chain. Fig. 3 shows self-editing underperforms on categories requiring geometric transformation or novel content synthesis, while external editors (Qwen-Image-Edit and Nano Banana) reach similar overall scores with complementary category-level strengths, highlighting the modular I → I interface.
- TDF3D-RoPE is the larger standalone component gain. Both TDF3D-RoPE and CoE improve overall performance, and their combination achieves the best score of 2.65; TDF3D-RoPE contributes the larger standalone gain by establishing pixel-level spatial correspondence across shots.
- Temporal Position Randomization and CoE interact. Without TPR, a fixed temporal index creates a spurious structural correlation that prevents CoE's intermediate keyframe from serving as a genuine appearance anchor; only their combination reaches peak performance.
- Data sampling weights matter. Each data source contributes complementary supervision and performance is best when balanced; moderately upsampling V↔I↔I improves chain exposure, while overly large weights reduce paired video supervision. I2I performs best at the default ratio.
Methodology in Plain English
Instead of asking a model to learn video-to-video editing directly (which needs hard-to-get paired video data), the authors split the job into three smaller steps: extract a representative frame from the source video (Video → Image), edit that frame with the instruction (Image → Image), and then regenerate the whole video from the edited frame while keeping the original motion and structure (Image → Video).
Everything is flattened into one long interleaved sequence of text and visual tokens: source video tokens, source keyframe tokens, edited keyframe tokens, and the target video slot, each preceded by its own text prompt. The edited image pair therefore sits in the model's context as a visual example — a pixel-level blueprint of what the edit should look like — and every output frame is generated with that example in view.
Two design pieces make this work. First, a position encoding (TDF3D-RoPE) gives each token a 3D position built from two added parts: a "between-shots" component that keeps different segments distinct, and a "within-video" component that gives spatially corresponding patches in different shots the same horizontal and vertical positions. That means a patch in the edited image and the matching patch in an output frame line up exactly, so the edit propagates pixel-faithfully. Images used in training are also given randomized temporal positions so the model cannot lazily associate temporal index with image modality.
Second, at inference the chain runs as two diffusion stages (Chain-of-Editing): first produce the edited keyframe, then feed that keyframe — inserted directly in latent space without a VAE decode/re-encode cycle — into the video generation stage. Conditioning shots are given timestep 0 (clean) while only the target shot receives the active diffusion timestep, and the whole multi-shot sequence goes through the DiT in one forward call. Because the chain is just extra context, more image-editing iterations can be run at test time to trade compute for fidelity without retraining.
Training mixes four data sources: image editing pairs, video editing pairs reformatted into sub-chains, video-based in-context image editing data (where GPT-4o writes the instruction and Seedream 4.5 applies it per frame, without paired video supervision), and full-chain samples. All are supervised with a single flow-matching objective. The backbone is a 3B MM-DiT initialized from the same text-to-video pretrained backbone as VINCIE, with a video VAE inflated from SD3 (×8 spatial, ×4 temporal downsampling, 16 channels) and Flan-T5 as text encoder; only the DiT is trained and it is shared across all sub-tasks and both inference stages. Training used 32 H100 GPUs for about 150 hours (42k steps, lr 5×10⁻⁵) in two stages, at 256×256 then 640×480.
Why This Matters
Impact on research. The work reframes video editing as an in-context conditioning problem rather than a data-collection problem, showing that mature image editing priors can substitute for expensive paired video supervision. It also introduces a modular I → I interface, meaning future advances in image editing transfer to video without retraining, and it establishes Chain-of-Editing as a training-free test-time scaling mechanism.
Real-world applications:
- Content creation and post-production, where editors restyle footage (for example summer-to-snow scene stylization) or apply localized compositional effects such as fire, water, or wind without task-specific fine-tuning.
- Object removal and subject replacement in video, where the paper shows V2V direct inference produces flickering and ghosting while Chain-of-Editing eliminates this instability by anchoring to a reference image.
- Subtitle editing and background replacement for localization or branding of existing footage.
- Attribute and identity editing at scale, using user-supplied or externally generated image edit pairs as the demonstration.
Industry relevance. The approach substantially reduces dependence on large-scale paired video editing data, which is the stated fundamental bottleneck for video editing. The reliance on an off-the-shelf image editor for the intermediate step means product pipelines can plug in whichever image editing model improves, and the result (3.08 overall, 24% relative gain over OpenVE-Edit) positions this as competitive with open-source baselines and closer to the proprietary Runway Aleph.
Future Directions
- Geometric and novel-viewpoint editing. Camera Edit remains the weakest category at 2.06 versus Runway Aleph's 4.53, and the paper attributes this to geometric operations outside the scope of appearance editing. Extending the chain to handle viewpoint change is a clear open problem.
- Synthesizing content absent from the source. Local Add (2.27) and Background Change (2.55) both require generating content not present in the source video, which the paper flags as having room for improvement.
- Better intermediate editors and longer chains. The paper shows external editors (Qwen-Image-Edit, Nano Banana) outperform self-editing with complementary category strengths, and that chains can be extended by iterating image editing multiple times; characterizing the optimal test-time compute budget is left open.
- Reducing reliance on paired video data further. Since a Pure V2V model scores only 2.32 while full data + CoE reaches 2.65, further work could explore how little paired video supervision is needed, and how best to balance the data mixture beyond the default proportions tested.
Target Audience
Researchers and engineers working on diffusion-based video generation and editing, in-context or interleaved multimodal generation, and position-encoding or long-context modeling for diffusion transformers. It is also relevant to practitioners building content-creation tools who need video editing without large paired video datasets, and to readers interested in test-time scaling strategies that trade compute for output fidelity. The appendix-level implementation details (position formulas, staged context construction, data annotation with GPT-4o and Seedream 4.5) make it useful for teams attempting reproduction, while the heavy dependence on diffusion and RoPE terminology makes it less suitable for beginners.
Authors’ abstract
Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.