Skip to content
AI.info

Research

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

Overview Research area: Computer Vision — training-free video editing with multimodal diffusion transformers (MM-DiTs). Technical level: Intermediate. The paper assumes familiarity with diffusion/flow

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
arXiv
2609.04190
Published
2026-09-03
Authors
Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed, Nabeel Bashir, Xiaona Zhou, Tianjiao Yu, Vedant Shah, Ismini Lourentzou

AI summary

Overview

  • Research area: Computer Vision — training-free video editing with multimodal diffusion transformers (MM-DiTs).
  • Technical level: Intermediate. The paper assumes familiarity with diffusion/flow-matching models and transformer attention, but its core ideas (reuse the previous frame's attention state, copy matched patches from an anchor frame, blend latents softly) are describable in plain terms.
  • Scope: A single training-free framework, EditVid, that performs instruction-guided and subject-guided video editing on top of a frozen image editor, evaluated on FiVE, IVEBench, a curated 50-video set, and a human preference study.

What This Paper Is About

Applying a strong image editor to each video frame independently produces flicker, identity drift, and unwanted background changes, because nothing ties the frames together. Training dedicated video editors is expensive, and existing training-free video methods either lag in edit quality or depend on auxiliary inputs such as depth maps and masks. The goal of this paper is to get temporally consistent, diverse video edits out of a frozen image-editing MM-DiT, with no video-specific training.

Key Contributions

  1. EditVid, a training-free framework that supports both instruction-guided and subject-guided video editing by treating an MM-DiT-based image editor as a strong prior across style transfer, attribute modification, object insertion, part-level editing, and subject replacement.
  2. A RoPE-aware local–global decomposition that identifies temporal context as the key factor in reusing MM-DiT representations: adjacent-frame key–value memory for short-range coherence, and confidence- and cycle-consistent token transfer for long-range preservation.
  3. Soft latent blending that derives continuous, timestep-dependent preservation weights from the discrepancy between source and edited trajectories, keeping instruction-irrelevant content intact while letting the requested edit dominate elsewhere.
  4. Comprehensive evaluation — quantitative, human, robustness, and cross-backbone — showing strong temporally consistent editing without video-specific training.

Main Findings

  • FiVE edit correctness: EditVid reaches 78.16 FiVE-Acc, versus 58.95 for the strongest evaluated training-free baseline (FlowDirector). Per-metric: FiVE-YN 69.45, FiVE-MC 86.87, FiVE-U 88.78, FiVE-∩ 67.54.
  • FiVE baseline spread: Other training-free comparisons report lower FiVE-Acc — StreamEdit (SF) 51.94, DMT 48.42, Wan-Edit 46.97, Pyramid-Edit 43.84, AnyV2V 38.02, VideoGrain 37.23, TokenFlow 27.43, VidToMe 26.77.
  • Traditional FiVE metrics: EditVid obtains PSNR 25.97, LPIPS 218.71, MSE 36.37, SSIM 83.31, CLIP_S 28.83, CLIP^edit_S 22.40, NIQE 4.17, MFS 85.31, with Structural Distance 17.10. The paper states these are the strongest among image/hybrid methods without auxiliary inputs for reconstruction, perceptual quality, and semantic alignment.
  • IVEBench: EditVid scores Total 0.6847, Quality 0.8114, Compliance 0.4587, Fidelity 0.7839 — the best training-free results in Total score, Instruction Compliance, and Video Fidelity, second overall in Total score (behind Kiwi-Edit-5B at 0.6936), and the best Video Fidelity across all methods.
  • Curated evaluation: On a 50-video set (24 general video-editing, 26 subject-guided), EditVid leads five of six criteria — subject-guided Prompt Following 7.769, Edit Quality 7.846, Background Consistency 9.558; general-video Prompt Following 9.458, Edit Quality 9.000, Background Consistency 9.500 (second-highest on that last criterion, behind StreamEdit's 9.792).
  • Human preference: A blind study with 22 participants over 20 video examples produced 440 participant–video evaluations and 1,760 criterion-level selections. EditVid received 51.8% overall preference against seven training-free baselines (StreamEdit 9.1, WANEdit 7.7, TokenFlow 7.7, VidToMe 7.3, FlowDirector 6.8, PyramidEdit 5.7, AnyV2V 3.9), plus 53.2% on prompt following, 49.3% on visual quality and temporal consistency, and 50.0% on background preservation.
  • Short memory beats long memory: On a 30-video FiVE subset, using the immediately preceding frame gives 77.50 FiVE-Acc and 91.63 MF-S at 4.096 s/frame and 44.65 GB, outperforming no temporal memory (72.50 / 91.39 / 3.495 s / 38.37 GB), two previous frames (72.50 / 89.76 / 4.146 s / 54.98 GB), and all previous frames (62.50 / 49.81 / 5.490 s / 90.16 GB).
  • Module ablation: The full model scores 81.67 FiVE-Acc and 84.88 MF-S (×10²). Removing soft latent blending drops FiVE-Acc to 80.00 (MF-S 84.59); removing spatio-temporal attention gives the same 80.00 FiVE-Acc and the lowest MF-S at 84.55; removing global token injection leaves FiVE-Acc at 81.67 but lowers MF-S to 84.72.
  • Correspondence filtering matters: Under viewpoint, occlusion, non-rigid, and structural stress, unfiltered anchor injection scores 72.53 overall MF-S — worse than removing global injection entirely (75.65) — while the full method reaches 76.23 and is best on viewpoint change (71.96), occlusion (66.27), and non-rigid motion (78.96).
  • Preservation metrics can mislead: FlowDirector and StreamEdit score well on traditional FiVE metrics by preserving the source conservatively, yet they apply the requested color change only partially, whereas FiVE-Acc rewards actual edit execution.

Methodology in Plain English

EditVid edits a video inside the latent space of a frozen image-editing MM-DiT (FLUX.2-Klein-9B, four denoising steps) and adds three mechanisms, none of which involve training.

First, sparse causal memory: when processing a frame, the model's attention is allowed to look at the key and value states of the immediately preceding frame, so motion and appearance flow forward one frame at a time. Because the context is exactly one frame long, the computational footprint does not grow with video length, and the method can be rolled out over long videos in chunks of 36, 24, or 16 frames depending on resolution.

Second, correspondence-based global token injection: before editing, the source video is inverted and patch tokens are extracted from the penultimate double-stream block at correspondence time t_corr = 0.25. Using frame 1 as the anchor, the system matches anchor patches to patches in every other frame by cosine similarity, keeping only matches that pass a similarity threshold and a cycle-consistency radius, and randomly dropping a fraction of the rest. During denoising, at a selected subset of injection layers, matched target-frame tokens are overwritten with the corresponding anchor-frame tokens. The injection happens after attention, so distant tokens never have to interact through RoPE-modulated cross-frame attention.

Third, soft latent blending: at each denoising step, the L1 difference between the edited latent and the inverted source latent produces a per-token map, normalized between low and high quantiles and optionally Gaussian-smoothed into a soft mask. Regions close to the source trajectory are pulled back toward the source, while regions with large latent changes follow the target edit, preserving backgrounds without crude masking.

Why This Matters

  • Research impact: The paper argues that modern image-editing MM-DiTs are an underexploited foundation for video editing, and that the local–global split is a general design principle — adjacent attention reuse where geometric continuity is reliable, correspondence-guided feature transfer where it is not. It also challenges reliance on preservation-style metrics, which can reward conservative edits that fail to follow instructions.
  • Real-world applications:
    • Attribute and style editing of existing footage, such as recoloring objects or restyling a scene consistently across many frames.
    • Part-level editing and object insertion for post-production and visual effects work without retraining a model per shot.
    • Subject replacement and personalized edits, where a reference subject's identity must be preserved throughout a clip.
    • Content creation and localization workflows where a single instruction must be applied uniformly to a long video.
  • Industry relevance: A training-free framework that runs on a frozen image editor lowers the cost of adopting video editing — no video-specific training data, no separate image-to-video propagation model, and bounded context that keeps computation from scaling with video length. The reported 4.096 s/frame and 44.65 GB figures for the preferred configuration also make the practical cost of deployment concrete.

Future Directions

  • Longer and higher-resolution video: The method is chunked at 36, 24, or 16 frames with cached cross-chunk memory; how quality holds up on much longer inputs is not established in the reported content.
  • Correspondence quality under hard motion: Unfiltered injection underperforms removing global injection entirely, so improving matching under strong viewpoint change, occlusion, and non-rigid motion remains an open problem.
  • Efficiency: Memory use climbs from 38.37 GB (no temporal memory) to 44.65 GB (immediately preceding frame) on the ablation setup, and the paper reports runtime and memory but no optimization of either.
  • Evaluation breadth: The curated human study and VLM evaluation cover 50 videos and 20 user-study examples; extending to broader edit taxonomies and more demanding subject-guided scenarios is a natural next step.

Target Audience

Researchers and practitioners in generative video editing, diffusion and flow-matching model design, and multimodal transformer architectures, especially those interested in inference-time techniques that avoid training. It is also relevant to applied engineers and creative-tool developers who need practical, cost-bounded video editing, and to readers studying how attention-state reuse and correspondence-based feature transfer interact with positional encodings such as RoPE.

Authors’ abstract

Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8\% overall preference for EditVid over 7 competing methods.

Read the original paper