Research
VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction
Overview Research area: Computer vision, specifically physics-aware video editing, 4D/physical scene reconstruction, and rigid-body simulation. Technical level: Advanced. The paper combines open-vocab

- arXiv
- 2609.35134
- Published
- 2026-09-28
- Authors
- Conghan Yue, Yuanjie Chen, Yue Han, Ya Gao, Yunyan Xiao, WeiYao Zhang, Zhineng Chen
AI summary
Overview
- Research area: Computer vision, specifically physics-aware video editing, 4D/physical scene reconstruction, and rigid-body simulation.
- Technical level: Advanced. The paper combines open-vocabulary detection and segmentation, monocular 3D reconstruction, 6DoF motion priors, rigid-body constraint solving, differentiable-style simulation search, and trajectory-conditioned video diffusion.
- Scope (one sentence): The paper defines physical counterfactual video editing (PCVE), builds a training-free pipeline (VideoPhysEdit) that reconstructs an executable rigid-body scene from a source video and uses simulated counterfactual trajectories to guide video generation, and releases a paired synthetic benchmark (PCVE-RigidBench) with a new metric (Physical Edit Score).
What This Paper Is About
Existing video editing tools handle visual consequences of an edit, such as changing shadows, reflections, or occlusions, but they generally do not handle the physical consequences: how later motion and object-to-object interactions should change. The authors define this as physical counterfactual video editing (PCVE), where the system receives a source video, a physical edit stated in natural language, and the frame at which that edit is executed, and must output a video that keeps the original evolution unchanged before that frame and then shows the physically altered evolution afterward. The goal is to replace implicit "guess what should happen" generation with explicit physical reasoning: reconstruct a scene that can actually be simulated, apply the edit as an intervention inside that simulation, and use the simulated trajectories to steer the video generator.
Key Contributions
-
Formulation of PCVE as a unified task. The paper defines physical edits as three types — inserting or removing an object, modifying an object's motion state, and altering physical parameters of an object or the scene — and defines executing such an edit at a specified frame as a physical intervention.
-
VideoPhysEdit, a training-free pipeline. The method reconstructs an executable physical scene from a source video by organizing observations into stable intervals and transition episodes, initializing object states and physical parameters from geometric and rigid-body constraints, then refining them so that simulation reproduces the observed motion and interactions. The edit is grounded in this scene, applied as an intervention, and the resulting simulated trajectories plus an edited reference image guide counterfactual video generation. Seven numbered modules (Stages 1–7) make up the pipeline.
-
PCVE-RigidBench. A synthetic benchmark containing 20 rigid-body scenes spanning impacts, rebounds, rolling, sliding, falls, and collision chains, simulated in PyBullet and rendered in Blender. It contains 129 editing tasks, each pairing a source video and a physical edit with a counterfactual target video and physical ground truth, covering insertion/removal and changes to initial velocity, mass, friction, or restitution, with interventions applied either at the first frame or partway through. Each task also includes a quantitative description (execution frame plus numerical or spatial change) and a qualitative description (direction and coarse timing).
-
Physical Edit Score (PES). A metric measuring the reduction in trajectory error against the counterfactual target relative to the unchanged source video, defined as PES = max(1 − Σ TE_i^pred / Σ TE_i^null, −1), evaluated only over objects present in the source video whose motion or presence changes after the intervention. A score of one means zero scored error relative to the counterfactual target, zero means no improvement over the unchanged source, and a negative score means worse than that baseline.
Main Findings
-
Best physical edit accuracy on PCVE-RigidBench. VideoPhysEdit achieves PES 0.376, the only positive score among the compared methods, and the paper reports it reduces trajectory error by 54.0% relative to the strongest competing method. It also achieves the highest Mask IoU (0.421) and the lowest trajectory error (66.70) in Table 1.
-
Competing methods score at or below zero. On Table 1, VACE scores PES −0.042 (TE 146.26, Mask IoU 0.273), Ditto −0.120 (TE 149.94, Mask IoU 0.264), MiniMax H3 −0.096 (TE 152.00, Mask IoU 0.250), and Seedance 2.5 −0.087 (TE 144.99, Mask IoU 0.231). The unchanged-source "No edit" baseline has PES 0.000 with TE 143.13 and Mask IoU 0.289.
-
Smaller losses on other affected objects. For objects whose motion changes as a downstream consequence of the edit, VideoPhysEdit is again the only method with a positive PES, reaching 0.276 (reported in Appendix Table 9).
-
Competitive visual fidelity. In Table 1, VideoPhysEdit records PSNR 27.51, SSIM 0.925, LPIPS 0.104, CLIP 0.929, and FVD 182.46, which the paper describes as close to Seedance 2.5 (PSNR 28.85, SSIM 0.928, LPIPS 0.080, CLIP 0.932, FVD 246.67) while achieving the best FVD.
-
Stronger results on object removal against a removal specialist. On the removal subset (Table 2), VideoPhysEdit scores PES 0.633, TE 52.99, Mask IoU 0.496, PSNR 26.31, SSIM 0.921, LPIPS 0.110, CLIP 0.926, FVD 230.29, reducing TE by 37.0% relative to VOID (PES 0.394, TE 84.07). VOID obtains the highest PSNR (29.22) in that table, which the paper attributes to its preservation of unaffected source regions; VideoPhysEdit achieves the best SSIM and FVD there.
-
Reconstruction quality propagates to counterfactual quality. In the stage analysis (Table 3), Mask IoU goes from 0.883 to 0.887 between Stage 3 and Stage 4; optimizing Stage 5 improves the factual Mask IoU from 0.375 to 0.678 and Stage 6 PES from 0.269 to 0.403; and Stage 6 to Stage 7 changes PES from 0.398 to 0.412, TE from 63.39 to 64.76, and Mask IoU from 0.391 to 0.412.
-
Visually plausible generation is not the same as a correct physical edit. The paper reports that baseline methods often retain the source motion or miss later consequences of the edit, and that explicitly describing downstream consequences in the instruction does not yield consistent improvements (Appendix C.5).
-
Qualitative real-video results. The paper reports that VideoPhysEdit applies to real-world scenes and better depicts downstream motion and interactions than the compared methods, including examples of mass changes altering post-impact marble motion, reduced toy-car speed preventing a later collision, projectile velocity edits, and removal before a collision.
-
Not reported in the supplied content: runtime, peak GPU memory, and the specific numbers from the simulation-search ablation are referenced as being in Appendix C but their values are not present in the provided text. The paper also notes two scenes in Appendix C.2.3 for which the pipeline produces no valid edited videos.
Methodology in Plain English
The pipeline is a sequence of seven stages and requires no training; all pretrained components use released checkpoints.
-
Find and track the relevant objects. A vision-language model (Qwen3-VL) reads the source video and the edit description, decides which object categories are involved, and identifies referenced objects. Grounding DINO detects instances in the first frame, and SAM 2 propagates per-object masks through the video while keeping consistent identities.
-
Describe the motion in simple pieces. Point tracks (CoTracker3) combined with the masks give 2D position, orientation, and a reliability score. The observed motion is segmented into stable intervals explained by simple models (constant-acceleration translation, rotation about a fixed axis, and stopping curves) and transition episodes around changes in motion, using a dynamic program with change detection.
-
Pick reference frames. A canonical frame is selected to serve as a shared world origin, and one motion anchor frame is chosen per stable interval; selection is driven by scores for boundary truncation, out-of-transition status, temporal stability, projected separation, texture clarity, and the number of stationary objects.
-
Rebuild geometry. From the canonical frame, the method estimates camera parameters and point clouds (VGGT with SuperGlue), recovers static scene planes, and fits each object with a textured sphere or box. Object scale, rotation, and translation are optimized with a placement loss combining 3D correspondence, silhouette IoU and Dice, regularization, and support consistency. Anchor scenes are then aligned into the same coordinate system using the static background.
-
Lift motion into 3D. The stable intervals, transition episodes, and anchor scenes are used to fit a continuous 6DoF motion prior against the source masks — simple models within stable intervals, boundary-constrained curves across transitions — allowing velocity changes at inferred impacts.
-
Solve for physical parameters (physical inversion). Collision proxies are built from reconstructed geometry, and rigid-body constraints are derived from support relations and the motion prior: stable intervals constrain force balance, friction, rolling, and energy; contact events constrain momentum balance, restitution, and friction. This initializes the parameter set η = (Θ, s₁, G^col). The parameters are then refined by simulation search: the loss L_mask^(h) = (1/|Ω_h|) Σ [1 − IoU(Ŝ, S)] between simulated and observed masks is minimized, adding one stable interval or transition episode at a time while keeping the best simulation plus up to two distinct alternatives, then re-refining the best saved simulation over all frames.
-
Apply the edit and generate video. The edit instruction is parsed into an Add, Delete, or Set operation (templates for benchmark instructions, a vision-language model otherwise), object references are bound to persistent identities, and relative quantities are resolved in the reconstructed scene. At the execution frame, the operation is applied to the factual state and the scene is re-simulated to obtain counterfactual states, which are converted into projected point trajectories plus an edited reference image. A pretrained video generator (Wan-Move) produces the continuation conditioned on those trajectories, the reference image, and a scene prompt.
Supporting tools include PyBullet for simulation, ObjectClear for Delete reference images, and Insert Anything and Cube3D for inserted-object appearance and geometry.
Why This Matters
Impact on research. The paper reframes physics-aware editing from "make the picture look physically plausible" to "make the simulation explain the observation, then let it drive generation." It provides a task definition, a paired benchmark, and a metric where none existed for this setting, and argues through its stage analysis that reconstruction accuracy is causally connected to counterfactual accuracy. The negative PES scores of strong commercial models and open-source editors provide a concrete signal that current generation-based approaches do not solve this problem.
Real-world applications:
- Visual effects and post-production: inserting, deleting, or re-parameterizing objects in a shot and having the resulting collisions, rebounds, and cascading interactions rendered consistently rather than hand-animated.
- Robotics and embodied AI: editing recorded manipulation or collision scenes to ask "what if mass, friction, or initial speed were different" as a source of controlled scenario variation.
- Content creation and advertising: producing alternate takes of a physical scene (a different bounce, a removed obstacle) from a single shot.
- Analysis and forensics: testing counterfactual explanations of an observed event — for example, whether a changed parameter could have prevented a collision.
Whether any of these are deployed is not reported in the paper; the evaluation is confined to PCVE-RigidBench plus qualitative real-video comparisons.
Industry relevance. The pipeline is training-free and assembles existing off-the-shelf checkpoints, which lowers the barrier to adoption for studios and product teams that already have video generation models in their stack. At the same time, the paper's findings imply that adding physical correctness to commercial video generators is not something that emerges automatically from scaling visual quality — an explicit physical scene is needed.
Future Directions
- Camera motion. The current setting works with the source camera; the paper explicitly lists camera motion as future work.
- Richer geometry and material models. Future work targets richer geometry, moving beyond simple sphere and box visual meshes and rigid-body scenes.
- Articulated and actively controlled agents. Extending the physical models to articulated bodies or agents that act on their own, rather than passive rigid bodies.
- Robustness and scalability. The paper notes failure cases (two scenes producing no valid edited videos) and describes resolution of ambiguous observations only progressively across stages; improving reliability under occlusion, boundary truncation, and missing evidence remains open.
Target Audience
Researchers and graduate students working on video editing, physics-aware generation, 4D reconstruction, and simulation-based scene understanding. It is also relevant to practitioners in visual effects, embodied AI, and video generation who need edits to respect downstream motion and contact. Readers should be comfortable with rigid-body dynamics, camera and 3D reconstruction terminology, and video diffusion conditioning; the paper is not a beginner-level introduction to video editing.
Authors’ abstract
Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored. We formulate this problem as physical counterfactual video editing (PCVE), which aims to generate a counterfactual video depicting the resulting motion and interactions given a source video, a physical edit, and its execution frame. PCVE is challenging because it requires understanding scene physics and inferring the downstream motion and interactions induced by a physical intervention, while paired factual and counterfactual data and dedicated evaluation metrics are lacking. We introduce VideoPhysEdit, a new training-free pipeline for PCVE in rigid-body scenes. It makes physical reasoning explicit through a novel physical scene reconstruction method that recovers a scene reproducing the observed motion and interactions under simulation, enabling the pipeline to apply physical edits as interventions and use the resulting trajectories to guide counterfactual video generation. We further construct PCVE-RigidBench, a synthetic benchmark with paired source and counterfactual target videos and physical ground truth, and introduce the Physical Edit Score. VideoPhysEdit achieves substantially higher physical edit accuracy than open-source methods and commercial models while maintaining competitive visual fidelity. Its Physical Edit Score is 0.376, the only positive score among the compared methods. Qualitative comparisons on real videos further show that VideoPhysEdit applies to real-world scenes and better depicts the downstream motion and interactions induced by the edits than the compared methods. Code: https://github.com/Hammour-steak/VideoPhysEdit