Skip to content
AI.info

Research

4Director: Controlling Video World Models with Rigid 3D Geometry

Overview Research area: Computer vision and generative video — specifically controllable video world models, where a pretrained video generator is steered by an explicit 3D/4D scene representation rat

4Director: Controlling Video World Models with Rigid 3D Geometry
arXiv
2610.02160
Published
2026-10-01
Authors
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

AI summary

Overview

Research area: Computer vision and generative video — specifically controllable video world models, where a pretrained video generator is steered by an explicit 3D/4D scene representation rather than text or 2D cues.

Technical level: Advanced. The paper assumes familiarity with latent video diffusion transformers, flow-matching objectives, SE(3) rigid transformations, and monocular geometry lifting. The prose below is plain-language, but the underlying method is specialist work.

Scope: The paper introduces 4Director, a video world model that generates videos from a single input image under user-prescribed 3D camera and object trajectories, by rendering a rigid 3D scene to a depth video and injecting it into a pretrained video generator (arXiv:2610.02160v1, Stability AI and University of Illinois Urbana-Champaign).

What This Paper Is About

Generating a video from one image is now possible, but directing it — telling the system exactly where the camera goes and how each object moves in 3D — is not. Existing controls are either 2D cues (boxes, dragged paths, point tracks) that are ambiguous in depth and rotation, or 3D cues (trajectories, 3D boxes, lifted point/Gaussian proxies) that carry no complete surface, so whatever geometry a turn or new viewpoint reveals is hallucinated independently in every frame. The goal of 4Director is to let a user prescribe full 3D motion for the camera and every object while a generative model supplies appearance, illumination, and non-rigid detail such as a walking figure's swinging limbs.

Key Contributions

  1. An explicit 4D scene representation with complete object geometry. Each object is reconstructed once from the input image as a complete canonical mesh and moved by one rigid transformation per frame, sharing a single coordinate frame with a background point cloud and the camera. Because the mesh has a surface on every side, geometry revealed by a turn or new viewpoint is rendered rather than regenerated.
  2. The Motion Adapter. A trainable DiT-style branch that injects a depth rendering of the rigid scene into a pretrained video generator (Wan2.1-VACE-14B), so the geometry fixes camera and object motion while the generator supplies appearance, illumination, non-rigid dynamics, and the background a camera move reveals.
  3. RealCOD-Rigid. A dataset of 20,774 clips annotated with complete object meshes, per-frame rigid transformations, and camera trajectories, produced by an automatic pipeline that recovers the rigid 3D scene of each monocular clip. Clips come from the 25,318 source clips of RealCOD-25K.
  4. Identity-Gated IoU (IG-IoU). A new object-control metric that credits mask IoU only in frames where a vision–language judge confirms the generated object is still the same object, addressing the fact that plain mask IoU rewards correct placement even when identity is lost.

Main Findings

  • 4Director leads on every quantitative metric in the joint camera-and-object control comparison (Table 1a). It records FID 44.1, FVD 370.4, CLIP-SIM 31.8, rotation error 3.65 degrees, translation error 0.122, recognition rate 94.9, and IG-IoU 60.4, versus the next-best scores of FID 45.8 and FVD 405.4 (SymphoMotion), 31.7 CLIP-SIM (VerseCrafter), 3.89 degrees and 0.132 (SymphoMotion), and 94.8 recognition rate and 54.8 IG-IoU (VerseCrafter).
  • The largest margin is in object control. IG-IoU rises from 54.8 for VerseCrafter, the next best, to 60.4, described as a 10.2 percent relative gain at an essentially unchanged recognition rate (94.9 against 94.8). The authors attribute this to the complete mesh fixing object extent and orientation.
  • Visual quality improves by 3.7 percent on FID and 8.6 percent on FVD over the next-best method (SymphoMotion), while matching VerseCrafter on CLIP-SIM. Translation error falls by 7.6 percent relative to the next best.
  • 2D-trajectory control degrades badly. MotionCtrl, driven by 2D trajectories, keeps the object recognizable in only 32.6 percent of frames, so its IG-IoU drops to 18.8.
  • VBench-I2V and the user study both favor 4Director. It leads on all eight applicable VBench-I2V dimensions, with margins over the next best staying within 2 points (for example motion smoothness 98.8 versus 98.7, dynamic degree 96.0 versus 94.0, imaging quality 64.6 versus 64.2). In the user study with 20 participants rating nine cases from 1 to 5, it scores 4.6 to 4.8 against at most 3.9 for the next best, with the largest margin on control accuracy (4.82 against 2.39).
  • Only 4Director follows a full turn of the source object in qualitative comparisons. On a duck, Perception-as-Control and VerseCrafter keep the original heading and MotionCtrl and SymphoMotion turn only slightly; on a tractor, SymphoMotion and VerseCrafter turn it only halfway, while Perception-as-Control loses it and MotionCtrl blurs it. The tractor also passes behind a fence, an occlusion 4Director reproduces while SymphoMotion and VerseCrafter keep it in front.
  • Objects can leave and re-enter the view without losing identity. In the camel example, 4Director leaves the frame empty while the camel is away, fills in the background it occupied, and brings back the same camel; MotionCtrl, SymphoMotion, and VerseCrafter keep the camel in frame throughout, and Perception-as-Control fails to preserve its shape on return.
  • Ablation shows both ingredients of the representation matter. Replacing rigid 3D geometry with a 2D box drops IG-IoU to 48.5, a 3D box (depth and orientation but no surface) to 50.3, and a 3D mesh (complete surface but no rotation) to 53.0, while recognition rate stays between 91.9 and 95.0 for every control — meaning the loss is mainly in placement, not identity. The four controls rank in the same order on FID, FVD, and IG-IoU.
  • Stated limitation. The representation controls motion only at the rigid-body level: an articulated or deforming object moves as one whole, so finer motion such as running, jumping, or limb movement cannot be prescribed and is decided by the generator, which may even keep a mostly articulated object in its first-frame pose.

Methodology in Plain English

The system takes four user inputs: an image, a text prompt, a camera trajectory, and one rigid trajectory per object the user marks.

Building the scene. Depth and intrinsics are estimated with MoGe-2, and SAM 2 turns user clicks on each object into a mask. The background — the image outside the masks — is back-projected into a static point cloud in the camera's reference frame. Each marked object is reconstructed by Pixal3D into a complete canonical mesh, including sides the image never shows, and aligned to its mask and depth. The user then prescribes trajectories as keyframes in a 3D viewer, with one rigid transformation per object per frame and one world-to-camera matrix per frame. A new object can be inserted as a mesh reconstructed from a separate image, rendered into the input image as the first frame.

Turning the scene into a control signal. For each frame the point cloud and the placed meshes are projected into a depth map taking the nearest surface at each pixel. The result is a depth video — 8-bit gray with near at 255 and far at 0, invalid pixels set to 128 — that encodes only viewpoint, object rigid transformations, and occlusion. Appearance, illumination, non-rigid dynamics, and background beyond the image are absent by construction.

Training the adapter. The Motion Adapter is a DiT-style branch built on Wan2.1-VACE-14B, initialized from released VACE weights, with eight context blocks — one for every fifth DiT block — that cross-attend to the prompt and inject linearly projected hints. The VAE, umT5 text encoder, and DiT serve as a frozen generic video prior. Training pairs come from RealCOD-Rigid: the control video rendered from a clip's recovered rigid scene is the condition, the original clip is the target, its first frame is the input, and its text description is the prompt. Because the control contains the rigid part of each surface point's motion but not the non-rigid residual, the adapter must learn to supply what the control omits. Training uses a standard flow-matching objective.

Building the training data. No dataset supplies clips paired with rigid 3D scenes, so the authors recover them automatically in four stages from RealCOD-25K clips that come with text prompts and SAM3 masks for one or two objects: (1) MegaSaM with MoGe-2 and UniDepthV2 estimates per-frame depth, the intrinsic matrix, and the camera trajectory; (2) Pixal3D reconstructs a canonical mesh from the object's first-frame crop, and a robust similarity fit places it, defining the first frame's transformation as the identity; (3) TAPIP3D tracks 3D points on each object, and the authors isolate the object's "stable core" — the tracks that move together rigidly — by alternating a robust fit with inlier selection, then fit each frame on that core alone; (4) the placed meshes are rendered into a control video. Clips where objects fail a 3D gate (residual thresholds of median at most 0.15, p90 at most 0.30, p95 at most 0.50 after normalizing by object diameter) or a silhouette gate (mesh-mask IoU at least 0.5) are excluded.

Evaluation. The backbone, the setting, and the metrics are all reported. The full adapter has 3.049 B trainable parameters and the LoRA variant about 15.3 M. Both train at 81 × 832 × 480 for three epochs on 3 × 8 GPUs with global batch 24, peak learning rate 5 × 10⁻⁵, 25 warmup steps, weight decay 0.01, seed 42. Inference uses 20 flow-matching sampling steps, classifier-free guidance 5.0, sigma shift 5.0, seed 42, 832 × 480, 81 frames at 16 fps, tiled VAE decoding, and hint scale 1.0. Evaluation uses the same 100 clips of RealCOD-25K excluded from training, comparing against MotionCtrl, Perception-as-Control, SymphoMotion, and VerseCrafter, each given the same motion converted into whatever control it expects. IG-IoU is computed over N = 16 frames spaced uniformly across frames where the object is visible, with SAM3 masks and Qwen3-VL as the identity judge; a frame failing the gate contributes zero but stays in the denominator.

Why This Matters

The paper's distinctiveness is that it treats scene geometry as fixed before generation and complete, rather than as a hint the generator is free to reinterpret. Image-plane controls cannot distinguish depth from size change or rotation from a silhouette change; lifted point and Gaussian proxies cover only the visible shell, so anything a turn reveals is invented anew in each frame. By giving every object a full mesh with a single rigid transform per frame, 4Director makes revealed geometry deterministic and makes identity preservation across occlusion and viewpoint change a property of the representation rather than a hope about the generator. The IG-IoU metric addresses a real gap in evaluation, where mask IoU alone can reward a correctly placed but wrong object.

Real-world applications:

  • Film and episodic production: previz and shot design where a director stages camera and actor trajectories in 3D before committing to a shoot.
  • Visual effects and virtual production: generating plates that follow a prescribed camera and object path, including occlusions such as a vehicle passing behind a fence.
  • Advertising and commercial content: rapid iteration on product shots, object insertion, and camera moves from a single still.
  • Interactive media and game prototyping: quickly visualizing how a scene reads from a designed camera path with moving objects.
  • Simulation and synthetic data: producing camera-consistent videos with known 3D object geometry for downstream training.

Industry relevance: The work comes from Stability AI with University of Illinois Urbana-Champaign, builds on a pretrained open video backbone (Wan2.1-VACE-14B), and targets professional video production workflows directly. Its explicit editing interface — insert an object from a reference image, draw trajectories, render — is closer to how production tools actually work than prompting alone.

Future Directions

  • Articulated geometry. The authors' stated next step is extending the rigid 3D scene with articulated parts so that running, jumping, and limb movement can be prescribed instead of left to the generator, which may keep a largely articulated object in its first-frame pose.
  • Scene scale. The recovered coordinate frame is described as a shared coordinate system of arbitrary scale, not a metric reconstruction; whether metric scale can be recovered reliably is not reported and remains open.
  • Authoring effort. Trajectories are authored as keyframes by a user in a 3D viewer, and each object must be clicked and reconstructed individually. How well this interface scales to crowded scenes with many interacting objects is not reported.
  • Ordering and symmetry. The paper notes that a 180-degree turn of a front–back symmetric object is not penalized by MaskIoU and is only shown qualitatively, which leaves evaluation of such cases open.
  • Compute and latency. The paper reports training hardware (3 × 8 GPUs, 3.049 B trainable parameters) but does not report inference runtime or generation latency, which matters for any interactive deployment.

Target Audience

Researchers and engineers working on controllable video generation, video world models, and diffusion-based image-to-video systems; graphics and 3D-vision researchers interested in coupling reconstruction with generative rendering; and technical practitioners in film, VFX, and content production evaluating 3D-driven control interfaces. Readers without a background in diffusion transformers and rigid-body geometry will find the high-level framing accessible, but the method and training details are advanced.

Authors’ abstract

Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

Read the original paper