Research
SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time
Overview Research area: Computer vision — generative rendering with video diffusion models, novel view synthesis, and space–time (4D) disentanglement for dynamic scenes. Technical level: Advanced. The

- arXiv
- 2512.25075
- Published
- 2025-12-31
- Authors
- Zhening Huang, Hyeonho Jeong, Xuelin Chen, Yulia Gryaditskaya, Tuanfeng Y. Wang, Joan Lasenby, Chun-Hao Huang
AI summary
Overview
Research area: Computer vision — generative rendering with video diffusion models, novel view synthesis, and space–time (4D) disentanglement for dynamic scenes.
Technical level: Advanced. The paper assumes familiarity with latent video diffusion models, Diffusion Transformers (DiT), 3D VAEs, camera extrinsics, and positional/rope embeddings. The problem statement itself is accessible to a general reader.
Scope: The paper proposes SpaceTimePilot, a video diffusion model that independently controls camera viewpoint and motion timing when re-rendering a single monocular dynamic video.
What This Paper Is About
A video is a 2D projection of a changing 3D world, and its appearance is governed by two separate factors: where the camera is (space) and how far the scene's motion has progressed (time). Existing camera-control video models keep time strictly monotonic, and existing 4D multi-view models only produce discrete sparse frames rather than continuous video. SpaceTimePilot's goal is to take one observed video of a dynamic scene and synthesize new videos in which the camera trajectory and the motion timeline can each be changed freely and independently — for example reverse playback, slow motion, or bullet time viewed from a new angle.
Key Contributions
- SpaceTimePilot, presented as the first video diffusion model that disentangles spatial and temporal factors to enable continuous, controllable novel view synthesis as well as temporal control from a single video.
- A temporal-warping training strategy that repurposes existing multi-view datasets to simulate diverse temporal variations (reversal, acceleration, freezing, segmental slow motion, and zigzag motion), so the model learns temporal control without explicitly constructed video pairs captured under different temporal settings.
- A precise camera–time conditioning mechanism: a dedicated "animation time" representation (sinusoidal embeddings of original frame indices compressed by two 1D convolution layers) plus source-aware camera conditioning that injects both source and target camera poses into the diffusion model.
- The Cam×Time dataset, a synthetic Space and Time full-coverage rendering dataset built in Blender: 180k videos rendered from 500 animations across 100 scenes and three camera paths, with source time 1:120 and target times drawn from {1, 2, …, 120}¹²⁰. Part of it is held out as a test set intended to serve as a benchmark.
Main Findings
- Temporal control wins across the board. On the withheld Cam×Time test split (50 scenes with dense full-grid trajectories), SpaceTimePilot achieves PSNR 21.75 / 20.87 / 20.85 (direction / speed / bullet time) and averages PSNR 21.16, SSIM 0.7674, LPIPS 0.1764. Baseline ReCamM+preshuffled averages PSNR 15.52, SSIM 0.6213, LPIPS 0.4529; ReCamM+jointdata averages PSNR 17.86, SSIM 0.7250, LPIPS 0.3073.
- Joint training with static-scene data is not enough. Following prior practice (training ReCamMaster with additional static-scene datasets) improves over frame shuffling, particularly in the bullet-time category, but a single temporal control pattern remains insufficient for robust temporal consistency.
- Visual quality is comparable rather than dominant. On VBench over 1800 generated videos, across six dimensions, SpaceTimePilot scores ImgQ 0.6486, BGCons 0.9199, Motion 0.9947, SubjCons 0.9325, Flicker 0.9781, Aesthetic 0.5315. It is highest on image quality, motion, and subject consistency, but its aesthetic score (0.5315) is below ReCamMaster (0.5332) and ReCamM+Aug (0.5385), and its flicker score (0.9781) is below the other three methods (0.9816, 0.9825, 0.9788).
- Camera accuracy improves substantially, including from an arbitrary first frame. On a real-world 90-video evaluation set built from OpenVideoHD with 20 camera trajectories per method (10 starting from the same pose as the source, 10 from different poses; 1800 generated videos total), SpaceTimePilot reaches RelRot 2.71, RelTrans 0.33, AbsRot 5.63, AbsTrans 0.34, Rot 4.09, RTA@15 35.19% and RTA@30 54.44%. TrajectoryCrafter scores 5.94 / 0.50 / 6.93 / 0.52 / 9.76 / 22.96% / 25.93%, and ReCamMaster 4.26 / 0.32 / 10.08 / 0.34 / 7.49 / 7.61% / 10.20%. ReCamMaster's RelTrans of 0.32 is marginally lower than SpaceTimePilot's 0.33; on every other listed metric SpaceTimePilot is best.
- More augmentation without source-camera conditioning hurts. ReCamM+Aug (retrained to allow non-identical first frames) yields higher errors (RelRot 3.66, AbsRot 11.74, Rot 13.88, RTA@15 3.89%, RTA@30 5.93%) than original ReCamMaster, while adding the source camera signal c_src gives the best overall performance — suggesting that without explicit reference to the source trajectory, exposure to more augmented videos with differing initial frames confuses the model.
- The compressor design matters. In the time-embedding ablation, uniform sampling (sinusoidal embeddings at the latent frame level) reaches PSNR 14.10 / SSIM 0.5981 / LPIPS 0.5039; 1D-Conv reaches 14.75 / 0.6134 / 0.4878; 1D-Conv + joint data reaches 15.41 / 0.6252 / 0.4830; 1D-Conv + Cam×Time reaches 21.16 / 0.7674 / 0.1764. Quality-wise, uniform sampling produces noticeable artifacts and an MLP compressor causes abrupt camera motion, whereas the 1D convolution locks animation time while allowing smooth camera movement.
- Qualitative disentanglement. In the shown comparison, only SpaceTimePilot correctly synthesizes both the camera motion and the animation-time state under reverse playback plus pan-right. TrajectoryCrafter is confused by the reverse frame shuffle, causing the camera pose of the last source frame to appear incorrectly in the first generated frame; ReCamMaster handles camera control but cannot modify temporal state.
- Longer sequences are feasible. A multi-turn autoregressive inference scheme chains 81-frame segments, each conditioned on a source video and the previously generated segment, enabling extended effects such as continuing a bullet-time rotation from 45° to 90°.
Methodology in Plain English
The model builds on a large text-to-video backbone (Wan-2.1 T2V-1.3B), which produces F′ = 21 latent frames that decode into F = 81 RGB frames through a 3D VAE, with a Transformer-based denoiser operating over multi-modal tokens.
The core idea is to give the network a second, separate control signal alongside the camera. The authors define an "animation time" vector t in R^F for both the source and target videos. These are turned into sinusoidal embeddings, then compressed from the fine-grained 81-frame space down to the 21-frame latent space by two 1D convolution layers, and finally added to the video token features together with the camera embedding. The authors report that reusing the existing frame-index RoPE embedding for temporal control is ineffective because it constrains camera and temporal motion at the same time.
The second change concerns camera conditioning. The prior approach assumes the first frame is identical between source and target and conditions only on the target trajectory. SpaceTimePilot instead estimates camera poses for both source and target videos with a pretrained pose estimator and injects both, concatenating target and source tokens along the frame dimension. This lets generation start from an arbitrary camera angle.
The third piece is data. Because no dataset offers paired videos of the same dynamic scene with continuous temporal variation, the authors warp the target video's frame order — reversal, acceleration, freezing, segmental slow motion, zigzag motion — while keeping the source as a standard forward reference. This creates explicit supervision for timing without new data collection. They then add Cam×Time, which renders every (camera, time) cell of a grid so that any two sampled sequences of F frames can form a source–target pair; one typical source choice is the grid diagonal. The default training mixture is ReCamMaster plus SynCamMaster with temporal warping, plus Cam×Time.
Only a subset of the network is trained: the camera embedder, the new animation-time embedder, the self-attention (full-3D attention), and the projector layers before cross-attention. Longer videos are produced by repeatedly generating a new segment conditioned on the previous one.
Why This Matters
Impact on research. The work shifts controllable video generation away from explicit 4D reconstruction pipelines (NeRFs, dynamic Gaussian splatting) toward direct generative rendering with an implicit 4D prior, and it argues that explicit paired supervision for temporal control can be manufactured by warping existing multi-view data. The Cam×Time dataset and its held-out test split are offered as a benchmark for fine-grained spatiotemporal modeling, and the paper positions space–time disentanglement as a distinct goal from camera control alone.
Real-world applications:
- Video post-production: re-angling and retiming a shot after it has been captured, without a reshoot.
- Film and sports effects: reverse playback, slow motion, and bullet-time sequences at arbitrary timesteps, including extended rotations built up over chained segments.
- Visual effects and virtual production: generating consistent novel camera paths around human and object motion in real footage.
- Immersive and interactive media: free exploration of a captured moment along both a camera axis and a time axis from a single monocular video.
Industry relevance. All but one author are affiliated with Adobe Research (the remaining affiliation is the University of Cambridge), and the target capabilities — camera re-posing, retiming, and motion effects — map directly onto editing and content-creation tooling rather than only academic benchmarking.
Future Directions
- Real-world temporal supervision. Temporal control is evaluated on the synthetic Cam×Time test split because ground-truth retimed frames for real dynamic footage are unavailable; camera accuracy is evaluated on real OpenVideoHD videos but with no temporal ground truth. Extending retiming evaluation to real captured dynamic scenes is an open problem.
- Scaling length and autonomy. The authors only begin to address longer video through a simple autoregressive, multi-turn, segment-chaining strategy with a 45°-to-90° bullet-time example; how far this scales to longer, coherent space–time trajectories is not resolved in the reported results.
- Stronger disentanglement guarantees. The paper shows joint training with static-scene datasets and augmentation without source-camera conditioning can degrade control, so the conditions under which temporal and spatial signals remain separated under new data mixtures or larger viewpoint changes remain an open question.
- Limitations. The main text provided does not include a dedicated limitations section, so specific failure modes, compute costs, and dataset licensing constraints are not reported.
Target Audience
Researchers and graduate students working on video diffusion models, novel view synthesis, 4D scene generation, and controllable generative rendering; engineers building video editing and VFX tooling; and practitioners who need camera and motion timing to be adjustable independently in captured footage. Readers without a diffusion-model background will still follow the problem framing and results, but the method section assumes familiarity with latent video diffusion architectures.
Authors’ abstract
We present SpaceTimePilot, a video diffusion model that disentangles space and time for controllable generative rendering. Given a monocular video, SpaceTimePilot can independently alter the camera viewpoint and the motion sequence within the generative process, re-rendering the scene for continuous and arbitrary exploration across space and time. To achieve this, we introduce an effective animation time-embedding mechanism in the diffusion process, allowing explicit control of the output video's motion sequence with respect to that of the source video. As no datasets provide paired videos of the same dynamic scene with continuous temporal variations, we propose a simple yet effective temporal-warping training scheme that repurposes existing multi-view datasets to mimic temporal differences. This strategy effectively supervises the model to learn temporal control and achieve robust space-time disentanglement. To further enhance the precision of dual control, we introduce two additional components: an improved camera-conditioning mechanism that allows altering the camera from the first frame, and CamxTime, the first synthetic space-and-time full-coverage rendering dataset that provides fully free space-time video trajectories within a scene. Joint training on the temporal-warping scheme and the CamxTime dataset yields more precise temporal control. We evaluate SpaceTimePilot on both real-world and synthetic data, demonstrating clear space-time disentanglement and strong results compared to prior work. Project page: https://zheninghuang.github.io/Space-Time-Pilot/ Code: https://github.com/ZheningHuang/spacetimepilot