Research
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
Light-X: Generative 4D Video Rendering with Camera and Illumination Control Overview Research area: Computer vision / generative video modeling — specifically controllable video generation combining n
- arXiv
- 2512.05115
- Published
- 2025-12-04
- Authors
- Tianqi Liu, Zhaoxi Chen, Zihao Huang, Shaocong Xu, Saining Zhang, Chongjie Ye, Bohan Li, Zhiguo Cao, Wei Li, Hao Zhao, Ziwei Liu
AI summary
Light-X: Generative 4D Video Rendering with Camera and Illumination ControlOverview
Research area: Computer vision / generative video modeling — specifically controllable video generation combining novel-view synthesis (camera control) with video relighting (illumination control).
Technical level: Advanced. The paper builds on diffusion transformers (DiT), latent video diffusion, point-cloud rendering, VAE encoders, Q-Former modules, and large-scale data curation pipelines.
One-sentence scope: The paper introduces Light-X, a video generation framework that jointly controls camera trajectory and illumination when re-rendering a monocular video, together with Light-Syn, a degradation-based pipeline that synthesizes the paired multi-view, multi-illumination training data such control requires.
What This Paper Is About
A monocular video records only a two-dimensional projection of a scene that is actually shaped by geometry, motion, and lighting together. Prior work has advanced along two separate tracks: video relighting methods (which change lighting but cannot move the camera) and camera-controlled video generation methods (which move the camera but cannot change lighting). This paper asks whether a single model can do both at once, and how to train such a model when paired multi-view, multi-illumination videos do not exist in the real world.
Key Contributions
- Light-X, described as the first framework for video generation with joint control of camera trajectory and illumination from monocular videos, supporting joint camera–illumination control, video relighting, and novel-view synthesis within a single model.
- Light-Syn, a degradation-based data pipeline with inverse geometric mapping that constructs paired training data (inputs, targets, and geometrically aligned conditioning cues) under controlled camera viewpoints and lighting, without requiring multi-view or multi-illumination captures.
- A disentangled conditioning scheme that explicitly separates geometry and motion from illumination cues, using dynamic point clouds projected along user-defined camera trajectories for geometry and a re-projected relit frame for lighting, plus an introduced Light-DiT layer for global illumination consistency.
- Extensive experiments reporting state-of-the-art performance in joint camera–illumination control and in video relighting under both text- and background-conditioned settings, supported by user studies with 57 participants.
Main Findings
- Joint camera–illumination control (Table 1): Light-X achieves FID 101.06, Aesthetic 0.623, Motion Preservation 2.007, and CLIP 0.989, versus TL-Free (FID 122.73, Aesthetic 0.595, Motion Pres. 3.356, CLIP 0.987), TC+LAV (FID 138.89, 0.574, 4.327, 0.986), LAV+TC (FID 144.61, 0.596, 5.027, 0.987), and TC+IC-Light (0.573, 6.558, 0.976). Light-X runs in 1.83 min, compared with 3.25 min (TC+IC-Light), 4.33 min (TC+LAV and LAV+TC), and 5.50 min (TL-Free).
- Joint control against real in-the-wild videos (Table 2): Light-X reaches PSNR 13.96, SSIM 0.582, LPIPS 0.378, FVD 45.91, better than TL-Free (13.49 / 0.547 / 0.418 / 54.44), LAV+TC (12.48 / 0.463 / 0.479 / 60.95), TC+LAV (12.18 / 0.470 / 0.508 / 73.78), and TC+IC-Light (10.96 / 0.456 / 0.474 / 58.85).
- Text-conditioned video relighting (Table 3): Light-X reports FID 83.65, Aesthetic 0.645, Motion Preservation 1.137, and CLIP 0.993 in 1.50 min, versus LAV (112.45 / 0.614 / 2.115 / 0.991, 2.50 min), IC-Light+AnyV2V (106.05 / 0.612 / 3.777 / 0.985, 1.67 min), and frame-wise IC-Light (0.632 / 3.293 / 0.983, 1.42 min). Evaluated on the first 16 frames, Light-X scores FID 77.97, Aesthetic 0.625, Motion Pres. 1.452, CLIP 0.992.
- Text-conditioned relighting against real videos (Table 4): Light-X achieves PSNR 13.84, SSIM 0.581, LPIPS 0.369, FVD 56.60, versus LAV (12.66 / 0.530 / 0.429 / 74.64) and IC-Light (11.75 / 0.517 / 0.422 / 67.50).
- Background-conditioned foreground video relighting (Table 5): Light-X scores FID 61.75, Aesthetic 0.680, Motion Preservation 0.220, CLIP 0.992, ahead of LAV (76.05 / 0.619 / 0.296 / 0.990), RelightVid (86.94 / 0.635 / 0.230 / 0.988), and IC-Light (0.645 / 0.374 / 0.987). On the first 16 frames, Light-X reaches FID 56.60, Aesthetic 0.682, Motion Pres. 0.199, CLIP 0.990.
- User study preference: With 57 participants, roughly 85–98 percent of participants preferred Light-X over each baseline across relighting quality (RQ), video smoothness (VS), ID preservation (IP), and 4D consistency (4DC) when compared against the individual baselines. Frame-wise IC-Light had a notable advantage in user preference on some video-relighting criteria (RQ 88.3, VS 90.3, IP 91.7 in Table 3).
- Ablation on data mixture (Table 6): Removing static data gives FID 123.35 / Aesthetic 0.594 / Motion Pres. 3.749 / CLIP 0.987; removing dynamic data gives 108.70 / 0.621 / 2.635 / 0.988; removing AI-generated data gives 102.09 / 0.613 / 2.498 / 0.988, versus Light-X at 101.06 / 0.623 / 2.007 / 0.989.
- Ablation on architecture (Table 6): Removing fine-grained lighting cues yields FID 143.02 / 0.602 / 2.242 / 0.989; removing global lighting control yields 103.13 / 0.612 / 2.348 / 0.989; replacing the conditioning with light–text concatenation yields 137.05 / 0.596 / 2.654 / 0.989.
- Ablation on training and conditioning strategy (Table 6): Using algorithm-generated ground truth instead of real video yields 137.83 / 0.524 / 4.066 / 0.986; relighting all frames instead of a single frame yields FID 71.10 / 0.571 / 4.238 / 0.986 (better FID but worse temporal coherence and higher cost); dropping the soft mask yields 148.51 / 0.545 / 2.879 / 0.988.
- Behavioral observations: Illumination strength was observed to gradually diminish as synthesized frames move further from the relit frame, motivating the global illumination control module. Relighting all frames "increases cost and reduces temporal coherence despite better FID."
Methodology in Plain English
Decoupling camera from lighting. Given a source video, the method builds two dynamic point clouds. The first comes from estimating per-frame depth and back-projecting each frame into 3D; projecting this cloud along a user-specified camera trajectory produces geometry-aligned views plus visibility masks. The second comes from relighting one frame with IC-Light according to a lighting prompt, keeping the rest blank, and lifting that sparse relit video into a point cloud using the same depths (not re-estimating depth from the relit image, so the two stay geometrically aligned). Projecting the relit cloud along the same trajectory yields relit views and masks that say where lighting information exists.
Conditioning and architecture. Six cues—source video, sparse relit video, projected source views and masks, projected relit views and masks—are encoded by a VAE and merged with sampled noise and text tokens, then passed through DiT blocks. A separate Q-Former extracts a global illumination token from the relit frame, injected via cross-attention in the introduced Light-DiT layer; the original DiT and Ref-DiT modules aggregate text-vision information and preserve 4D consistency with the source video.
Synthesizing training data. Because paired multi-view, multi-illumination video does not exist, Light-Syn takes an in-the-wild video as the target, degrades it to make the input, records the degradation, and applies the inverse transformations to carry the target's geometry and lighting onto the degraded video. Training data come from static scenes (8k), dynamic scenes (8k), and AI-generated videos (2k). Extensions add HDR environments (16k samples, using DiffusionLight for environment lighting) and reference-image conditioning (about 1k samples each for text- and background-conditioned settings). Soft masks act as domain indicators, with alpha_ref = 0.25 and alpha_hdr = 0.50.
Flexibility. Because camera and lighting conditioning are decoupled and masked, the same model supports camera control alone (substituting the original frame for the relit frame), illumination control alone (setting projected views to the source video and using fully visible masks), and background-image-conditioned foreground relighting. Other illumination hints such as environment maps and reference images can also be supplied as conditioning inputs.
Evaluation setup. The model is trained at 384 × 672 resolution on 49 frames for 16,000 iterations with learning rate 2 × 10^-5 and batch size 8 on eight H100 GPUs. Evaluation uses 200 collected videos from Pexels, Sora, and Kling; background-conditioned relighting uses 10 background images and 30 foreground videos producing 300 combinations. None of these videos was used for training in any compared method.
Why This Matters
Impact on research. The paper reframes controllable video generation as a disentanglement problem rather than two separate tasks, and pairs that framing with a data-synthesis strategy that sidesteps the impossible requirement of capturing multi-view, multi-illumination video. It also reports improved runtime over combined baselines (1.83 min versus 3.25–5.50 min), suggesting joint control need not come at a computational premium.
Real-world applications:
- Immersive AR/VR, where a user revisits footage from a new viewpoint under new lighting.
- Filmmaking pipelines that need to re-shoot or re-light a scene without re-shooting it.
- Relighting existing monocular footage under text prompts, background images, or HDR environment maps.
- Reference-image-driven lighting transfer onto video content, in the manner of a style-transfer source.
Industry relevance. Any pipeline that consumes user-captured monocular video — content creation tools, visual effects, e-commerce and virtual production — stands to benefit from a single model that handles both viewpoint and lighting changes, especially given the reported runtime advantage and the ability to accept multiple illumination modalities (text, background, HDR map, reference image) within one model.
Future Directions
- Stronger lighting priors. The method relies on single-image relighting priors such as IC-Light, and the authors note that suboptimal prior quality in some scenes propagates into video generation quality.
- Better geometry and wider camera motion. Because point clouds are the novel-view prior, inaccurate depth estimation biases geometry; the framework also struggles with very wide camera motions (e.g., 360 degrees) due to limited 3D cues and constrained generation length.
- Longer and cheaper generation. The authors propose stronger backbones such as Wan2.2, progressive point-cloud expansion for larger camera ranges, and techniques such as Diffusion Forcing and Self Forcing to extend video length; multi-step denoising remains computationally expensive.
- Fine detail fidelity. Handling fine details such as hands remains challenging, matching a limitation common to other video diffusion approaches.
Target Audience
Researchers and graduate students working on video diffusion models, controllable video generation, and computational relighting; engineers building content-creation, VFX, or AR/VR pipelines who need joint viewpoint and lighting manipulation; and practitioners interested in data curation strategies that manufacture training pairs from unpaired in-the-wild video. The paper is written for readers already comfortable with diffusion transformers, latent video diffusion, and novel-view synthesis terminology.
Authors’ abstract
Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighting, a key step toward generative modeling of real-world scenes is the joint control of camera trajectory and illumination, since visual dynamics are inherently shaped by both geometry and lighting. To this end, we present Light-X, a video generation framework that enables controllable rendering from monocular videos with both viewpoint and illumination control. 1) We propose a disentangled design that decouples geometry and lighting signals: geometry and motion are captured via dynamic point clouds projected along user-defined camera trajectories, while illumination cues are provided by a relit frame consistently projected into the same geometry. These explicit, fine-grained cues enable effective disentanglement and guide high-quality illumination. 2) To address the lack of paired multi-view and multi-illumination videos, we introduce Light-Syn, a degradation-based pipeline with inverse-mapping that synthesizes training pairs from in-the-wild monocular footage. This strategy yields a dataset covering static, dynamic, and AI-generated scenes, ensuring robust training. Extensive experiments show that Light-X outperforms baseline methods in joint camera-illumination control and surpasses prior video relighting methods under both text- and background-conditioned settings.