Research
V-DPM: 4D Video Reconstruction with Dynamic Point Maps
V-DPM: 4D Video Reconstruction with Dynamic Point Maps Overview Research area: Computer vision — feed-forward 3D/4D reconstruction from images and video, specifically extending point-map representatio

- arXiv
- 2601.09499
- Published
- 2026-01-14
- Authors
- Edgar Sucar, Eldar Insafutdinov, Zihang Lai, Andrea Vedaldi
AI summary
V-DPM: 4D Video Reconstruction with Dynamic Point MapsOverview
Research area: Computer vision — feed-forward 3D/4D reconstruction from images and video, specifically extending point-map representations to dynamic scenes.
Technical level: Advanced. The paper assumes familiarity with camera extrinsics, point maps, scene flow, transformer backbones, and bundle-adjustment style optimisation.
Scope: The paper introduces V-DPM, a multi-view/video extension of Dynamic Point Maps (DPMs) built on top of the pretrained VGGT static reconstructor, and evaluates it on 4D reconstruction, dense 3D tracking, video depth, and camera pose benchmarks.
What This Paper Is About
Point maps such as DUSt3R's viewpoint-invariant representation encode 3D shape and camera parameters, but they assume the scene is static. Dynamic Point Maps (DPMs) remove that assumption by adding time invariance, yet prior DPM work handled only image pairs and required optimisation-based post-processing for more than two views.
V-DPM is the authors' answer: a design that extends DPMs to an entire video snippet in a single feed-forward pass, so that a network can recover 3D shape, 3D motion (scene flow), camera intrinsics, and camera motion together.
Key Contributions
-
A multi-image/video extension of Dynamic Point Maps. The authors formulate the problem in terms of two sets of point maps — time-varying point maps and time-invariant point maps at a chosen reference timestamp — and show that this reduces the redundant (N^3) (or, after fixing a common viewpoint, (N^2)) set of maps to (2N-1) maps per feed-forward pass.
-
A natural extension path for existing multi-view static reconstructors. Because the backbone predicts the same kind of time-varying point maps that models like VGGT already output for static scenes, a static network can be extended to dynamic reconstruction with minimal architectural change.
-
A time-conditioned decoder design. Added transformer decoder blocks with alternating frame and global attention, conditioned via adaptive LayerNorm (adaLN) on a target-time token, produce the time-invariant point maps and can be re-run for different target times while reusing backbone computations.
-
A demonstrated state-of-the-art 4D reconstructor from modest fine-tuning. VGGT, which was trained for static reconstruction only and never saw dynamic data before fine-tuning, is adapted into an effective V-DPM predictor using a mixture of static and synthetic dynamic data.
Main Findings
-
Large margin on 2-view 4D reconstruction. In Table 1 (End-Point Error on four predicted point maps (P_0(t_0)), (P_0(t_1)), (P_1(t_0)), (P_1(t_1))), V-DPM reports 0.029, 0.030, 0.032, 0.032 on PointOdyssey; 0.018, 0.019, 0.018, 0.018 on Kubric-F; 0.023, 0.024, 0.024, 0.023 on Kubric-G; and 0.064 on all four Waymo entries, at a margin of 2 frames. At a margin of 8 frames V-DPM reports 0.029, 0.031, 0.032, 0.030 (PointOdyssey), 0.017, 0.039, 0.033, 0.025 (Kubric-F), 0.022, 0.049, 0.045, 0.029 (Kubric-G), and 0.065, 0.067, 0.065, 0.064 (Waymo). The paper states that St4RTrack and TraceAnything trade places on PointOdyssey and Kubric, while V-DPM achieves roughly 5 times lower error than both.
-
Advantage grows on long sequences. In the 10-frame dense tracking experiment (Table 2, Tracking EPE), V-DPM reports 0.032 (PointOdyssey), 0.027 (Kubric-F), 0.035 (Kubric-G), and 0.042 (Waymo), compared with DPM at 0.114, 0.088, 0.109, 0.103; TraceAnything at 0.152, 0.107, 0.126, 0.119; and St4RTrack at 0.137, 0.153, 0.201, 0.167. The paper explains this as DPM being unable to use temporal context from pairs alone.
-
Video depth: competitive but not the best. On Sintel and Bonn (Table 3), V-DPM records Abs Rel 0.247 and (\delta<1.25) 69.4 on Sintel, and 0.057 and 97.3 on Bonn. The concurrent (\pi^3) is stronger on both (0.210 / 72.6 and 0.043 / 97.5), which the authors attribute to scale: (\pi^3) could train on 14 public datasets plus an internal dynamic dataset, whereas V-DPM uses 6.
-
Camera pose: competitive, again behind (\pi^3). On Sintel (Table 4) V-DPM records ATE 0.105, RPE trans 0.048, RPE rot 0.67; on TUM-dynamics it records 0.057, 0.017, 0.34. (\pi^3) records 0.074, 0.040, 0.282 and 0.014, 0.009, 0.312 respectively. DPM is listed as having no Sintel values.
-
Qualitative robustness. Figure 7 compares 4D reconstructions from 10-frame snippets on DAVIS. The authors report that both DPM and St4RTrack fail on the fishtank sequence, and only V-DPM plausibly reconstructs a tennis player's body pose at the final timestep, with smoother and more self-consistent trajectories.
-
Design choices matter. An ablation trained for 35 epochs on Kubric-G with two views at margin 8 reports (P_0(t_1)) / (P_1(t_0)) errors of 0.0500 / 0.0472 for the full model, 0.0518 / 0.0476 with decoder depth reduced to two blocks, 0.0524 / 0.0484 when conditioning by adding the time token instead of adaLN, and 0.0538 / 0.0502 when using a conditioned copy of the DPT head instead of extra transformer layers.
-
Training efficiency. V-DPM is fine-tuned on 16 GH200 GPUs for 60 epochs with AdamW at a base learning rate of (1.5 \times 10^{-4}) and cosine decay. Fine-tuning covered snippets up to 20 frames, though the authors found it generalises to roughly 50 frames at test time; longer sequences of hundreds of frames are handled with a sliding window and bundle-adjustment fusion.
Methodology in Plain English
The authors build on the Dynamic Point Map idea, which says that for a pair of images, four point clouds are enough to describe everything about the scene: the 3D shape at each of the two times, expressed from a common viewpoint. One looks up where a pixel's 3D point was at time A versus time B to get its motion.
The problem is scale. With (N) images and free choice of viewpoint and time, the number of possible point maps grows quickly. The authors cut this down in two steps. First, since all relative-viewpoint maps can be derived by a rigid transform once cameras are known, they fix one common viewpoint. That leaves (N^2) maps, still too many. Second, they split the task: the network backbone predicts one time-varying point map per input image (same viewpoint, each at its own timestamp), and then dedicated decoder layers predict a second set of maps that place every image's points at one shared reference timestamp. That second step is what turns the time-varying reconstruction into a time-invariant one and implicitly produces correspondences across time — which is what scene flow needs.
Because the first-stage outputs look exactly like what a static multi-view reconstructor such as VGGT already produces for static scenes, the static model can be fine-tuned rather than retrained. The authors add a time-conditioned transformer decoder that repeatedly refines the backbone features to align all frames to a chosen reference frame's features. The target timestamp is supplied as an extra token, whose processed representation modulates the decoder blocks through adaptive LayerNorm and gates the attention outputs.
Practically, the backbone runs once per video snippet; changing the reference timestamp only requires re-running the decoder, which keeps the cost of reconstructing the scene at many times low. Training mixes static datasets (ScanNet++, BlendedMVS) with dynamic ones (Kubric-F, Kubric-G, PointOdyssey, Waymo), sampling windows of 5, 9, 13, or 19 frames and adjusting batch size to fit memory (batch 4 at 5 frames, batch 1 at 19 frames). A per-example-then-per-batch loss normalisation prevents dense static annotations from swamping the sparse dynamic 3D tracks.
Why This Matters
The work suggests that the expensive part of learning 4D reconstruction is not the dynamic data itself but the representation and the reuse of a strong static prior. Since static 3D datasets are easy to obtain and auto-annotate, a training recipe that combines them with a small amount of synthetic 4D data is a practical path forward.
Real-world applications:
- Visual effects and post-production, where scene flow and camera parameters are needed to insert or manipulate objects in moving footage.
- Robotics and manipulation, where a robot must reason about how objects move and deform (the paper shows a robot manipulation example).
- Video generation and world modelling, where a model of how points move provides structure for synthesis.
- Vision-based control, where dense 3D tracks of every pixel in a snippet inform action.
Industry relevance: Any pipeline currently relying on DUSt3R- or VGGT-style static reconstruction — AR/VR capture, autonomous driving (Waymo appears in the training data), film and broadcast tooling — can in principle substitute a dynamic-capable model, and the "fine-tune a static backbone" recipe lowers the barrier to doing so.
Future Directions
-
Scaling data and backbone. The authors explicitly state that V-DPM is only outperformed on static 3D and camera tasks by (\pi^3), and expect that scaling up training data and adopting a stronger, more recent backbone would close the gap; they also note V-DPM could be integrated on top of a network like (\pi^3).
-
Longer sequences. Training only covered snippets up to 20 frames; the model generalises to about 50 at test time, and beyond that requires sliding-window optimisation. Extending native temporal context is an open direction.
-
Reducing reliance on synthetic 4D data. The paper frames the combination of large static datasets with small synthetic 4D sets as a template, raising the question of how far that mix can be pushed before dynamic data becomes the bottleneck.
-
Evaluation scale. The authors name the scale of their evaluation as a limitation of the work, constrained by available resources.
Target Audience
Researchers and engineers working on feed-forward 3D and 4D reconstruction, multi-view geometry, dynamic scene representation, and dense point tracking. It is most useful to readers already familiar with DUSt3R, MASt3R, VGGT, or MonST3R, and to practitioners who want to add dynamic capability to an existing static reconstruction pipeline without training a new model from scratch.
Authors’ abstract
Powerful 3D representations such as DUSt3R invariant point maps, which encode 3D shape and camera parameters, have significantly advanced feed forward 3D reconstruction. While point maps assume static scenes, Dynamic Point Maps (DPMs) extend this concept to dynamic 3D content by additionally representing scene motion. However, existing DPMs are limited to image pairs and, like DUSt3R, require post processing via optimization when more than two views are involved. We argue that DPMs are more useful when applied to videos and introduce V-DPM to demonstrate this. First, we show how to formulate DPMs for video input in a way that maximizes representational power, facilitates neural prediction, and enables reuse of pretrained models. Second, we implement these ideas on top of VGGT, a recent and powerful 3D reconstructor. Although VGGT was trained on static scenes, we show that a modest amount of synthetic data is sufficient to adapt it into an effective V-DPM predictor. Our approach achieves state of the art performance in 3D and 4D reconstruction for dynamic scenes. In particular, unlike recent dynamic extensions of VGGT such as P3, DPMs recover not only dynamic depth but also the full 3D motion of every point in the scene.