Research
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images Overview Research area: Computer vision — dense 4D scene reconstruction, dynamic 3D Gaussian Splatting, feedforward multi-task perception

- arXiv
- 2602.24290
- Published
- 2026-02-27
- Authors
- Junhwa Hur, Charles Herrmann, Songyou Peng, Philipp Henzler, Zeyu Ma, Todd Zickler, Deqing Sun
AI summary
UFO-4D: Unposed Feedforward 4D Reconstruction from Two ImagesOverview
Research area: Computer vision — dense 4D scene reconstruction, dynamic 3D Gaussian Splatting, feedforward multi-task perception (geometry, motion, and camera pose estimation).
Technical level: Advanced. The paper assumes familiarity with 3D Gaussian Splatting, differentiable rasterization, vision transformers, photometric self-supervision, and standard reconstruction metrics (EPE, Abs. Rel., δ<1.25, ATE, RPE, PSNR).
Scope: The paper presents a single feedforward network that takes two unposed images plus camera intrinsics and directly outputs dynamic 3D Gaussians and relative camera pose, evaluated on geometry, motion, and pose benchmarks plus a new 4D interpolation application.
What This Paper Is About
Recovering dense 4D structure — camera pose, 3D geometry, and 3D motion — from casually captured images is hard because the problem is ill-posed and large-scale densely annotated 4D training data is scarce (synthetic data has a domain gap; the real-world Stereo4D dataset provides sparse and sometimes unreliable annotations). Prior work either relies on slow per-scene test-time optimization or on fragmented feedforward models that solve one task at a time, which prevents them from exploiting the tight coupling between geometry and motion.
UFO-4D's goal is to reconstruct a dense, explicit 4D representation from just a pair of unposed images in a single forward pass, and to use differentiable rendering of that representation as a self-supervised signal that compensates for sparse ground truth.
Key Contributions
- A unified model for unposed feedforward 4D reconstruction from two images, built on a dynamic 3D Gaussian Splatting representation that jointly estimates 3D geometry, 3D motion, and relative camera pose.
- A robust semi-supervision framework that leverages differentiably rendered image, point, and motion outputs to mitigate the scarcity of annotated data.
- 4D spatio-temporal interpolation of image, depth, and motion as a new application of the feedforward output.
- State-of-the-art performance on 3D geometry and 3D motion benchmark datasets.
Main Findings
-
Geometry estimation: On the Stereo4D test split, UFO-4D reports EPE 0.6593, Abs. Rel. 0.1063, and δ<1.25 of 88.38, beating DynaDUSt3R (EPE 0.811, Abs. Rel. 0.112, δ<1.25 86.07), ZeroMSF (1.6619), St4RTrack (1.4668), MonST3R (1.5038), and MASt3R (1.8869). On Bonn it reaches EPE 0.1622, Abs. Rel. 0.064, δ<1.25 96.04; on KITTI, EPE 2.568, Abs. Rel. 0.090, δ<1.25 88.92; on Sintel final, EPE 3.8660, Abs. Rel. 0.319, δ<1.25 61.94. The paper notes the method substantially outperforms competitors on Stereo4D and remains very competitive on the other datasets, attributing per-dataset trade-offs to differences in each method's training data mixture.
-
Scene flow estimation: On Stereo4D forward flow, UFO-4D reports 3D EPE 0.0488 and δ3D^0.05 of 83.0618; backward flow, EPE 0.0500 and δ3D^0.05 of 82.6700. On KITTI forward flow, EPE 0.1372 and δ3D^0.05 of 47.7964. The paper states this is more than three times lower EPE on Stereo4D and KITTI than the best competing method, and attributes the gain to clearer motion boundaries and better separation of moving objects from background.
-
Camera pose estimation: Reported for Stereo4D, Bonn, and Sintel. On Stereo4D, ATE 0.0101, RPE_trans 0.0142, RPE_rot 0.1794; on Bonn, ATE 0.0020, RPE_trans 0.0028, RPE_rot 0.2709; on Sintel, ATE 0.0122, RPE_trans 0.0172, RPE_rot 0.2980. The paper states the fully feedforward approach outperforms methods that rely on iterative solvers such as MonST3R and St4RTrack, and that direct pose estimation removes the need for inference-time post-regression.
-
Loss ablation (Stereo4D, 256×256, 40k iterations, batch size 12): The full model achieves PSNR 23.929, Point EPE 0.827, rasterized Point EPE 0.841, Motion EPE 0.069, rasterized Motion EPE 0.064. Removing the image gradient raises PSNR to 24.268 but degrades Point to 0.903, rasterized Point to 0.899, Motion to 0.072, and rasterized Motion to 0.070. Removing motion and point rendering drops PSNR to 20.675 and degrades Point to 0.911, rasterized Point to 0.943, Motion to 0.071, rasterized Motion to 0.069. Removing motion, point, and image rendering yields Point 0.886 and Motion 0.071, with no PSNR or rasterized values reported. The paper states the rendering losses are crucial for sharp motion boundaries.
-
Architecture comparison (trained under the authors' protocol, 256×256, 60k iterations): The Dynamic 3DGS representation gives Stereo4D Motion EPE 0.058 and Point EPE 0.781, KITTI Motion 0.244 and Point 3.553, and Bonn Point 0.211. A per-pixel point-and-motion representation (equivalent to DynaDUSt3R and ZeroMSF) gives Stereo4D 0.057 and 0.809, KITTI 0.283 and 4.567, Bonn 0.196. A per-pixel point-only representation (equivalent to MonST3R) gives Stereo4D Point 0.804, KITTI Point 4.562, Bonn Point 0.201. The paper reports gains on Stereo4D and substantial gains on KITTI, but slightly behind performance on Bonn, which it attributes to textureless regions where large-scale overlapping Gaussians introduce error-prone spatial dependencies.
-
Opacity as learnable confidence: An emergent behavior, visualized on a heavy occlusion/disocclusion example, is that opacity acts as a learned confidence value. The model assigns high opacity to regions that are disoccluded between views, and for mutually visible regions selects only one of the two corresponding Gaussians, producing a compact 4D representation.
-
4D interpolation: Given an input pair, the model renders images, depth, and motion maps at interpolated times between the two frames, at the canonical camera coordinate and at arbitrary camera views. A qualitative example is shown on DAVIS.
-
Training setup: Training samples Stereo4D, PointOdyssey, and Virtual KITTI 2 with probabilities 60%, 20%, and 20%; 10% of the full Stereo4D dataset is used by subsampling frames. The network is initialized from NoPoSplat (Gaussian head) and MASt3R (rest) pretrained weights, with the camera head trained from scratch. Image width is 512, training runs 120k iterations with a global minibatch size of 16, taking around three days on four NVIDIA A100 40GB GPUs.
Methodology in Plain English
The system takes two images and their camera intrinsics and predicts two things: the relative camera pose between them, and one dynamic 3D Gaussian for every pixel in both images. Each Gaussian carries a 3D center, a 3D motion vector, rotation, size, spherical-harmonic color, and opacity. Gaussians from the first image get forward motion and Gaussians from the second get backward motion, so translating them along their velocity vectors represents the scene at any time between the two frames. Everything is expressed in the first camera's coordinate system, which serves as the canonical space.
Architecturally, the model follows a DUSt3R/NoPoSplat-style design: a shared-weight encoder turns each image into tokens, those tokens are concatenated with an intrinsic token (from camera intrinsics through a linear layer) and a learnable pose token, and a ViT-style decoder with cross-attention matches information between the two images. Separate DPT-based heads predict Gaussian centers, Gaussian attributes, and velocities, while a three-layer MLP predicts the relative pose as translation plus quaternion.
The core trick is the renderer. The standard 3D Gaussian alpha-blending formula is extended so that instead of blending colors only, the same depth-sorted weights blend the Gaussians' 3D centers (producing dense pointmaps) and their motion vectors (producing dense 3D scene flow). This makes images, geometry, and motion all differentiable outputs of the same primitives. Training combines a supervised loss on motion, points, and pose with a self-supervised loss made of a photometric term (MSE plus LPIPS between input and rendered images) and an edge-aware smoothness term on rendered points and motion. Because the geometry and motion losses act on the same Gaussians, supervision on one regularizes the other, which the authors argue compensates for sparse ground truth. Downstream 2D tasks fall out of the 3D outputs by projection: depth is the last channel of the point, optical flow comes from projecting 3D scene flow, and moving objects are segmented by thresholding scene flow.
Why This Matters
The work argues that a single explicit, differentiable 4D representation lets view synthesis and reconstruction supervise each other, avoiding both the cost of test-time optimization and the fragmentation of task-specific heads. It also shows that camera pose can be regressed directly in a feedforward manner with better accuracy than PnP-plus-RANSAC pipelines on the reported benchmarks.
Real-world applications identified or implied by the paper:
- Robotics, where pose, geometry, and motion need to be recovered from casually captured images.
- Autonomous driving, using benchmarks such as KITTI and Virtual KITTI 2.
- 3D/4D generative AI and content creation, via novel-view and novel-time interpolation of appearance, geometry, and motion.
- Downstream perception products built on the outputs: depth estimation, optical flow, scene flow, and moving-object segmentation from the same prediction.
Industry relevance: the paper positions feedforward 4D reconstruction as an alternative to pipelines that need hours of per-scene optimization and pre-computed camera poses or optical flow, which matters for any real-time or on-device setting. The paper does not report latency, on-device, or deployment measurements.
Future Directions
- Scaling from two frames to long sequences: the authors note a naive extension is memory-intensive because the number of Gaussian Splats grows linearly, and suggest compact scene representations as a direction.
- Handling more complex dynamics: the current linear-motion and constant-brightness assumptions suit short time intervals; learnable non-linear motion models and time-varying Gaussian attributes are proposed for more complex dynamics and photometric changes.
- Adding rendered images, motion, and geometry at novel views and times as extra supervision objectives when multi-view video annotations are available, to improve spatio-temporal consistency.
- Understanding the Bonn regression: the authors leave open why the Gaussian representation trails a per-pixel point representation in textureless regions, and discuss the error-prone spatial dependencies of large overlapping Gaussians.
Target Audience
Researchers and engineers working on 4D reconstruction, dynamic Gaussian Splatting, feedforward multi-view/pose estimation, and differentiable rendering; practitioners in robotics, autonomous driving, and generative 3D content who need pose, geometry, and motion from image pairs; and readers already familiar with DUSt3R/MASt3R-style architectures, MonST3R, DynaDUSt3R, ZeroMSF, and St4RTrack who want to see how a shared dynamic Gaussian representation changes the accuracy trade-offs across geometry, motion, and pose.
Authors’ abstract
Dense 4D reconstruction from unposed images remains a critical challenge, with current methods relying on slow test-time optimization or fragmented, task-specific feedforward models. We introduce UFO-4D, a unified feedforward framework to reconstruct a dense, explicit 4D representation from just a pair of unposed images. UFO-4D directly estimates dynamic 3D Gaussian Splats, enabling the joint and consistent estimation of 3D geometry, 3D motion, and camera pose in a feedforward manner. Our core insight is that differentiably rendering multiple signals from a single Dynamic 3D Gaussian representation offers major training advantages. This approach enables a self-supervised image synthesis loss while tightly coupling appearance, depth, and motion. Since all modalities share the same geometric primitives, supervising one inherently regularizes and improves the others. This synergy overcomes data scarcity, allowing UFO-4D to outperform prior work by up to 3 times in joint geometry, motion, and camera pose estimation. Our representation also enables high-fidelity 4D interpolation across novel views and time. Please visit our project page for visual results: https://ufo-4d.github.io/