Research
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis Overview Research area: Computer vision — sparse-view novel view synthesis (NVS), combining 3D visual geometry foundatio

- arXiv
- 2609.33253
- Published
- 2026-09-27
- Authors
- Kangjie Chen, Xiangyu Li, Dongbin Zhang, Chaoda Zheng, Shijia Chen, Jinhao Deng, Hongbin Lin, Choo Sin Wai, Minqi Wang, Minghao Yang, Dake Zhong, Guorui Song, Yu Zhang, Xianming Liu, Boyang Wang
AI summary
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View SynthesisOverview
Research area: Computer vision — sparse-view novel view synthesis (NVS), combining 3D visual geometry foundation models with pretrained video diffusion.
Technical level: Advanced. The paper assumes familiarity with diffusion/flow-matching models, transformer backbones (DiT), camera projection geometry, and point-cloud reconstruction metrics.
Scope: The paper introduces VGGT-Diff, a geometry-routed multi-view diffusion framework that injects visual geometry latents from VGGT-Ω into a pretrained video diffusion model (Wan2.1-I2V-14B) to synthesize novel views from only six source images.
What This Paper Is About
Novel view synthesis from sparse inputs sits between two imperfect options: reconstruction-based methods preserve the geometry they can see but produce artifacts in unobserved regions, while diffusion-based methods generate plausible content but often ignore or override the scene's actual 3D structure, causing structural drift and cross-view inconsistency. This paper's goal is to get both — strong completion of unseen regions and fidelity to observed geometry — by routing uncertainty-aware geometry features from a visual geometry foundation model into a pretrained video diffusion model, rather than using geometry merely as pre-rendered RGB or as a full scene reconstruction.
Key Contributions
-
Geometry-routed generation. A multi-view diffusion framework whose confidence-aware Visual Geometry Router (VGR) transforms VGGT-Ω features, 3D points, and per-token confidence into query-aligned latent conditions, grounding generative completion in observed scene structure. The router builds a hard depth-selected anchor plus confidence-weighted front and secondary "layered" hypotheses to retain evidence near occlusion boundaries.
-
Geometry-grounded consistency. Point-Track Residual Consistency (PTRC) aligns the predicted-clean residuals of the same physical 3D point across jointly generated views, using VGGT-derived correspondences, so cross-view drift is reduced without forcing view-dependent appearance to match.
-
Geometry-condition regularization. Routed geometry is stochastically retained, attenuated, or removed during training to prevent over-reliance on imperfect projections. The dropped-condition branch also supports an optional matched Geometry-Prior CFG at inference that strengthens geometry-aware denoising.
-
A joint multi-view diffusion design. Source and query views are placed in a single DiT sequence with temporal ordering removed from the view axis, and query slots are denoised jointly; flow matching is supervised on query slots only (query RGB is never exposed as a condition).
Main Findings
- Top-ranked on DL3DV-Benchmark. Across 6,188 DL3DV targets, VGGT-Diff ranks first in PSNR, LPIPS, and DreamSim and second in SSIM, using only ~1k training scenes compared with far larger budgets for competing methods (e.g., LVSM 67.5K, SEVA 80K, DepthSplat 77.5K, AnySplat 254K, FrameCrafter 1K).
- Absolute DL3DV numbers. VGGT-Diff: PSNR 18.104, SSIM 0.508, LPIPS 0.222, DreamSim 0.062. The strongest diffusion baseline, FrameCrafter (same Wan2.1-I2V-14B prior), reaches PSNR 17.180, SSIM 0.445, LPIPS 0.223, DreamSim 0.066 — a gain of 0.924 dB PSNR and 0.063 SSIM.
- Zero-shot transfer to Mip-NeRF 360. VGGT-Diff ranks first in LPIPS and DreamSim and second in PSNR: PSNR 16.296, SSIM 0.317, LPIPS 0.315, DreamSim 0.091 (vs. FrameCrafter 15.640 / 0.279 / 0.365 / 0.111). E-RayZer and DepthSplat lead on PSNR and SSIM respectively, but the paper argues their higher perceptual errors indicate those gains come at a cost in perceptual quality.
- Robust across pose difficulty. On 192 pose-stratified DL3DV targets (32 per bin), VGGT-Diff PSNR is 20.920 / 18.391 / 15.853 for interpolation near/mid/far and 22.396 / 18.793 / 17.793 for extrapolation near/mid/far. It leads five bins and trails SEVA by only 0.020 dB in interpolation-near; its margin over FrameCrafter ranges from 0.582 to 1.228 dB and exceeds 0.9 dB in five bins.
- Improved cross-view geometric consistency. Across 52 DL3DV scenes, eight jointly generated views were reconstructed with two frozen reconstructors, VGGT-Ω (the model used for conditioning) and Pi3 (never used in training or conditioning). VGGT-Diff ranks first under both, reducing Chamfer-L1 by 26.7% with VGGT-Ω and 16.4% with Pi3 relative to FrameCrafter, and improving F-score@1% / F-score@2% by 11.45 / 12.48 and 9.49 / 11.41 percentage points respectively. Reported values: VGGT-Ω — Chamfer 0.0733, F@1% 0.4231, F@2% 0.6041; Pi3 — 0.1174, 0.3089, 0.5021.
- Visual geometry is the vital ingredient. Removing all VGGT-Ω-dependent components causes the largest ablation degradation: PSNR falls 1.730 dB, SSIM by 0.056, and LPIPS rises by 0.056 (17.294 / 0.501 / 0.357 versus the 19.024 / 0.557 / 0.301 full checkpoint without CFG). The authors conclude camera rays plus the Wan prior alone cannot replace visual geometry.
- Features beat rendered RGB. Replacing routed appearance-bearing features with point-rendered RGB, while keeping the geometry, router, and PTRC, costs 0.471 dB PSNR and raises LPIPS by 0.015 (18.553 / 0.536 / 0.316).
- PTRC matters. Removing PTRC costs 0.438 dB PSNR and 0.027 SSIM (18.586 / 0.530 / 0.312) — the largest SSIM loss among single-component ablations.
- Layered refinement is a smaller but real gain. Removing it while keeping the hard anchor costs 0.129 dB PSNR and 0.009 SSIM (18.895 / 0.548 / 0.303), indicating hard routing captures most of the benefit and layered refinement corrects ambiguous locations.
- Condition regularization is the primary robustness mechanism. Removing it costs 0.302 dB PSNR and 0.019 SSIM while raising LPIPS by 0.008 (18.722 / 0.538 / 0.309). Applying Geometry-Prior CFG to the same checkpoint improves PSNR by 0.115 dB, SSIM by 0.006, and reduces LPIPS by 0.005 (19.139 / 0.563 / 0.296) without retraining.
Methodology in Plain English
- Inputs and output. Given six posed source images and a set of requested query cameras, the system generates images for those query cameras. Source and query views are packed into a single diffusion sequence so information can flow between target views, but the temporal ordering of the view axis is dropped since the inputs are a set of camera observations, not a video timeline.
- Base model and conditioning. The system builds on a pretrained Wan2.1-I2V-14B video diffusion transformer. Beyond the noisy diffusion state, the model sees three conditions: clean source latents with blank query slots, dense Plücker ray maps encoding each pixel's camera, and the routed geometry features. These are concatenated and mapped into the pretrained token space by an expanded input projection; new camera and geometry inputs are initialized so they have no effect at the start of fine-tuning. No scene-specific text is used — training and inference share one fixed empty textual context.
- Using geometry as routing, not as reconstruction. A frozen VGGT-Ω assigns each source feature a 3D point and a confidence. Features are projected into each query camera using its extrinsics, focal lengths, and principal point, and routed to the corresponding query-grid location, transferring source-observed appearance evidence rather than pre-rendered RGB or explicit geometric predictions. A depth-selected hard anchor (z-buffered, confidence-weighted) provides the primary condition; a zero-initialized residual refiner adds confidence-weighted front and secondary hypotheses to handle occlusion boundaries and small geometric errors. Locations with no supporting source evidence are simply left for the diffusion prior to complete.
- Consistency supervision. Flow matching is applied only to query slots. On top of it, PTRC reuses VGGT-Ω point predictions to build tracks between every pair of query views, keeping a point only when it projects inside both views, is in front of both cameras, and passes a z-buffer visibility test. The smoothed-L1 difference between the two views' predicted-clean residuals at those track locations is penalized, weighted by normalized VGGT-Ω confidence. The total loss is flow matching plus λ_PTRC·PTRC, with λ_PTRC = 0.1.
- Training robustness. Geometry conditions are randomly retained, attenuated, or removed during training so the model never treats imperfect projections as an infallible scene reconstruction. Geometry-Prior CFG at inference uses the dropped-condition branch as a reference and extrapolates the difference between fully conditioned and reference velocity in early denoising.
- Training recipe. The VAE and VGGT-Ω are frozen; the Wan DiT, expanded input projection, visual-feature adapter, and Visual Geometry Router are trained. Training uses 980 scenes from the 1K-scene DL3DV-10K split (one unavailable scene and 19 overlapping with the 140-scene DL3DV-Benchmark removed), for 147 epochs at 192×336 then 60 epochs at 480×832, in BF16 with global batch size eight. Learning rates are 10^-5 and 5×10^-6 across the two stages, with 10^-4 for new modules. A 6-source-to-variable-N curriculum moves from N ∈ {1,2,4} to N ∈ {4,8,12,16}.
- Evaluation. Comparison uses 6,188 DL3DV-Benchmark targets and zero-shot Mip-NeRF 360, with identical six-view sources and target cameras for all methods, each target generated independently using 50 flow-sampling steps, following the FrameCrafter protocol with outputs resized or center-cropped to 480×480. Metrics are PSNR, SSIM, LPIPS, and DreamSim; Table 1 uses LPIPS-AlexNet, while controlled ablations on the fixed 192-case protocol use LPIPS-VGG.
Why This Matters
Research impact. The paper reframes what geometry foundation models are for in generative NVS: not as a source of complete reconstructions or pre-rendered RGB, but as uncertainty-aware routing evidence that conditions and constrains a generative prior. It also demonstrates a consistency mechanism (PTRC) that regularizes errors rather than predictions, which the ablations show matters more for SSIM than any other single component. The cross-reconstructor evaluation with Pi3 — a model never used for conditioning — is a notable methodological step for arguing that consistency gains are not an artifact of the evaluator.
Real-world applications:
- 3D content capture from casual photos. Turning a handful of phone or drone photos into consistent novel views for AR/VR, product visualization, or real-estate walkthroughs.
- Autonomous driving and simulation. The authors are affiliated with XPeng Motors, and driving scenes involve exactly the sparse, wide-baseline viewpoint changes this work targets — useful for rendering training and validation scenarios from limited captures.
- Robotics and embodied simulation. Generating consistent views of a scene for policy training or for verifying a reconstructed environment.
- Film, VFX, and virtual production. Filling in camera angles that were never shot while staying anchored to the geometry actually captured on set.
Industry relevance. The result is achieved with ~1k training scenes and an off-the-shelf video diffusion backbone, which is a relatively practical recipe for industrial pipelines that cannot afford million-scene training budgets. The method is also a drop-in style adaptation: it preserves the pretrained image pathway and adds new modules, releasing code at https://github.com/chenkangjie1123/VGGT-Diff.
Future Directions
- Scaling data and optimization. The paper notes that additional half-resolution experiments on the full DL3DV-10K yield substantial gains from scaling data and optimization, detailed in the Appendix — suggesting the current 980-scene setting is not the ceiling.
- Robustness to even sparser inputs. Appendix Sec. A.10 examines robustness to fewer source views; how the geometry router behaves as the source set shrinks or geometry quality degrades is an open question.
- Extending geometry-routed multi-view generation. Sec. A.12 lists promising directions for extension, and Secs. A.3–A.5 examine joint multi-target generation and the feature depth/optimization scope of VGGT-Ω, implying the routing and consistency mechanisms could be tuned further.
- Continuous-trajectory generation. Sec. A.11 clarifies the VAE protocols used for quantitative evaluation versus continuous-trajectory generation, pointing toward longer, smoother generated view paths beyond the discrete target sets evaluated here.
- A caveat to address. The authors caution that cross-view metrics use references reconstructed from target RGB rather than physical scans, so they measure relative consistency and reconstructability rather than absolute 3D accuracy.
Target Audience
Researchers and engineers working on novel view synthesis, 3D reconstruction, and generative diffusion models, particularly those interested in grounding video diffusion priors with geometry foundation models. It is most valuable to readers already comfortable with diffusion transformers, camera projection math, and point-cloud evaluation metrics; practitioners building sparse-view capture or rendering pipelines with limited training data will also find the training recipe and ablation breakdown directly useful.
Authors’ abstract
We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.