Skip to content
AI.info

Research

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Overview Research area: Computer vision, specifically novel view synthesis (NVS) and 3D scene generation from sparse images. Technical level: Advanced. The paper assumes familiarity with diffusion mod

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
arXiv
2609.35734
Published
2026-09-28
Authors
Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang, Mu Huang, Chunhua Shen, Mulin Yu, Bo Dai

AI summary

Overview

  • Research area: Computer vision, specifically novel view synthesis (NVS) and 3D scene generation from sparse images.
  • Technical level: Advanced. The paper assumes familiarity with diffusion models, latent spaces, 3D foundation models, Gaussian/point-cloud representations, and flow matching.
  • Scope: The paper introduces GeoVerse, a framework that generates novel views directly inside the geometric latent space of a 3D foundation model, augmented with frozen video-diffusion features and a persistent point-cloud memory to keep multi-view and multi-round synthesis consistent.

What This Paper Is About

Novel view synthesis has to do two things at once: faithfully reproduce parts of a scene that were observed, and plausibly invent parts that were never seen. Reconstruction-based methods are good at the first but fail at the second, while video generative models are good at the second but drift and lose 3D consistency when applied sequentially across many views.

GeoVerse tries to get both properties at once by running generation inside the latent space of a 3D geometric foundation model (DA3), while injecting appearance priors from a pretrained video diffusion model (Wan2.2 VACE) and anchoring every new prediction against a global spatial memory of everything already observed or generated.

Key Contributions

  1. A framework that bridges geometric latent diffusion and video generative priors. GeoVerse injects multilevel features extracted once from a frozen Wan2.2 VACE model into the geometric latent diffusion model through a ControlNet-style adapter with zero-initialized projections, so the pretrained denoiser is initially unmodified.

  2. A global spatial memory mechanism for long-sequence consistency. A colored point-cloud memory accumulates observed and synthesized content in a shared coordinate frame and is reprojected into each target view as target-aligned RGB-D guidance plus a validity mask, anchoring successive predictions and suppressing cumulative drift.

  3. Scaled training and few-step distillation. Training is expanded from 4 to 15 diverse real and synthetic multi-view datasets for cross-domain generalization, and reflow distillation reduces Level-1 generation to three denoising steps, with a 3 + 1 budget (three Level-1 steps, one Level-0 cascade step) cutting inference from over 100 seconds to under 10 seconds.

Main Findings

  • Visual quality: GeoVerse reaches state-of-the-art results across all four benchmarks. On DL3DV it reaches 19.61 PSNR versus 17.38 for GLD, a 2.23 dB gain. On Mip-NeRF 360 it reaches 20.02 PSNR versus 17.57 for GLD, a 2.45 dB gain. On RealEstate10K it reaches 21.29 PSNR versus 19.66 for GLD.
  • Geometric consistency: ATE on Mip-NeRF 360 drops from 0.105 (GLD) to 0.071, a 32.4% improvement. On DL3DV, ATE falls from 0.058 (GLD) to 0.028.
  • Long-sequence stability: On ScanNetV2 across three generation rounds, PSNR rises from 15.33 dB (GLD) to 19.65 dB, with key scene elements retaining cross-round consistency.
  • Inference efficiency: GeoVerse synthesizes six target views in roughly nine seconds on the three main benchmarks (9.24 s on DL3DV, 9.29 s on RealEstate10K, 9.12 s on Mip-NeRF 360), described in the experiments as about 18× faster than GLD and in the conclusion as over 17× speedup. GLD's reported times are 169.29 s, 163.50 s, and 156.01 s respectively.
  • Denoising budget trade-off (Table 2, DL3DV): Reducing Level-1 from three steps to one lowers runtime from 9.24 s to 6.07 s but worsens LPIPS from 0.339 to 0.367 and loses fine brick textures despite higher PSNR (20.55 vs 19.61), showing PSNR alone does not reflect perceptual quality. Reducing the Level-0 cascade from 49 steps to one yields LPIPS of 0.342 versus 0.339 at much lower latency (52.55 s versus 9.24 s for the cascade sweep).
  • Ablation, Wan2.2 prior: Removing the video prior causes the most severe visual degradation, dropping ScanNetV2 PSNR from 19.65 to 16.83 and worsening ATE from 0.009 to 0.022.
  • Ablation, spatial memory: Disabling the spatial memory drops PSNR to 17.87 and ATE to 0.014, degrading both visual fidelity and cross-view alignment.
  • Ablation, supervision losses: Omitting RGB supervision lowers PSNR to 18.58; removing depth supervision lowers PSNR to 19.05 and weakens geometric consistency (ATE 0.010).
  • Architecture alternative: A Mixture of Transformers video-KV variant reaches 19.24 PSNR with a scaled time of 36.29 versus 33.11 for the reference; the paper states that isolating its relative advantages requires comparison under matched training settings.
  • Where baselines still win: The paper notes that some baselines achieve lower reprojection or translation errors on individual datasets, while GeoVerse maintains better global scene coherence. For example, on RealEstate10K GLD reports a lower reprojection error (0.609 versus 0.631), and CAMEO is faster (14.28 s versus 9.29 s is not the case here — CAMEO is slower than GeoVerse on RealEstate10K but faster than other baselines on several benchmarks).

Methodology in Plain English

GeoVerse builds on Geometry Latent Diffusion (GLD), which performs diffusion directly inside the multilevel feature space of a frozen Depth Anything 3 (DA3) encoder rather than in a standard image or video latent space. GeoVerse keeps that setup but changes three things.

First, instead of GLD's zero-padded target conditioning, the input sequence merges context images with projections of a spatial memory, which the frozen DA3 encoder processes up to a chosen "synthesis boundary" — Level-1, corresponding to DA3 Transformer blocks (b0, b1, b2, b3) = (5, 7, 9, 11). Level-1 features are generated by a flow-matching denoiser whose velocity prediction is conditioned on camera Plücker-ray embeddings, Wan2.2 features, projected depth guidance, and a validity mask. A Level-0 cascade then recovers finer features.

Second, the model borrows appearance know-how from a frozen Wan2.2 VACE video diffusion model. Intermediate features are extracted in a single forward pass from blocks (k0, k1, k2, k3) = (0, 5, 10, 15), aligned in resolution and channel count, and injected as additive residuals through a ControlNet-style branch. Noise with σ = 0.1 is added only to valid target projections; context RGB stays unperturbed. Because the residual projections are zero-initialized, the pretrained denoiser starts out unchanged.

Third, a colored point-cloud memory accumulates the scene. Context observations initialize it; each generated round is scale-aligned via a camera-center similarity transform, back-projected into 3D points, and fused into the memory. For each target camera, depth-aware point splatting renders this memory into RGB-D hints and a validity mask, which are fed back as conditioning. Valid projections anchor known structure; masked regions tell the generative prior where to invent content.

Training supervises both the latents and the decoded outputs, using flow-matching losses at Level-1 and Level-0, a mean valid-pixel ℓ1 RGB loss, the same loss on high-gradient regions, and a depth loss after scale alignment. Context views are weighted α = 0.25 and target views α = 1. Training proceeds in two stages: a 300k-iteration adaptation phase at learning rate 3 × 10⁻⁵, then up to 50k iterations of reflow distillation at 1 × 10⁻⁵, on 32 NVIDIA A800 GPUs with AdamW, global batch size 32, (β1, β2) = (0.9, 0.95), gradient-norm clipping at 1.0, and EMA decay 0.9995, using DA3-Base at 504 × 504 resolution with eight ordered views per sample (one to four selected as context views). Evaluation uses N_C = 2 context views and N_T = 6 target views.

Why This Matters

The paper argues that the field has been stuck between two poles — rigid geometry that cannot imagine missing content, and expressive video generation that forgets the scene over time. GeoVerse is a concrete demonstration that a geometric latent space can host video-learned priors without giving up explicit camera control, and that a persistent 3D memory can carry consistency across many rounds of expansion rather than just a single pass. It also shows the approach is fast enough (roughly nine seconds per round) to matter practically, which most high-fidelity video-based alternatives are not.

Real-world applications:

  • Real estate and virtual tourism: Generating walkthroughs of properties or locations from only a handful of photographs, including rooms or angles that were never photographed.
  • Film and VFX previsualization: Producing consistent camera moves through a scene from sparse reference plates before committing to expensive shoots.
  • AR/VR and telepresence: Filling in the parts of a reconstructed environment that a user's head movement reveals but the original capture never covered.
  • Robotics and simulation: Synthesizing multi-view training data for embodied agents where only limited sensor viewpoints exist.

Industry relevance: the efficiency numbers matter commercially, since ~9 s per six-view batch is close to interactive, and the framework builds on frozen, off-the-shelf components (DA3, Wan2.2 VACE) rather than requiring models trained from scratch. Content creation, gaming, architecture, and any pipeline that consumes sparse captures into explorable 3D scenes could plausibly adopt it.

Future Directions

  • Joint geometric and photometric stability: The authors identify textureless or complex regions, where initial depth errors and generated appearance drift accumulate through memory updates, as the main remaining failure mode. Improving joint stability is stated as a key future objective.
  • Escaping the DA3 appearance ceiling: Peak visual fidelity is bounded by the appearance capacity of the DA3 feature space; the paper suggests that closing the gap to full-scale video diffusion models without losing the speed advantage is an open problem.
  • Better video-prior injection: The Mixture of Transformers alternative was not conclusively better or worse than additive residual injection, and the paper explicitly says its relative advantages need evaluation under matched training settings.
  • Path toward video world models: The conclusion frames the work as a foundation for building multi-view consistent video world models, which implies extending the approach from static scenes to dynamic, temporally evolving environments.

Target Audience

Researchers and graduate students working on novel view synthesis, 3D reconstruction, diffusion models, and generative 3D scene understanding. It is also relevant to practitioners building 3D content pipelines who need to know how much inference cost and consistency they can expect from a geometric-latent approach augmented with video priors. Readers without background in diffusion latent spaces, 3D foundation models, or camera pose metrics such as ATE and RPE will find the methodology sections demanding.

Authors’ abstract

Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.

Read the original paper