Skip to content
AI.info

Research

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

Overview Research area: Computer vision — self-supervised representation learning for 3D geometry from multi-view images, specifically novel view synthesis (NVS) as a pretraining signal. Technical lev

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
arXiv
2610.03717
Published
2026-10-02
Authors
Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg

AI summary

Overview

Research area: Computer vision — self-supervised representation learning for 3D geometry from multi-view images, specifically novel view synthesis (NVS) as a pretraining signal.

Technical level: Advanced. The paper is written for readers comfortable with transformer architectures, attention mechanisms, camera projection geometry, and standard 3D vision benchmarks. Its core ideas (decoder capacity, prediction targets) are understandable at an intermediate level, but the method relies on PRoPE-style pose conditioning, RoPE2D, DINOv3 feature layers, and multi-task probing protocols.

Scope: The paper diagnoses why feed-forward NVS models produce weak transferable geometric representations and proposes SNAP, an encoder-decoder design that fixes two architectural choices, then evaluates it on five downstream geometric and robotics tasks.

What This Paper Is About

Novel view synthesis models must implicitly reason about 3D structure to render unseen viewpoints, so in principle they should learn reusable geometry representations without needing expensive annotated 3D data. In practice, existing encoder-based NVS methods produce features that transfer poorly to downstream geometric tasks. The authors argue this is not a supervision problem but an architectural one, and they build SNAP to show that limiting decoder expressivity and changing the reconstruction target produces much stronger geometric representations.

Key Contributions

  1. SNAP, a self-supervised multi-view spatial reasoning backbone that combines frozen 2D semantic latent spaces (DINOv3 ViT-B patch features) with an asymmetric NVS prediction objective and latent feature targets, learning geometrically transferable representations without dense 3D supervision.

  2. Mechanistic analyses that isolate the independent impact of decoder expressivity (via receptive field sweeps from 1x1 to 16x16) and prediction space (via sweeps over DINOv3 initialization layers and prediction target layers) on geometric transfer.

  3. Empirical evaluation across five tasks — point correspondence, visual localization, relative pose estimation, relative depth estimation, and robot manipulation — matching or exceeding self-supervised baselines and remaining competitive with geometry-supervised methods that require large-scale annotated 3D data.

  4. A demonstration that the architectural lessons generalize beyond a single teacher model: Appendix D reports correspondence improvements with DINOv3, SAM, and MAE teachers alike.

Main Findings

  • Decoder expressivity is inversely related to geometric transfer. As the decoder's receptive field (RF) grows from 1x1 to 16x16, feature reconstruction error falls but point correspondence error rises, and visual localization and pose estimation degrade monotonically. The authors describe this as the critical crossover: reconstruction quality is a poor proxy for representation quality.

  • Full decoder self-attention hurts. At RF = 16x16 the window spans the full feature map, matching the full target-view self-attention used in prior NVS models (LVSM, LagerNVS). At RF = 1x1, target-query self-attention is reduced to a pointwise residual MLP, and similarity maps become sharply localized and stable across viewpoints.

  • Latent targets outperform pixel targets. Lower-layer DINOv3 targets are easier to predict (best reconstruction error came from the input:9, target:2 pairing) but yield worse downstream geometry. Predicting higher-layer DINOv3 features provides a better supervisory signal for transferable spatial structure.

  • An input-layer trade-off. Input layer 4 gave the best in-distribution PCK and MRR, improving over layer 11 by roughly 5.8% and 5.2% respectively, but its representation had approximately 53% higher intrinsic dimensionality (measured with the TwoNN estimator). The authors chose layer 11 as input, accepting slightly lower in-distribution accuracy for a more compressed, invariant representation.

  • Point correspondence (Table 1). On RealEstate10K, SNAP reached 15.48 PCK@5px, 36.29 PCK@10px, and 18.41 APE, outperforming VGGT (12.85 / 32.78 / 21.79) and Muskie (12.09 / 34.43 / 20.90) while operating at 256x256 versus VGGT's 518x518. On DL3DV, SNAP was competitive with VGGT (4.20 / 13.33 / 39.26 versus 4.15 / 13.65 / 39.34) using no dense 3D supervision. NVS baselines (ERayZer, LVSM, RayZer) consistently underperformed.

  • Visual localization (Table 2). On RE10K, SNAP achieved R5 80.32, T5 15.47, A20 26.41, A30 34.72, establishing a frontier among self-supervised models on most metrics.

  • Relative pose estimation (Table 2). On RE10K, SNAP recorded R5 78.81, T5 41.42, A20 53.07, A30 62.30 — high rotation accuracy among NVS-based methods and competitive on translation, showing the local decoder bottleneck does not degrade large-scale geometric inference.

  • Relative depth (Table 3). On Hypersim, SNAP achieved AbsRel 0.7093 and δ1 64.6815, outperforming NVS-based encoders; VGGT with explicit 3D supervision was better (0.2080 / 94.6036). On the real-world 7Scenes dataset, SNAP reached AbsRel 0.9257 and δ1 52.0218.

  • Robot manipulation under viewpoint shift. On three MimicGen tasks under camera shifts from 0.0 to 1.0 radians, SNAP outperformed its own frozen DINOv3 teacher, generally outperformed the task-specific baseline Adapt3R, and tracked closely with the geometry-supervised VGGT. Semantic encoders (DINOv3, ResNet) collapsed rapidly beyond 0.2 radians.

  • Memory efficiency. With 256x256 images and patch size 16, SNAP's encoder processes 256N+8 tokens including registers, versus VGGT's 1369N+5 tokens at 518x518 with patch size 14. The authors state it would require more than 12 views for SNAP to approach these baselines' attention cost. GQA is configured with 16 query heads and 8 KV-groups.

  • The gradient term is not auxiliary. Without the finite-difference loss term, the decoder can satisfy the feature loss with spatially-uniform predictions, collapsing the geometric signal the encoder must encode.

Methodology in Plain English

SNAP works in two stages. In the first stage, a pose-free multi-view encoder takes N reference images and runs each through a frozen DINOv3 ViT-B featurizer. The resulting patch features are projected to the transformer working dimension, tagged with 2D image-grid positions via RoPE2D, and processed by full cross-view self-attention so that every patch can attend to every patch in every view. Critically, no camera parameters are given at this stage. The output is a unified set of scene tokens Z that must implicitly encode the 3D scene.

In the second stage, a small pose-conditioned decoder is asked to predict the DINOv3 feature map of a novel target view. It does this by cross-attending a broadcast learned query against Z, with queries and keys rotated by PRoPE using the target and context camera intrinsics and extrinsics. This makes attention score high when world-space rays are co-aligned, behaving like a soft epipolar search.

Two constraints do the heavy lifting. First, a BlockNN mask forbids target queries from attending to each other outside a small spatial block; the optimal setting is a 1x1 block, which eliminates target-view spatial communication entirely. Because the decoder cannot smooth or hallucinate locally, it must rely on the encoder, which pressures geometric structure into Z. Second, supervision happens in DINOv3 feature space rather than pixel space, using an L1 feature loss plus a finite-difference gradient loss over the token grid, with both predicted and target features normalized onto the unit hypersphere.

For evaluation, the encoder is frozen and only lightweight task-specific heads are trained for five tasks. The pretraining mixture totals roughly 117,000 sequences: about 67,000 from RealEstate10K, about 10,000 from DL3DV-10K, and about 40,000 from Co3Dv2, with 20% of DL3DV and Co3Dv2 held out. Training samples up to N=6 reference views and one target view per sequence at 256x256, using a 16-layer encoder with embedding dimension 1024 and 8 blocks of alternating PRoPE cross-attention and masked self-attention.

Why This Matters

Impact on research. The paper reframes a common assumption: that better reconstruction implies better representations. It shows the two can be in direct tension, and that the fix is architectural rather than more data or more supervision. It also offers a path to geometric pretraining that avoids expensive dense 3D annotation, which is the main bottleneck for geometry foundation models like DUSt3R, MASt3R, VGGT, and D4RT.

Real-world applications:

  • Robotics manipulation, where policies built on frozen SNAP features remained robust under camera shifts of up to 1.0 radian while standard 2D semantic backbones collapsed beyond 0.2 radians.
  • Visual localization for mobile devices and augmented reality, where a query image must be matched against a reference map and its absolute pose recovered.
  • Novel view synthesis for photo and video post-production, AR/VR content, and free-viewpoint video.
  • Dense 3D reconstruction from casually captured video, where SNAP offers an alternative to pipelines requiring structure-from-motion and dense ground truth.
  • Relative depth estimation on real-world held-out imagery such as 7Scenes, useful for scene understanding where absolute metric depth is unavailable.

Industry relevance. The lower memory footprint is directly practical: 256N+8 tokens versus 1369N+5 tokens for a comparable geometry-supervised model means the approach runs under limited compute budgets. Since the encoder is frozen and only small heads are trained for each task, deployment across multiple downstream applications can reuse a single pretrained backbone.

Future Directions

  • Dynamic video. The paper explicitly states that current pretraining relies on posed video sequences and static, textured scenes; extending to dynamic video is called a critical next step for establishing this as a foundation model.

  • Scaling laws. The authors note their findings are bound by a fixed compute and data budget that has yet to reveal the full scaling behavior of the architecture, leaving open how performance grows with more data and compute.

  • Mapping the expressivity/representation frontier. The receptive field sweep covers 1x1 through 16x16, and the layer sweep covers DINOv3 layers 0-11; how these trade-offs behave for other decoder designs, larger models, or different teachers (SAM and MAE were tested only in Appendix D) remains open.

  • Beyond static, textured indoor and object scenes. The training mixture draws on RealEstate10K, DL3DV, and Co3Dv2, so generalization to outdoor, untextured, or extreme-baseline conditions is not established.

Target Audience

Researchers and engineers working on 3D vision, novel view synthesis, self-supervised representation learning, and geometric foundation models will find the architectural analysis most valuable. Robotics researchers interested in viewpoint-robust visual backbones for imitation learning are a secondary audience, as are practitioners who need strong geometric features without access to large annotated 3D datasets. Readers without background in transformer attention or camera geometry will find the mechanism sections dense, though the core argument about decoder capacity is accessible without that background.

Authors’ abstract

This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. https://snap-nvs.github.io

Read the original paper