Research
LVSPM: Long Sequence View Synthesis and Pose Estimation Model
LVSPM: Long Sequence View Synthesis and Pose Estimation Model Overview Research area: Computer vision / 3D vision — specifically pose-free novel view synthesis and camera pose estimation from uncalibr
- arXiv
- 2610.10960
- Published
- 2026-10-07
- Authors
- Xi Chen, Yachi Zhang, Linghao Chen, Minghua Liu, Hao Su, Zexiang Xu, Xiaoshuai Zhang
AI summary
LVSPM: Long Sequence View Synthesis and Pose Estimation ModelOverview
Research area: Computer vision / 3D vision — specifically pose-free novel view synthesis and camera pose estimation from uncalibrated image collections.
Technical level: Intermediate. The paper builds on ideas from neural rendering (NeRF, 3D Gaussian Splatting, LVSM) and geometry-based transformers (DUSt3R, VGGT), so familiarity with camera poses, Plücker rays, and feed-forward 3D reconstruction helps, though the core idea is explained conceptually.
Scope: The paper introduces LVSPM, a single feed-forward model that jointly predicts camera poses and renders novel views from long, unposed image sequences using only RGB image and camera pose supervision.
What This Paper Is About
Nearly all view synthesis methods assume camera poses are already known, typically obtained from structure-from-motion (SfM) pipelines that are expensive and fragile in the real world. Existing pose-free alternatives are usually limited to sparse inputs (often fewer than ten views) or require dense 3D supervision such as point maps, which are costly to obtain. LVSPM aims to solve both problems at once: jointly estimate world-aligned camera poses and synthesize novel views from unposed sequences while scaling to hundreds of input views, trained with only RGB images and camera poses.
Key Contributions
-
A unified pose-free model that scales to 256 input views. LVSPM is presented as the first pose-free view synthesis model to scale to 256 input views, built on a LaCT (Large-Chunk Test-Time Training) backbone with learnable camera tokens that predicts camera poses and novel views in a single sequence-to-sequence framework.
-
Novel view synthesis supervision as a substitute for dense 3D labels. The authors show that NVS supervision supplies enough geometric signal for accurate pose estimation. Their ablation reports that removing NVS supervision reduces AUC3 from 71.9 to 50.9, while the full model surpasses geometry-based methods trained with dense point-map supervision.
-
An equi-temporal evaluation protocol. Input views are sampled from the first N/N_max portion of a video sequence, so that more views correspond to larger scene coverage rather than denser sampling of the same region, isolating scene-scale effects from sampling density.
-
State-of-the-art results on RE10k, Co3Dv2, and DL3DV for both long-sequence pose estimation and view synthesis, surpassing even pose-dependent baselines such as DepthSplat while scaling to large scenes where prior methods collapse.
Main Findings
-
Pose estimation on RealEstate10k: At 16 views, LVSPM reaches 40.0 AUC3 and 53.8 AUC5 versus VGGT's 14.0 AUC3 and 25.3 AUC5 (described as nearly 3× and 2× the precision). At 128 views, the gap widens to 61.3 vs. 20.6 AUC3 and 72.5 vs. 34.1 AUC5. Against DA3, LVSPM achieves higher AUC30 and AUC5 across all view counts and wins 10 out of 12 metrics, while using only RGB images and camera poses rather than dense geometric supervision. Flare is third-best at longer sequences with 35.9 AUC3 at 128 views.
-
Pose estimation on Co3Dv2: LVSPM outperforms VGGT at all metrics — 39.2 vs. 33.9 AUC3 and 50.1 vs. 45.2 AUC5 at 16 views, and 68.8 vs. 62.1 AUC3 and 76.6 vs. 71.4 AUC5 at 128 views. It surpasses DA3 on AUC3 at 64/128 views and AUC5 at 128 views, though DA3 remains a strong reference particularly in the sparse-view setting.
-
Pose estimation on DL3DV-10K: At 16 views, LVSPM achieves 88.4 AUC3 versus VGGT's 74.7 (+13.7 absolute) and DA3's 84.0, with 91.6 AUC5 versus 81.9 and 88.7. At 256 views, LVSPM maintains 90.1 AUC3 versus 83.6 (VGGT) and 83.8 (DA3), while other baselines collapse: Fast3R drops to 8.0 AUC3 and Cut3R to 19.2 AUC3.
-
Novel view synthesis on DL3DV-10K: At 16 views, LVSPM reaches 25.77 PSNR, surpassing the pose-dependent DepthSplat (24.80) and the pose-free AnySplat (23.15). Under the equi-temporal protocol, LVSPM's PSNR decreases from 25.77 to 22.09 (−3.68 dB) across a 16× increase in scene scale, while AnySplat drops from 23.15 to 18.55 (−4.60 dB) and DepthSplat collapses from 24.80 to 18.39 at just 64 views (−6.41 dB) before failing due to model constraints. The gap over AnySplat widens from +2.62 dB at 16 views to +3.54 dB at 256 views, and LPIPS behaves similarly (ours 0.152 → 0.250 vs. AnySplat 0.158 → 0.339).
-
Ablation on NVS supervision: Removing RGB/NVS supervision on ASE causes AUC3 to fall from 71.9 to 50.9, which the authors describe as a nearly 30% degradation, indicating that NVS supervision effectively replaces dense 3D supervision.
-
Ablation on chunk size and depth: Halving the TTT chunk size from 8224 to 4112 tokens with two updates per layer causes a slight NVS drop and noticeable pose degradation (67.9 AUC3 vs. 71.9 for the full model). Reducing from 24 to 12 TTT layers yields 59.3 AUC3. The authors hypothesize that repeated updates on smaller chunks promote catastrophic forgetting.
-
Ablation on synthetic pretraining: On DL3DV, training without synthetic pretraining ("w/o Syn") yields 79.9 AUC3 and 20.08 PSNR, versus 80.2 AUC3 and 21.12 PSNR with the full setup.
-
Inference speed: LVSPM scales linearly with input views and processes 256 images in under 2 seconds on an H100 GPU, whereas baselines require over 10 seconds. It renders at 58.8 FPS regardless of the number of target views.
-
Zero-shot and sparse-view evaluation: LVSPM generalizes zero-shot to Tanks and Temples and MipNeRF360 under dense (64–128 views) and sparse (3–6 views) settings, and to RE10K, MipNeRF360, and LLFF for sparse-view evaluation, where it is compared with MVSplat (posed) and NoPoSPlat (unposed). The paper reports that LVSPM surpasses AnySplat on RE10K zero-shot, and when fine-tuned on RE10K it outperforms baselines trained exclusively on that dataset; full tables are in the supplementary material.
Methodology in Plain English
LVSPM treats pose estimation and view synthesis as one sequence-prediction problem rather than using explicit 3D structures. Each input image is split into patches and projected into tokens by a linear layer. Since the input poses are unknown, the model does not use Plücker ray encoding for input views (unlike pose-required models); instead, every input image gets a learnable "camera token" that carries pose information. The first view is the canonical reference frame with zero translation and identity rotation and uses a special token, while other views share a different token. Scene scale is normalized by setting the distance between the first and furthest points to unit length.
Target views use Plücker ray tokens computed from their camera parameters, plus another learnable token, so the model knows where to render. All tokens pass through 24 LaCT blocks, each combining per-view windowed self-attention, an MLP-based test-time training (TTT) layer, and a feed-forward MLP. The TTT layer updates its internal MLP weights on the fly using keys and values from input tokens only — target tokens are read-only. That asymmetry means the TTT hidden state acts as a compressed scene representation queried by camera tokens (for pose) and ray tokens (for rendering), and each novel view can be rendered independently, which is why rendering speed is independent of the number of targets.
Prediction is deliberately lightweight: a two-layer MLP decodes target tokens into RGB, and a separate two-layer MLP decodes each camera token into a 9-dimensional output (4-dim quaternion, 3-dim translation, 2-dim horizontal and vertical fields of view). Training mixes photometric loss (MSE plus LPIPS), pose loss (L1 on translation and rotation), and intrinsic loss (L1 on fields of view). The model (312M parameters, dimension 768, 12 attention heads) is trained on 64 NVIDIA H100 GPUs over three stages: synthetic ASE pretraining for 90k iterations at 32 input/target views and 128×128 resolution; a 60k-iteration mix of ASE with DL3DV-10K, ScanNet++, Hypersim, and Co3Dv2; and a final stage at 512×288 with 128 input and 64 target views.
Why This Matters
Impact on research: The paper argues that view synthesis supervision alone can replace dense 3D labels (point maps, depth) for accurate pose estimation, which lowers the supervision cost of building 3D perception systems. It also demonstrates that TTT-based architectures generalize from 128 training views to 256 inference views without finetuning, and introduces the equi-temporal protocol as a cleaner benchmark for long-sequence methods.
Real-world applications:
- Reconstructing and rendering large real-estate or architectural scenes from casually captured photo/video collections where SfM preprocessing is impractical.
- Immersive content capture, where long sequences are needed for wide scene coverage and free-viewpoint rendering.
- Robotics and AR/VR scene understanding from unposed image streams, reducing dependence on fragile camera calibration.
- Content creation from uncalibrated mobile footage, since the model renders at 58.8 FPS and processes 256 images in under 2 seconds.
Industry relevance: The method needs only RGB images and camera poses for training, avoiding expensive dense 3D annotation pipelines, which matters for companies building scalable 3D capture, mapping, or spatial computing products.
Future Directions
- Fully self-supervised training to further reduce supervision requirements, as the authors propose.
- Robustness in hard capture conditions — performance may degrade under extreme lighting changes or in weakly textured scenes where reliable geometric reasoning is difficult.
- Closing remaining gaps with specialized geometry models, which the authors acknowledge may still win in certain edge cases despite LVSPM's joint pose-plus-synthesis capability.
- Reducing training pipeline complexity, since the multi-stage training process introduces additional complexity.
Target Audience
Researchers and engineers working on 3D vision, novel view synthesis, and camera pose estimation, particularly those interested in pose-free feed-forward models and long-sequence scaling. It is also relevant to practitioners building 3D capture or spatial computing systems who want to avoid SfM preprocessing and dense 3D labels, and to readers following test-time training architectures such as LaCT.
Authors’ abstract
We present LVSPM, a generalizable model that jointly estimates camera poses and synthesizes novel views from uncalibrated image collections. Trained with only RGB images and pose supervision, LVSPM avoids dense 3D ground truth and employs test-time training (TTT) layers to scale seamlessly to hundreds of input views. On RealEstate10k, Co3Dv2, and DL3DV, LVSPM surpasses VGGT in pose estimation across 16-256 views, with especially large margins at strict thresholds. For novel view synthesis under a practical protocol where more views cover larger scenes, LVSPM achieves state-of-the-art pose-free quality---surpassing even pose-dependent models in PSNR---and still maintains high quality as scene scale grows, while baselines collapse. The code is available at https://burningdust21.github.io/Projects/LVSPM .