Skip to content
AI.info

Research

Visual Implicit Geometry Transformer for Autonomous Driving

Overview Research area: Computer vision for autonomous driving — specifically 3D geometric perception, multi-view reconstruction, and occupancy estimation from surround-view cameras. Technical level:

Visual Implicit Geometry Transformer for Autonomous Driving
arXiv
2602.05573
Published
2026-02-05
Authors
Arsenii Shirokov, Mikhail Kuznetsov, Danila Stepochkin, Egor Evdokimov, Daniil Glazkov, Nikolay Patakin, Anton Konushin, Dmitry Senushkin

AI summary

Overview

Research area: Computer vision for autonomous driving — specifically 3D geometric perception, multi-view reconstruction, and occupancy estimation from surround-view cameras.

Technical level: Advanced. The paper assumes familiarity with Vision Transformers, cross-attention, bird's-eye-view (BEV) representations, volume rendering, and occupancy prediction benchmarks.

Scope: The paper presents ViGT, a calibration-free, self-supervised transformer that estimates a continuous 3D occupancy field from surround-view camera images and renders it into point clouds, occupancy fields, or voxel grids.

What This Paper Is About

Autonomous driving requires accurate metric 3D geometry, but existing approaches either depend on precise camera calibration, require expensive 3D voxel annotations, or produce pixel-aligned depth/point predictions that are hard to reconcile across views. The authors set out to build a single, scalable model that infers a scene-centric, metric-scale 3D occupancy field directly from multiple camera images — without calibration and without manual labels — and that transfers across different sensor rig configurations.

Key Contributions

  1. A calibration-free image-to-BEV transformer (ViGT). A ViT-Large image encoder feeds a calibration-free "implicit BEV projection" module that learns the camera-to-BEV mapping from data via cross-attention rather than using known intrinsics and extrinsics, followed by a query-based implicit decoder that outputs occupancy probabilities for arbitrary 3D points.

  2. A continuous, scene-centric occupancy representation. Instead of voxel grids or per-view depth maps, the model learns a compact continuous BEV field that can be queried at any point and rendered into multiple downstream representations, aligning geometry from all cameras into one metric coordinate frame.

  3. A scalable self-supervised training recipe. Training uses only synchronized multi-view images and LiDAR, converting LiDAR rays into free-space (negative) and surface (positive) query labels, with a stratified plus symmetric-interval sampling strategy that emphasizes object boundaries.

  4. Large-scale validation across five datasets. The model is trained jointly on NuScenes, Waymo, NuPlan, ONCE, and Argoverse 2 with balanced sampling, and is evaluated on Occ3D-NuScenes and on pointmap estimation across all five datasets.

Main Findings

  • Best average rank on pointmap estimation. Across the five datasets and both metrics, ViGT achieves an Average Rank of 1.8, ahead of Cut3R (3.5), VGGT (3.9), Stream3R (4), Monst3R (4.8), Mast3R (6.4), DA3 (6.4), DUSt3R (6.9), and RenderOcc (7.3). Baselines were evaluated with depth-map (D) or pointmap (P) rendering; ViGT and RenderOcc use points rendered (PR) from predicted occupancy fields.

  • Strongest results on NuScenes and Argoverse 2 pointmap estimation. On NuScenes, ViGT reaches AbsRel 0.068 and Chamfer Distance 1.807, improving on the second-best method (Stream3R with depth maps) by 0.105 in AbsRel and 1.83 in CD. On Argoverse 2 it reaches AbsRel 0.131 (outperforming the best competitor by 0.031) and CD 2.965 (outperforming the best competitor by 0.716).

  • Mixed but competitive results elsewhere. On Waymo, ViGT obtains the second-best CD of 2.431, trailing the best method by only 0.05, with AbsRel 0.121. On ONCE it has the best AbsRel (0.169) and third-best CD (5.821). On NuPlan it has the best AbsRel (0.118) and second-best CD (3.298).

  • Competitive on Occ3D-NuScenes without any 3D annotation or calibration. ViGT achieves F1 0.7115 and IoU 0.5658, the third-best overall result on the benchmark and the best among self-supervised and 2D-supervised methods. PanoOcc (F1 0.8347, IoU 0.7271) and FB-Occ (F1 0.8181, IoU 0.7022) lead overall but use dense 3D ground truth; RenderOcc reaches F1 0.6442 / IoU 0.4824, Sparse-Occ 0.6271 / 0.4680, Offset-Occ 0.6240 / 0.4637, and Self-Occ 0.6552 / 0.4960. Relative to Self-Occ, ViGT improves F1 by 0.056 and IoU by 0.07, and is the only calibration-free entry.

  • The implicit BEV projection learns geometrically correct mappings. Attention visualization shows consistent forward (camera-to-BEV) and inverse (BEV-to-camera) correspondences, and the attention responses partition BEV space according to each camera's field of view for rigs of 6 cameras (NuScenes) and 7 cameras (Argoverse 2).

  • Architecture ablations favor two cross-attention blocks and the last four encoder layers. On NuScenes-only ablation training (4 GPUs, 100k iterations), CA + CA gives CD 2.699 / F1 0.713 / IoU 0.566 versus CD 3.051 / F1 0.697 / IoU 0.547 for a single cross-attention block and CD 2.814 / F1 0.697 / IoU 0.548 for CA + SA + CA. Using the last four ViT-L layers (CD 2.699) outperforms layers 5, 11, 16, 23 (CD 3.093, F1 0.70, IoU 0.55).

  • Symmetric negative sampling near surfaces matters. Stratified + symmetric interval sampling achieves CD 2.699 / F1 0.713 / IoU 0.566, beating random sampling (2.895 / 0.701 / 0.552) and plain stratified sampling (3.039 / 0.697 / 0.547).

  • Robustness to missing cameras. Appendix results show consistent occupancy predictions within visible regions when some or most surround-view cameras are removed.

Methodology in Plain English

The model has three parts. First, a ViT-Large encoder processes each camera image independently and keeps the token features from the last four layers, giving four token sequences per image. Second, a calibration-free projection module moves those tokens into bird's-eye-view space: for each encoder layer, a set of BEV queries attends to the image tokens through two sequential cross-attention blocks, and the resulting four BEV representations are merged and upsampled with DPT into one 256×256, 256-channel BEV feature map. No camera intrinsics or extrinsics are used anywhere in this step — the geometry is learned from data. Third, an implicit decoder (inspired by ImplicitIO and built on a Convolutional Occupancy Network) takes any 3D query point, interpolates its feature from the BEV grid, concatenates the normalized coordinates, and outputs an occupancy probability.

Training needs no human labels. Each LiDAR ray defines two classes of points: those between the sensor and the reflecting surface are free space, and those within a thin shell of thickness τ (set to 0.1 m) at the surface are occupied. Negatives are sampled by splitting the free segment [0, d) into K = 5 bins and sampling uniformly inside each bin, plus an extra symmetric set drawn from [d − τ, d) near the surface so the model learns object boundaries. Each sample uses 150K positive and 150K negative query points (120K from stratified bins, 30K from the symmetric interval). The loss is plain binary cross-entropy.

Optimization uses AdamW for 200K iterations with a cosine schedule, peak learning rate 5×10⁻⁵, 10K warmup iterations, batch size 6 per GPU, shortest image side resized to 192 pixels, and gradient norm clipping at 1.0, running on 128 A100 GPUs for five days. Because the output is a continuous field rather than a fixed grid, it can be rendered into whatever form a task needs: point clouds by volumetric integration along rays, or binary voxel grids by sampling several points per voxel and taking the maximum occupancy probability.

Why This Matters

Impact on research. The paper argues that domain-specific geometric foundation models for driving should be scene-centric and metric rather than pixel-aligned and scale-invariant. It shows that a calibration-free transformer trained only on image-LiDAR pairs can be competitive with supervised occupancy methods and can lead on pointmap estimation across five datasets, suggesting that explicit geometric inductive biases and manual annotation may not be necessary for scalable 3D perception.

Real-world applications:

  • Collision checking and motion planning, which need explicit free-space versus occupied-space estimates in a metric coordinate frame.
  • Multi-sensor fleet deployment, where a single model must work across vehicles with different camera counts and mounting positions — the paper demonstrates 6-camera (NuScenes) and 7-camera (Argoverse 2) rigs.
  • Annotation cost reduction: supervision comes from LiDAR that autonomous platforms already collect, avoiding voxel-level semantic labels.
  • Fault tolerance: predictions stay consistent within visible regions when cameras drop out, relevant to sensor degradation in the field.

Industry relevance. The calibration-free, self-supervised design targets exactly the bottlenecks that make perception stacks expensive to scale — per-vehicle calibration and per-dataset labeling — and the authors released source code at https://github.com/whesense/ViGT.

Future Directions

  • Extending the representation to semantics. The evaluation deliberately collapses the 17 semantic categories of Occ3D-NuScenes into a single occupied label; recovering semantic classes from this self-supervised field is an open step.
  • Temporal modeling. The paper notes that prior occupancy work has extended to occupancy forecasting over time, but ViGT as presented estimates a single-scene field; adding temporal context is a natural extension.
  • Closing the gap to fully supervised methods. ViGT trails PanoOcc and FB-Occ on Occ3D-NuScenes (F1 0.7115 versus 0.8347 and 0.8181), so improving self-supervised accuracy to match 3D-annotated training remains open.
  • Efficiency and deployment characteristics. Training cost (128 A100 GPUs for five days) is substantial, and the paper does not report inference latency, parameter count, or memory footprint — figures that would be needed to judge on-vehicle feasibility.

Target Audience

This paper is aimed at computer vision and autonomous driving researchers working on 3D perception, multi-view reconstruction, occupancy prediction, and BEV representation learning. It is also relevant to engineers building production driving stacks who are concerned with calibration overhead, annotation cost, and generalization across sensor configurations. Readers need a working understanding of transformers and 3D geometry; the paper is not an introductory treatment.

Authors’ abstract

We introduce the Visual Implicit Geometry Transformer (ViGT), an autonomous driving geometric model that estimates continuous 3D occupancy fields from surround-view camera rigs. ViGT represents a step towards foundational geometric models for autonomous driving, prioritizing scalability, architectural simplicity, and generalization across diverse sensor configurations. Our approach achieves this through a calibration-free architecture, enabling a single model to adapt to different sensor setups. Unlike general-purpose geometric foundational models that focus on pixel-aligned predictions, ViGT estimates a continuous 3D occupancy field in a birds-eye-view (BEV) addressing domain-specific requirements. ViGT naturally infers geometry from multiple camera views into a single metric coordinate frame, providing a common representation for multiple geometric tasks. Unlike most existing occupancy models, we adopt a self-supervised training procedure that leverages synchronized image-LiDAR pairs, eliminating the need for costly manual annotations. We validate the scalability and generalizability of our approach by training our model on a mixture of five large-scale autonomous driving datasets (NuScenes, Waymo, NuPlan, ONCE, and Argoverse) and achieving state-of-the-art performance on the pointmap estimation task, with the best average rank across all evaluated baselines. We further evaluate ViGT on the Occ3D-nuScenes benchmark, where ViGT achieves comparable performance with supervised methods. The source code is publicly available at \href{https://github.com/whesense/ViGT}{https://github.com/whesense/ViGT}.

Read the original paper