Research
Unified Sensor Simulation for Autonomous Driving
Unified Sensor Simulation for Autonomous Driving (XSIM) Overview Research area: Computer vision / 3D scene reconstruction and sensor simulation for autonomous driving, built on 3D Gaussian Splatting (

- arXiv
- 2602.05617
- Published
- 2026-02-05
- Authors
- Nikolay Patakin, Arsenii Shirokov, Anton Konushin, Dmitry Senushkin
AI summary
Unified Sensor Simulation for Autonomous Driving (XSIM)Overview
Research area: Computer vision / 3D scene reconstruction and sensor simulation for autonomous driving, built on 3D Gaussian Splatting (3DGS) and the 3DGUT projection framework.
Technical level: Advanced. Understanding the contributions requires familiarity with Gaussian Splatting, camera projection models, rolling-shutter sensing, and the Unscented Transform.
Scope (one sentence): The paper presents XSIM, a unified sensor simulation framework that renders both rolling-shutter cameras and spinning LiDAR in a single formulation by extending 3DGUT splatting with phase modeling and dual-opacity Gaussians.
What This Paper Is About
Autonomous driving needs large, diverse sensor data, but collecting real-world driving data is expensive and slow. Simulating new views and trajectories from existing driving logs is a cheaper alternative, and 3D Gaussian Splatting is a strong engine for this. The problem is that real driving sensors are not simple pinhole cameras: cameras and LiDARs both capture scenes in rolling-shutter fashion over time, and spinning spherical LiDARs break the standard projection assumption when a Gaussian particle spans the azimuth boundary. XSIM's goal is to render these complex sensors accurately and consistently within one framework.
Key Contributions
- XSIM framework. A sensor simulation framework for autonomous driving that extends 3DGUT splatting and renders LiDAR and camera sensors in a unified manner with generalized rolling-shutter modeling.
- Phase modeling for spherical rasterization. A mechanism that explicitly accounts for temporal and shape discontinuities of Gaussians projected by the Unscented Transform at azimuth borders, where a single particle may project into two separate 2D Gaussians with different covariances and depths.
- Extended 3D Gaussian representation. Each Gaussian is given two distinct opacity parameters — one for camera (σ_c) and one for LiDAR (σ_L) — jointly optimized and regularized toward consistency, addressing mismatches between geometry and color distributions.
- Comprehensive evaluation. Experiments on three autonomous driving datasets (Waymo Open Dataset, Argoverse 2, PandaSet) showing state-of-the-art results on RGB image metrics and LiDAR Chamfer Distance.
Main Findings
- Consistent state-of-the-art results. XSIM reports the best performance across all three datasets in both scene reconstruction and novel-view synthesis settings, per Table 1.
- Waymo reconstruction gains. XSIM reaches 30.75 PSNR, 0.9030 SSIM, 0.2228 LPIPS and 0.08 CD, versus SplatAD at 27.74 PSNR, 0.8650 SSIM, 0.2807 LPIPS and 0.82 CD — gains of +3.01 PSNR and +3.8% SSIM, with LPIPS reduced by 20.6%.
- Novel-view synthesis gains. XSIM achieves +2.74 PSNR on Waymo and +1.04 PSNR on Argoverse over the previous state-of-the-art method. Waymo NVS numbers are 29.80 PSNR, 0.8904 SSIM, 0.2236 LPIPS, 0.18 CD.
- Large LiDAR error reductions. XSIM reports an x8.8 Chamfer Distance error reduction on Waymo reconstruction and an x4.5 error reduction on Waymo novel-view synthesis.
- PandaSet LPIPS is second-best. On PandaSet, LPIPS remains competitive and ranks second, with a minor gap to the best-performing method; XSIM's other metrics lead (reconstruction: 29.05 PSNR, 0.8839 SSIM, 0.1872 LPIPS, 0.20 CD).
- Azimuth-boundary artifacts are the motivation for phase modeling. Near the azimuth discontinuity border, standard 3DGUT projection produces partially missing and distorted range image renders; without phase modeling, a particle spanning φ = ±π projects into an overly large, incorrectly shaped 2D Gaussian.
- Ablation confirms each component matters (novel-view synthesis, six Waymo scenes, half of the split). Full model: 30.03 PSNR, 0.8945 SSIM, 0.2122 LPIPS, 0.21 CD. Removing camera rolling shutter drops to 28.95 PSNR; removing LiDAR opacity gives 29.55 PSNR / 0.25 CD; removing LiDAR rolling shutter gives 29.05 PSNR / 0.31 CD; removing phase modeling gives 29.32 PSNR / 0.28 CD.
- Rolling-shutter solve is cheap. The point-observation-time equation has no closed-form solution but is solved iteratively with the Newton–Raphson method, which the authors observe converges 1–2 iterations faster than fixed-point iteration.
- Qualitative improvements. XSIM better preserves characteristic LiDAR ring patterns, renders pedestrians accurately, and produces smooth, dense depth maps; it also renders more consistently when the ego-vehicle trajectory is laterally shifted by 3 meters.
Methodology in Plain English
Base engine. The framework builds on 3DGUT, which projects Gaussian particles onto arbitrary camera models using the Unscented Transform rather than the linearized single-point projection of EWA splatting. Approximate 2D conics are used only for assigning particles to image tiles; the actual volumetric integration is performed in 3D by evaluating each particle at its point of maximum response along the camera ray.
Generalized rolling shutter. Because sensor readings are captured row by row, every image point (u, v) has its own capture time τ(u, v) = τ_start + u·τ_u + v·τ_v. Both the camera and the dynamic actors are assumed to move with constant linear and angular velocities over the exposure, which gives closed-form expressions for how a world point and the camera pose change with time. Since the observation time of a point is unknown before projection, the authors solve for the time η that is consistent with the pixel's capture time using Newton–Raphson iterations.
Phase modeling. The azimuth of a spherical projection is periodic. Instead of keeping only the principal solution (k = 0), the authors explicitly consider k ∈ {−1, 0, +1}, performing auxiliary projections shifted by ±π. Each interval gets its own 2D Gaussian conics and depth estimates, and valid projections are passed to tiling. This removes false tile–particle intersections and reduces depth-sorting errors near azimuth boundaries.
Dual opacity. Accurate geometry wants a single opaque Gaussian per surface, while accurate appearance wants multiple semi-transparent Gaussians for specular and translucent effects. XSIM gives each Gaussian separate camera and LiDAR opacities, regularized toward each other with L_opacity = Σ|σ_c,i − σ_L,i|.
Training. Scene nodes (static, rigid, deformable, and SMPL-based human nodes) are optimized jointly from driving logs by sampling images and the closest-in-time LiDAR sweeps each iteration, using the loss L = λ·L1 + (1−λ)·L_SSIM + L_depth + L_opacity + L_reg, with λ = 0.2. Rendering uses custom CUDA kernels with a shared rasterization forward/backward pass across camera models. Gaussian initialization uses LiDAR sweeps, bounding box annotations, and camera colors, with inverse-distance sphere sampling to fill regions outside LiDAR coverage.
Why This Matters
Impact on research. The paper identifies a concrete failure mode of Unscented-Transform-based projection — cyclic azimuth and time discontinuities at the spherical-camera boundary — and offers an explicit fix. It also argues that a single opacity parameter is insufficient when geometry and appearance supervision come from different sensor modalities, which is a design lesson for any multi-sensor Gaussian representation.
Real-world applications:
- Generating augmented autonomous driving datasets for training and evaluation without new real-world data collection.
- Testing perception and planning algorithms against controlled ego-vehicle trajectory changes, such as the lateral 3-meter lane shift demonstrated in the paper.
- Simulating LiDAR and camera renders of the same scene consistently, which supports sensor-fusion development and cross-sensor validation.
- Producing smooth, dense depth maps from camera rendering, useful for geometry-dependent downstream tasks.
Industry relevance. Autonomous driving companies depend on simulation to reduce data-collection cost and to cover rare or dangerous scenarios. A framework that handles the actual distortions of shipping camera and LiDAR sensors — rather than idealized pinhole models — is more directly useful for sim-to-real work. The code is publicly available at https://github.com/whesense/XSIM.
Future Directions
- The paper assumes constant linear and angular velocities for the camera and all dynamic actors over the exposure; testing whether this assumption, and the Newton–Raphson solve, holds under more aggressive or non-constant motion is a natural next question.
- Phase modeling considers k ∈ {−1, 0, +1}; the paper does not report behavior for particles or sensors that would require a wider range of azimuth periods.
- Evaluation covers Waymo Open Dataset, Argoverse 2, and PandaSet; generalization to other sensor configurations and driving domains is not reported.
- On PandaSet, XSIM's LPIPS ranks second rather than first, and the paper gives no analysis of that gap — a useful target for follow-up work.
- The truncated content ends mid-description of the scene-node rendering pipeline, so the full set of regularization terms (L_reg) and the complete hyperparameter details are only referenced as being in the appendix.
Target Audience
Researchers and engineers working on 3D Gaussian Splatting, neural rendering, and autonomous driving simulation who already understand camera models, splatting rasterization, and rolling-shutter sensing. It is also useful for practitioners building synthetic data pipelines who want to know exactly which distortions must be modeled to make rendered LiDAR and camera data trustworthy. Readers without a graphics or Gaussian-splatting background will find the projection and Unscented Transform details demanding.
Authors’ abstract
In this work, we introduce \textbf{XSIM}, a sensor simulation framework for autonomous driving. XSIM extends 3DGUT splatting with a generalized rolling-shutter modeling tailored for autonomous driving applications. Our framework provides a unified and flexible formulation for appearance and geometric sensor modeling, enabling rendering of complex sensor distortions in dynamic environments. We identify spherical cameras, such as LiDARs, as a critical edge case for existing 3DGUT splatting due to cyclic projection and time discontinuities at azimuth boundaries leading to incorrect particle projection. To address this issue, we propose a phase modeling mechanism that explicitly accounts temporal and shape discontinuities of Gaussians projected by the Unscented Transform at azimuth borders. In addition, we introduce an extended 3D Gaussian representation that incorporates two distinct opacity parameters to resolve mismatches between geometry and color distributions. As a result, our framework provides enhanced scene representations with improved geometric consistency and photorealistic appearance. We evaluate our framework extensively on multiple autonomous driving datasets, including Waymo Open Dataset, Argoverse 2, and PandaSet. Our framework consistently outperforms strong recent baselines and achieves state-of-the-art performance across all datasets. The source code is publicly available at \href{https://github.com/whesense/XSIM}{https://github.com/whesense/XSIM}.