Research
Shoot-Bounce-3D: Single-Shot Occlusion-Aware 3D from Lidar by Decomposing Two-Bounce Light
Overview Research area: Computer vision and computational imaging — specifically single-photon lidar, inverse light transport, and single-shot 3D scene reconstruction (published at SIGGRAPH Asia 2025

- arXiv
- 2512.06080
- Published
- 2025-12-05
- Authors
- Tzofi Klinghoffer, Siddharth Somasundaram, Xiaoyu Xiang, Yuchen Fan, Christian Richardt, Akshat Dave, Ramesh Raskar, Rakesh Ranjan
AI summary
Overview
Research area: Computer vision and computational imaging — specifically single-photon lidar, inverse light transport, and single-shot 3D scene reconstruction (published at SIGGRAPH Asia 2025 Conference Papers).
Technical level: Intermediate to Advanced. The core ideas (time of flight, multi-bounce light, neural rendering) are explained in the paper, but the method assumes familiarity with lidar sensing, transient histograms, and neural reconstruction.
One-sentence scope: The paper introduces Shoot-Bounce-3D (SB3D), a learned pipeline that separates ("demultiplexes") the mixed two-bounce light in a single lidar capture taken with all laser spots firing at once, and uses that separated signal to recover metric depth, occluded geometry, and specular surfaces.
What This Paper Is About
Single-photon lidar can measure not only light that bounces once off a surface and returns to the sensor, but also light that bounces two or more times (multi-bounce light). That extra light carries information about depth, hidden/occluded objects, and mirrors. The problem is that prior methods which exploit this signal assume a laser scans one scene point at a time, whereas real consumer devices illuminate many points simultaneously (multiplexed illumination), which mixes all the signals together and breaks existing approaches. The goal of this paper is to learn, in a data-driven way, how to unmix that mixed signal so a single shot is enough to reconstruct 3D geometry, including areas hidden behind occluders and regions with specular materials such as mirrors.
Key Contributions
- Data-driven demultiplexing. From a single lidar measurement where multiple scene points are illuminated at once, the authors propose a learned method that decomposes the two-bounce signal into the separate contributions of two-bounce time of flight and shadows for each individual laser spot.
- Occlusion-aware 3D from one shot. They show that the demultiplexed two-bounce time of flight and shadow maps can be fed to an existing neural reconstruction method (PlatoNeRF) to render depth from novel views, revealing occluded areas despite specular materials in the scene.
- A large-scale multi-bounce lidar dataset. They build a dataset of roughly 100,000 simulated multi-bounce transient measurements of indoor scenes, which the paper states is the first large-scale transient dataset (prior datasets contained around 5,000 simulated transients).
- Generalizable multi-bounce lidar features. They find that the features learned for demultiplexing time of flight can be frozen and reused by a newly trained decoder to segment specular objects, which they present as a step toward a generalizable multi-bounce transient representation.
Main Findings
- Depth estimation on 6k test samples (MAE lower is better, F1 Boundary higher is better): Shoot-Bounce-3D achieves 0.0228 MAE and 0.6238 F1 Boundary, compared to Bounce Flash Lidar at 0.4922 / 0.0138, CompletionFormer at 0.4394 / 0.0066, Depth Anything V2 at 0.1640 / 0.1999, and Depth Pro at 0.1089 / 0.2930.
- Baseline weaknesses described by the authors: Bounce Flash Lidar produces noisy depth maps and cannot resolve depth for shadowed pixels because it has no learnable mechanism for denoising or infilling; CompletionFormer struggles with depth from only 25 points, and even with 100 points still reaches only 0.149 m MAE. Depth Pro produces qualitatively similar depths to ground truth but struggles to preserve scale even after rescaling with anchor points, and shows higher depth error around edges.
- Specular segmentation (Pixel MAE lower is better, IoU higher is better): Shoot-Bounce-3D reaches 0.0010 pixel MAE and 86.52% IoU, versus EBLNet at 0.0117 and 81.21% (EBLNet was retrained on the authors' dataset for a fair comparison).
- Occlusion-aware 3D reconstruction (MAE lower is better, F1 Boundary higher is better), averaged over four scenes with 80 novel test views each: Shoot-Bounce-3D reaches 0.0983 MAE / 0.2725 F1 Boundary, versus ZeroNVS at 0.5619 / 0.0090. A PlatoNeRF Oracle, which uses ground-truth inputs, reaches 0.0950 / 0.3317 — meaning the learned demultiplexing comes close to the oracle.
- Demultiplexing accuracy: The mean absolute test error for two-bounce pathlength was 0.2736 m, which the authors note is higher than depth error because two-bounce paths are longer. Predicted shadow maps reached 0.0214 pixel MAE and 95.3% IoU.
- "Light in flight" visualization: By combining predicted two-bounce time of flight and shadows into a transient measurement, the method can render light-propagation videos per illumination point.
- Transfer to a new task: Features learned for demultiplexing time of flight, when frozen and paired with a randomly initialized decoder, accurately predict specular segmentations in simulation, which the authors interpret as a sign the features may be a generalizable representation for single-photon lidar.
- Real-world proof of concept: The authors capture a new dataset with multiplexed illumination using 16 laser spots. Because PlatoNeRF cannot handle multiplexing, it is restricted to a single capture with 1 illumination point; the authors report their method performs better, especially in specular regions. A comparison to PlatoNeRF with more captures is deferred to the supplement.
Methodology in Plain English
The authors treat the problem as an unmixing task and split it into three stages rather than trying to directly invert the whole physical model, which they describe as highly ill-posed.
Stage one — depth, as a stand-in for time of flight. Predicting a separate two-bounce time-of-flight map for every laser spot would be an enormous output, so instead they train an encoder–decoder to predict dense depth from the raw transient cube. Depth is a lower-dimensional quantity, and once you know the depth of each pixel you can compute the two-bounce path length for each laser spot using simple ray geometry (the path laser → virtual source → scene point → sensor). They assume the laser spot position is known from the one-bounce return. The depth network is trained with a data fidelity loss combining an SSIM term and an L1 term with weight α = 0.15, plus an edge-aware smoothness term weighted by β = 10⁻³ that uses the time-integrated transient as a proxy for image edges, since no RGB image is available.
Stage two — shadows. Occlusions are detected through the shadows that hidden objects cast from each illumination spot. The authors found that directly predicting many binary shadow masks from the multiplexed measurement works poorly, and conditioning on the laser spot index also works poorly, because the network input measures the light that arrives while the output predicts light that is absent. They fix this mismatch with the concept of a "shadow transient": a calibrated capture (the scene with occluded objects removed) minus the actual measurement. Rather than removing objects physically, they estimate that calibrated capture from the predicted two-bounce time of flight using a sum of delta functions, setting the unknown two-bounce intensity to 1. Both the measured transient and this estimated calibrated capture are fed to the network, which is trained with a binary cross-entropy loss.
Stage three — 3D reconstruction. With per-spot two-bounce time-of-flight maps and shadow masks in hand, they train PlatoNeRF, an existing neural reconstruction model, without modification. PlatoNeRF normally needs a separate time-of-flight map and shadow mask per laser spot; SB3D supplies them from a single multiplexed measurement.
Extra task and data. They also freeze the pretrained time-of-flight encoder and train a randomly initialized decoder to predict binary specular segmentation masks with binary cross-entropy supervision. All of this is trained on their simulated dataset built on the Aria Synthetic Environments dataset, containing 97,432 scenes assembled from roughly 8,000 unique objects, with one 256×256 transient rendered per scene at 128 ps temporal resolution and renderings at 4, 25, and 100 illumination points. Experiments use 25 illumination points, with roughly 87k samples for training and 6k for test metrics.
Why This Matters
This is described as the first work to bring data-driven methods to two-bounce transients. It targets a mismatch between what laboratory lidar methods assume (one laser point at a time) and what consumer devices actually do (many points at once), and shows a learning-based route around a problem the authors argue is hard to invert analytically. The work also matters as an early example of applying large-scale data priors and representation learning to single-photon lidar, a sensor modality the authors note lacks the large-scale datasets and machine learning approaches that RGB images now have.
Real-world applications mentioned or implied by the framing:
- Autonomous vehicles, which need reliable 3D perception.
- Extended reality, where headsets need single-shot depth of the surrounding scene.
- Consumer devices such as mobile phones and tablets that already carry high-resolution SPAD sensors and use multiplexed point illumination (the paper cites the iPhone as an example of point illumination).
- Scenes containing mirrors, windows, and other specular surfaces, where RGB-based methods can hallucinate "portals" or depth holes.
Industry relevance: The author list spans MIT and Meta, and the paper is published at SIGGRAPH Asia 2025, connecting academic computational imaging research with industry work on consumer sensing hardware. The authors state that code and dataset are released on the project webpage, which they intend as a foundation for future machine learning work on single-photon lidar.
Future Directions
- Handling real sensor artifacts. The authors explicitly set aside practical challenges of high-resolution SPADs such as cross talk, hot pixels, blooming, and dead time, listing them as beyond the scope of this work.
- Extending beyond the simulated setting. Results are primarily simulated on their dataset, with real-world results described as proof-of-concept on 16 laser spots; scaling up real captures and comparing against PlatoNeRF with more captures remains in the supplement.
- Generalizable transient representations. The finding that time-of-flight features transfer to specular segmentation is presented as a first step; whether these features form a general-purpose representation for single-photon lidar is left open.
- Other sensor and illumination regimes. The authors do not consider low-resolution SPADs with diffuse illumination, and they assume objects are purely specular or diffuse — both restrictions that future work would need to relax.
Target Audience
Researchers and practitioners in computational imaging, computer vision, and computer graphics working on lidar, time-of-flight sensing, inverse light transport, neural rendering, or single-photon sensors. It will also be useful to engineers at companies building depth-sensing consumer devices (phones, tablets, headsets) or autonomous-vehicle perception systems, and to machine learning researchers interested in how large simulated datasets and learned priors can be applied to non-RGB sensing modalities. Readers need some background in lidar and 3D reconstruction; the paper's own framing of the light transport is accessible, but the method details assume familiarity with transient measurements and neural reconstruction models.
Authors’ abstract
3D scene reconstruction from a single measurement is challenging, especially in the presence of occluded regions and specular materials, such as mirrors. We address these challenges by leveraging single-photon lidars. These lidars estimate depth from light that is emitted into the scene and reflected directly back to the sensor. However, they can also measure light that bounces multiple times in the scene before reaching the sensor. This multi-bounce light contains additional information that can be used to recover dense depth, occluded geometry, and material properties. Prior work with single-photon lidar, however, has only demonstrated these use cases when a laser sequentially illuminates one scene point at a time. We instead focus on the more practical - and challenging - scenario of illuminating multiple scene points simultaneously. The complexity of light transport due to the combined effects of multiplexed illumination, two-bounce light, shadows, and specular reflections is challenging to invert analytically. Instead, we propose a data-driven method to invert light transport in single-photon lidar. To enable this approach, we create the first large-scale simulated dataset of ~100k lidar transients for indoor scenes. We use this dataset to learn a prior on complex light transport, enabling measured two-bounce light to be decomposed into the constituent contributions from each laser spot. Finally, we experimentally demonstrate how this decomposed light can be used to infer 3D geometry in scenes with occlusions and mirrors from a single measurement. Our code and dataset are released at https://shoot-bounce-3d.github.io.