Research
SemanticBridge - A Dataset for 3D Semantic Segmentation of Bridges and Domain Gap Analysis
Overview Research area: 3D computer vision, specifically point cloud semantic segmentation applied to civil infrastructure (bridges), plus a study of sensor-induced domain gap. Technical level: Advanc

- arXiv
- 2512.15369
- Published
- 2025-12-17
- Authors
- Maximilian Kellner, Mariana Ferrandon Cervantes, Yuandong Pan, Ruodan Lu, Ioannis Brilakis, Alexander Reiterer
AI summary
Overview
Research area: 3D computer vision, specifically point cloud semantic segmentation applied to civil infrastructure (bridges), plus a study of sensor-induced domain gap.
Technical level: Advanced. The paper assumes familiarity with point cloud representations, sparse 3D convolutions, kernel-point convolutions, transformer-based point networks, and standard segmentation metrics such as mIoU.
Scope: The paper releases SemanticBridge, an annotated TLS/MLS point cloud dataset of 20 bridges, benchmarks three 3D segmentation architectures on it, and measures how much performance drops when the test data comes from a different scanner than the training data.
What This Paper Is About
Supervised 3D semantic segmentation needs annotated data, and no public dataset targets bridge scenes with bridge-specific components such as abutments and pillars. This paper introduces SemanticBridge, containing 20 bridges scanned in the UK and Germany with stationary (TLS) and mobile (MLS) laser scanners, annotated over nine classes. It then uses that data to quantify the "domain gap" — the accuracy loss that occurs when a model trained on one scanner is evaluated on data from a different scanner.
Key Contributions
- A new bridge segmentation dataset. SemanticBridge contains 20 bridges captured with three different scanners (Faro Focus 3D X330 and Leica RTC360 as TLS, Leica BLK2GO as MLS), totaling 245M points with nine semantic classes including abutment and pillar classes that the authors state previous datasets do not cover. The authors describe it as the largest annotated laser-scanned point cloud dataset of bridges available, and in the conclusion as the first annotated point cloud dataset of bridges.
- A controlled domain gap study. Seven bridges were surveyed with both a stationary and a mobile scanner, so the same bridge can be tested with each sensor while the model and target bridge stay fixed, isolating the sensor as the only variable.
- A comparative architecture benchmark. Three state-of-the-art 3D architectures — UNet3D with sparse convolutions and residual blocks, KPConv, and PointTransformerv2 (PTv2) — were trained and evaluated on the dataset.
- Public release of code and data at https://github.com/mvg-inatech/3d_bridge_segmentation.
Main Findings
- Baseline performance on stationary-scanner data is similar across models. On the TLS test split, overall mIoU is 0.707 for UNet3D, 0.705 for KPConv, and 0.635 for PTv2.
- Switching sensors degrades every model. Evaluated on the same bridges captured by the mobile scanner, mIoU is 0.574 for UNet3D, 0.595 for KPConv, and 0.575 for PTv2.
- The measured gap ranges from 6.9 to 11.4 percentage points of mIoU. KPConv drops the most at 11.4 points, UNet3D by 10.8 points, and PTv2 the least at 6.9 points.
- The hardest classes are Traffic Sign, Pillar, and Abutment. Traffic Sign is the weakest for all models; on the TLS baseline, UNet3D reaches 0.186 IoU on Traffic Sign versus 0.064 for KPConv, nearly tripling it.
- UNet3D is strongest on Pillar on the TLS baseline, leading KPConv by 3.9 percentage points and PTv2 by 19.4 percentage points.
- Domain gap impact is class-dependent. The largest single drop is KPConv on Pillar, at 26.60 percentage points. The smallest change for a bridge-structure class is PTv2 on Pillar, at 0.50 percentage points. PTv2 on Traffic Sign actually improves by 7.80 percentage points when using MLS data.
- Ground and vegetation are barely affected. The paper reports that the "Underground" and "High Vegetation" classes "barely show changes in the prediction metric" across sensors.
- PTv2 is the weakest overall despite its reputation. It achieves the best result only on the vegetation class, and by just 0.7 percentage points over the next best model, even though its reported performance on the S3DIS indoor dataset is superior to the other compared models.
- On average across the three models, degradation is concentrated on the "Top Surface", then "Superstructure", then "Abutment" classes, in that order.
- Dataset scale. The released splits contain 160,943,804 TLS training points, 84,153,822 TLS test points, 43,066,128 MLS training points, and 29,485,178 MLS test points. The average number of subclouds used in a single epoch was 3575.
Methodology in Plain English
Data capture. Three scanners were used. The Faro Focus 3D X330 and Leica RTC360 are stationary tripod scanners (TLS) with 360/300 degree horizontal/vertical field of view, quoted accuracies of ±2 mm and 1.9 mm, and scan ranges of 0.6–330 m and 0.5–130 m respectively. The Leica BLK2GO is a handheld mobile scanner (MLS) with 360/270 degree FOV, a 0.5–25 m range, and a quoted accuracy of ±3 mm. Ten bridges were recorded with the Faro TLS, ten with the Leica TLS, and seven with the MLS. For each bridge the authors report height, length, width, number of spans, span lengths, number of pillars, and the crossing type (road, river, or railing).
Annotation. Stationary-scanner clouds were annotated manually in CloudCompare. Because manual annotation is expensive, the mobile-scanner clouds were labeled by interpolation: the two clouds for a bridge were aligned, the Euclidean distance from each unlabeled MLS point to the labeled TLS points was computed, and the label was assigned by majority vote among the k = 8 nearest neighbors. Points whose minimum distance exceeded a threshold of θ = 0.5 m were dropped to avoid mislabeling. The interpolated labels were then manually verified.
Classes. Nine classes: Unlabeled, Ground, High Vegetation, Abutment, Superstructure, Deck, Railing, Traffic Sign, and Pillar. Abutment, Superstructure, Deck, Railing, and Pillar cover the bridge itself; Ground, High Vegetation, and Traffic Sign cover the surroundings. The class scheme was aligned with SynBridge so the datasets can be combined.
Preprocessing and splitting. All clouds were downsampled with a 1 cm³ voxel grid to remove redundancy. The training split uses only stationary-scanner data from 15 bridges (eight Faro clouds, seven Leica clouds); the test split holds out five bridges (two Faro, three Leica). For the domain gap experiment, three of the seven MLS-scanned bridges (test indices 13, 17, 19) were evaluated, chosen because the other four MLS bridges correspond to bridges already seen in training.
Training setup. Because full bridge scenes are too large to segment in one pass, each scene is split into subclouds by sampling a random center point and cropping a bounding box with b_x = b_y = b_z = 3, repeated until every point appears in at least one subcloud. Input to the networks is the point coordinates x, y, z plus color features r, g, b, with an extra constant channel so black points are still represented. The starting voxel size is 8 cm. UNet3D doubles voxel size per layer with feature dimensions 32→64→128→256; KPConv starts at 64 features with 15 kernel points and a convolution radius growing from 0.2; PTv2 uses voxel multipliers [×3.0, ×2.5, ×2.5, ×2.5], embeddings of 48 channels expanding to 48→96→192→384→512, k = 16 neighbors, and encoder/decoder block depths of [2, 2, 6, 2] and [1, 1, 1, 1]. Implementations use PyTorch and NumPy, SGD with momentum 0.9 and weight decay 0.0001, initial learning rate 0.001 with exponential decay γ = 0.98, cross-entropy loss with label smoothing ε = 0.1, on a single NVIDIA A100 GPU with batch size six. Augmentations include dropping points, random shifts, z-axis rotation, scaling, axis flips, and Gaussian noise on both coordinates and colors.
Evaluation. Performance is reported as mean Intersection over Union (mIoU), per class and overall. Three evaluations are run: TLS-only test data, the paired TLS clouds of bridges 13/17/19, and the paired MLS clouds of the same bridges, with the differences between the last two reported as the domain gap.
Why This Matters
This is one of the few works that both releases a bridge-specific 3D dataset and quantifies, rather than only names, the sensor-induced domain gap. The finding that identical models on identical bridges lose up to 11.4 points of mIoU purely from a scanner change is a concrete warning for anyone planning to deploy a trained segmentation model on data from new hardware.
Real-world applications:
- Automated condition assessment and digital twin creation for bridge assets, using an existing bridge inventory model as a starting point.
- Tracking damage and changes over repeated inspection cycles, which requires consistent component-level labels.
- Feeding bridge management systems with structured 3D component data for maintenance planning and budget allocation.
- Benchmarking and developing domain adaptation methods that let a model trained on stationary scans transfer to cheaper mobile scanning campaigns.
Industry relevance. Because mobile laser scanning dramatically reduces acquisition time compared to tripod-based scanning, but produces sparser data, the domain gap is exactly the trade-off infrastructure operators face. Quantifying it lets agencies decide when mobile scanning is worth the accuracy cost, and gives researchers a benchmark for closing that gap. The work was funded by the German BMDV as part of the mFUND project "Partially automated creation of object-based inventory models using multi-data fusion of multimodal data streams and existing inventory data - mdfBIM+" (FKZ: 19FS2021B).
Future Directions
- Exploiting the four unused MLS bridges. The four remaining mobile-scanned bridges were deliberately excluded from the benchmark because their TLS counterparts were already in the training set, and the authors suggest they could be used to study whether and how they can reduce the domain gap.
- Extending to more bridge types and regions. Every bridge in the dataset is of the same type, so the authors warn that applying the trained networks to an arch bridge could perform poorly due to learned bias; likewise, all data comes from the UK and Germany, and designs in other regions may follow different regulations and construction practices.
- Broadening the sensor study. Only two sensor types are represented, and the authors argue a wider variety of sensors and more diverse environments are needed for a comprehensive understanding of the domain gap.
- Tuning the architectures for this task. The networks were deliberately not optimized for bridge segmentation; the authors expect gains from adjusting hyperparameters such as the initial voxel size or subcloud size, and they note the Traffic Sign class remains weak while being frequently covered in other datasets.
Target Audience
Researchers in 3D deep learning and point cloud semantic segmentation who need an infrastructure-focused benchmark; civil and structural engineering researchers working on bridge inspection, digital twins, and structural health monitoring; practitioners evaluating whether mobile laser scanning is accurate enough for component-level segmentation; and domain adaptation researchers looking for a controlled real-world sensor-shift benchmark.
Authors’ abstract
We propose a novel dataset that has been specifically designed for 3D semantic segmentation of bridges and the domain gap analysis caused by varying sensors. This addresses a critical need in the field of infrastructure inspection and maintenance, which is essential for modern society. The dataset comprises high-resolution 3D scans of a diverse range of bridge structures from various countries, with detailed semantic labels provided for each. Our initial objective is to facilitate accurate and automated segmentation of bridge components, thereby advancing the structural health monitoring practice. To evaluate the effectiveness of existing 3D deep learning models on this novel dataset, we conduct a comprehensive analysis of three distinct state-of-the-art architectures. Furthermore, we present data acquired through diverse sensors to quantify the domain gap resulting from sensor variations. Our findings indicate that all architectures demonstrate robust performance on the specified task. However, the domain gap can potentially lead to a decline in the performance of up to 11.4% mIoU.