Research
TrueCity: Real and Simulated Urban Data for Cross-Domain 3D Scene Understanding
Overview Research area: 3D computer vision, specifically semantic segmentation of urban point clouds and synthetic-to-real (sim-to-real) domain shift. Technical level: Intermediate. The paper is reada

- arXiv
- 2511.07007
- Published
- 2025-11-10
- Authors
- Duc Nguyen, Yan-Ling Lai, Qilin Zhang, Prabin Gyawali, Benedikt Schwab, Olaf Wysocki, Thomas H. Kolbe
AI summary
Overview
Research area: 3D computer vision, specifically semantic segmentation of urban point clouds and synthetic-to-real (sim-to-real) domain shift.
Technical level: Intermediate. The paper is readable for anyone familiar with basic machine learning and point cloud concepts, but it assumes some familiarity with LiDAR acquisition, segmentation metrics (IoU, mIoU, OA), and 3D city modeling standards.
Scope: The paper introduces TrueCity, a benchmark dataset that pairs cm-accurate annotated real-world mobile laser scanning point clouds with a semantic 3D city model and simulated point clouds of the same four streets in Ingolstadt, Germany, and uses it to quantify the synthetic-to-real domain gap for point cloud semantic segmentation.
What This Paper Is About
Semantic segmentation of 3D point clouds in urban scenes is limited by a shortage of high-quality, annotated real-world data, which pushes researchers toward simulated data that is scalable and perfectly labeled. The core problem is that simulated scenes are usually designer-crafted fictions that do not match any real location, so the synthetic-to-real domain gap cannot be measured fairly. TrueCity addresses this by providing real and simulated point clouds that represent the same physical city, labeled with a shared class taxonomy, so that domain shift can be studied per class and per model.
Key Contributions
- A new list of 12 urban semantic segmentation classes harmonized with the international standards CityGML 2.0 and OpenDRIVE 1.4, replacing the ad hoc taxonomies common in prior point cloud datasets.
- The TrueCity benchmark dataset itself: cm-accurate annotated real-world point clouds, a derived semantic 3D city model structured according to CityGML 2.0, and annotated simulated point clouds representing the same city.
- A synchronization setup that enables consistent evaluation of the sim-to-real gap, including spatially separated (not interleaved) synthetic and real training segments and a fixed real-only validation and test split.
- An extensive empirical study across eight baselines from point-based, kernel-based, and transformer-based families, quantifying domain shift and identifying which classes are sensitive or insensitive to the synthetic-real mix.
Main Findings
- No prior benchmark offers synchronized real and simulated point clouds at the same site. The authors state that DELIVER and Paris-CARLA-3D include simulation data, but their synthetic scenes do not represent the corresponding real sites; in Paris-CARLA-3D, Paris scans are paired with CARLA-generated fictitious towns, and the underlying real 3D city assets are unavailable.
- Synthetic data helps, but how much depends on the model's inductive bias. Point Transformer v3 improves mIoU from 14.13 at 100S–0R to 25.30 at 50S–50R and 25.24 at 0S–100R, while Point Transformer v1 moves from 16.30 at 100S–0R to 28.89 at 0S–100R, indicating a stronger reliance on real sensor statistics.
- Locality-driven models lean heavily on real data. RandLA-Net rises from 8.98 at 100S–0R to 17.71 at 0S–100R mIoU, and Superpoint Transformer from 14.31 to 19.61 as the real fraction increases.
- A modest synthetic prior can help hierarchical set abstraction. PointNet++ peaks at 25.38 mIoU with 25S–75R versus 23.39 with 0S–100R.
- Real data can be partially replaced by simulation without loss. Point Transformer v3 achieves 65.94% OA with 50% synthetic data, an 8.5% relative improvement over the all-real setting (60.75%).
- Some classes show minimal domain gap. WallSurface is captured well synthetically: Point Transformer v1 reaches 53.15% IoU without real data and improves only to 67.10% with full real supervision (a 26.2% relative gain), and KPConv attains 73.55% with 75% real data versus 59.97% synthetic-only (22.6% improvement). RoadSurface reaches around 80% for PointNet++ and Point Transformer v1 with 50% real data (relative gains of 33.3% and 28.7%), while KPConv rises from 30.64% to 69.26% with just 25% real data (126.1% improvement).
- A small amount of real data can close much of the gap for vegetation. SolitaryVegetationObject improves from 15.54% to 35.35% for Point Transformer v1 with 25% real data, from 45.74% to 70.51% for KPConv, and from 0% to 50.60% for PointNet++.
- Door, RoofSurface, and BuildingInstallation are the weakest classes across all methods, with IoU rarely exceeding 8%. Gains from real data are unstable; Point Transformer v1 improves Door from 0.01% to 0.66% but the gains often collapse when real data dominates.
- Noise benefits most consistently from real data. PointNet++ climbs from 0% to 31.70%, KPConv from 0.31% to 34.73%, and Point Transformer v1 from 0.16% to 35.04%.
- CityFurniture shows a U-shaped trend. KPConv declines from 18.25% to 1.85% at intermediate real-data proportions but recovers to 14.57% with fully real supervision, suggesting mid-range ratios create competing cues.
- Window remains hard because simulation cannot model glass. Without real data, Window IoU is only 3.28% for Point Transformer v1 and 3.61% for PointNet++, rising to 31.29% and 14.60% respectively with full real data.
- Class definitions vary widely across existing urban datasets. The authors report that the minimum number of classes is eight and the maximum is 50 among the datasets they analyze.
Methodology in Plain English
The researchers chose a real urban environment: the inner city of Ingolstadt, a mid-sized German town, covering four streets totaling about 500 m with adjacent facades of 2- to 4-story buildings, vegetation, street furniture, vehicles, and pedestrians.
For the real data, a mobile mapping system (MoSES, mounted on a minivan) recorded 113 million points at densities up to 3,000 pts/m² with a relative accuracy of 1–3 cm. Positioning combined an inertial measurement unit, odometer sensors, and differential GPS with RTK correction from the German SAPOS service. Points were georeferenced in UTM Zone 32N (EPSG:25832). Labeling was a three-step manual and semi-automatic process: divide the cloud into spatially connected components, separate ground surfaces (roads, sidewalks) using cloth simulation filtering, then manually assign class IDs.
For the synthetic data, the team manually modeled buildings and facades from the real point clouds following the CityGML standard, curated road networks and roadside objects as an OpenDRIVE dataset, converted that to CityGML 2.0, replaced coarse tree geometry with detailed 3D assets, and translated everything into a local coordinate reference system. They then re-simulated the real scanning route in the CARLA driving simulator with a 128-laser sensor model (500,000 points/s, 20 Hz rotation, 15° upper and −25° lower field of view, 360° horizontal field of view, 100 m range) plus a Gaussian range error with ρ = 0.02 m to make the scans more realistic.
The simulation covers only static objects, so Vehicle and Pedestrian are absent from synthetic clouds. Experiments mixed synthetic and real training data at 100S–0R, 75S–25R, 50S–50R, 25S–75R, and 0S–100R by assigning contiguous spatial segments rather than interleaving points, keeping street length and coverage fixed while raw point totals varied modestly (112.9–137.6M). Test and validation data were always real-only. Eight baselines were trained under a unified setup of 100 epochs, constant learning rate 10⁻⁴, mini-batch size 32, AdamW, and a 2,048-point budget per item, on NVIDIA H40, L40, and RTX 6000 Ada GPUs. Superpoint Transformer deviated, using SGD with learning rate 0.01, weight decay 10⁻⁴, and batch size 4. The authors report that they deliberately omitted data augmentation to avoid adding a confounding variable to the domain gap analysis.
Why This Matters
Impact on research: TrueCity is described as the first urban semantic segmentation benchmark with cm-accurate annotated real-world point clouds, semantic 3D city models, and annotated simulated point clouds representing the same city. Because the real and synthetic data are synchronized and share a standardized class taxonomy, researchers can now measure per-class domain shift instead of relying on fictitious, de-synchronized scenes, which the authors argue removes a major uncertainty factor in sim-to-real studies.
Real-world applications:
- Municipal and national 3D city model maintenance, given the paper's note that approximately 216.5 million building models are available worldwide as open CityGML datasets as of 2024.
- Automated extraction of road space, facade, and street furniture information for high-definition mapping and driving simulation workflows that rely on OpenDRIVE.
- Urban planning and infrastructure monitoring that benefit from semantically labeled street-level point clouds.
- Training data generation pipelines for survey and mapping companies, where simulation can substitute for part of the costly real annotation effort.
Industry relevance: The standardized class list links segmentation output directly to CityGML 2.0 and OpenDRIVE 1.4, formats already used by public authorities and geospatial industries. The finding that real data can be partially replaced by simulation (for example, Point Transformer v3 reaching 65.94% OA with 50% synthetic data) has direct cost implications for organizations that pay for mobile laser scanning and manual annotation.
Future Directions
- Incorporating radiometry. The current experiments use geometry only. The authors note that radiometric values could be added through image projection, since images and trajectories from the mobile mapping acquisition are provided, and they call for future work on how object material affects simulation and the resulting domain gap.
- Scaling the dataset. The authors acknowledge that the high acquisition cost and accuracy mean TrueCity cannot reach the scalability of image-based datasets such as ScanNet, leaving open how to broaden coverage while retaining accuracy.
- Dynamic objects. Because the real capture is from one timestamp, only static objects are included; the authors suggest that dynamic objects could be simulated in various traffic scenarios using the three provided data subsets.
- Material-complex classes. Simulation fails on objects like glass, as shown by the low Window IoU without real data, raising the question of how simulators must improve to model ray-penetrable and reflective materials.
Target Audience
This paper is most useful for researchers and engineers working on 3D semantic segmentation, LiDAR simulation, and sim-to-real domain adaptation. It is also relevant to geomatics and 3D city modeling practitioners who work with CityGML and OpenDRIVE standards, and to teams in mapping, surveying, and autonomous driving who need to decide how much real annotated data is actually necessary when synthetic data is available. Readers seeking a benchmark for evaluating cross-domain generalization in urban point cloud segmentation will find the dataset and evaluation protocol directly applicable.
Authors’ abstract
3D semantic scene understanding remains a long-standing challenge in the 3D computer vision community. One of the key issues pertains to limited real-world annotated data to facilitate generalizable models. The common practice to tackle this issue is to simulate new data. Although synthetic datasets offer scalability and perfect labels, their designer-crafted scenes fail to capture real-world complexity and sensor noise, resulting in a synthetic-to-real domain gap. Moreover, no benchmark provides synchronized real and simulated point clouds for segmentation-oriented domain shift analysis. We introduce TrueCity, the first urban semantic segmentation benchmark with cm-accurate annotated real-world point clouds, semantic 3D city models, and annotated simulated point clouds representing the same city. TrueCity proposes segmentation classes aligned with international 3D city modeling standards, enabling consistent evaluation of synthetic-to-real gap. Our extensive experiments on common baselines quantify domain shift and highlight strategies for exploiting synthetic data to enhance real-world 3D scene understanding. We are convinced that the TrueCity dataset will foster further development of sim-to-real gap quantification and enable generalizable data-driven models. The data, code, and 3D models are available online: https://tum-gis.github.io/TrueCity/