Skip to content
AI.info

Research

Semantic Enrichment of CAD-Based Industrial Environments via Scene Graphs for Simulation and Reasoning

Overview Research area: Robotics, specifically 3D scene graph generation and semantic understanding of industrial environments for simulation and high-level reasoning. Technical level: Intermediate. T

Semantic Enrichment of CAD-Based Industrial Environments via Scene Graphs for Simulation and Reasoning
arXiv
2601.06415
Published
2026-01-10
Authors
Nathan Pascal Walus, Ranulfo Bezerra, Shotaro Kojima, Tsige Tadesse Alemayoh, Satoshi Tadokoro, Kazunori Ohno

AI summary

Overview

  • Research area: Robotics, specifically 3D scene graph generation and semantic understanding of industrial environments for simulation and high-level reasoning.
  • Technical level: Intermediate. The paper combines well-known tools (DBSCAN clustering, LVLM prompting, USD-based CAD loading) in a multi-stage pipeline, so no single component is exotic, but readers need familiarity with scene graphs, CAD meshes, and clustering.
  • Scope: The paper proposes an offline pipeline that converts an unstructured CAD model of a complex industrial room into a multi-layered 3D scene graph with semantic labels and inferred functional relations among pipes, valves, and gauges.

What This Paper Is About

Industrial CAD models give robots and simulators an exact geometric and visual description of a facility, but they carry no semantic labels, no relational structure, and no functional information, which limits high-level scene understanding and robot training. Existing 3D scene graph work has focused on indoor and domestic settings rather than factories or power plants, and has not addressed messy structures such as long, intertwined pipe networks. The goal of this paper is to build a detailed 3D scene graph directly from a CAD environment and to use it to automatically infer which functional units (valves, gauges) relate to which others inside pipe systems.

Key Contributions

  1. A novel pipeline for scene graph generation from CAD models that uses an LVLM (gpt-4o) for semantic enrichment and DBSCAN for spatial clustering to handle the complexity of industrial environments.
  2. A multi-layered scene graph structure capturing both low-level geometric relationships and high-level functional dependencies between components such as pipes, gauges and valves, even in cluttered settings.
  3. A method (Algorithm 1) for automatically inferring functional relations within pipe systems from the generated scene graph, enabling advanced scene reasoning and accessibility for LLM agents.
  4. An experimental evaluation on a real robot test field CAD environment, including quantitative semantic accuracy over three pipe structures and a comparison of gpt-4o-generated semantics against manually created ground truth semantics.

Main Findings

  • Environment scale: The test environment comprises 8,327 meshes. After excluding 9 meshes identified as modeling artifacts, mesh grouping (merging meshes ≤ 1 cm³ with larger neighbors) reduced the count to 2,068 meshes, an approximately 75% reduction.
  • Clustering results: DBSCAN with ϵ = 0.01 and min_samples = 1 produced 39 clusters. Excluding the outer fences left 15 clustered structures. The algorithm isolated freestanding structures and grouped complex multi-pipe systems correctly, but two main errors occurred: one pipe structure was incorrectly split into two clusters, and the central access platform was clustered together with a pair of tanks despite no visible physical connection.
  • Semantic labeling volume: Each of the 2,068 meshes was assigned two semantic labels, using a total of 7 'group' labels and 44 distinct 'name' labels. Structural elements made up the majority of the distribution; pipe-related meshes made up around one quarter.
  • Semantic accuracy (Table I): For structure S₁ (174 meshes): 38.5% mesh label accuracy, 79.9% group label accuracy. For S₂ (82 meshes): 30.5% mesh label accuracy, 73.2% group label accuracy. For S₃ (200 meshes): 42.0% mesh label accuracy, 84.0% group label accuracy.
  • Functional unit detection (Table I): S₁ had 5 fully found valves, 7 partially found valves, 0 missed valves, 12 fully found gauges, 0 missed gauges. S₂ had 2 fully found valves, 4 partially found, 0 missed, 4 fully found gauges, 2 missed gauges. S₃ had 0 fully found valves, 3 partially found, 1 missed valve, 3 fully found gauges, 2 missed gauges.
  • Robustness to partial errors: Even with partially incorrect semantics, the position and sequence of functional relations was correctly identified in structure S₁, demonstrating robustness to minor labeling errors.
  • Systematic mislabeling: The ratio of gaskets to flanges came out logically inconsistent — a 1:2 ratio would be expected (one gasket between two flanges), but the results showed a strong over-prediction of gaskets, suggesting gpt-4o misclassified many flanges as gaskets despite being given bounding box dimensions.
  • Functional analysis limitations: Algorithm 1 cannot resolve the order of functional units connected solely to one common mesh (such as two gauges on one pipe segment in S₁). The scene graph also captures only connectivity, not directionality — for the tank in S₃, it remains ambiguous which pipes are inputs and which are outputs. The environment did not feature complex branching pipes, leaving applicability there open.
  • Qualitative comparison: In S₁, ground truth semantics led to all functional units being identified, while gpt-4o-based semantics caused wheel valves to be incorrectly identified as two separate functional units; both semantics showed uncertainty regarding the order of the gauges. In S₃, one valve was completely mislabeled as 'connection assembly' and therefore not considered a functional unit.

Methodology in Plain English

The input is a CAD model in Universal Scene Descriptor (usd) format representing a single room (excluding walls and ceiling) from a robot test field location in Fukushima, Japan. NVIDIA's Isaac Lab is used to access the environment, and the whole pipeline is programmed in Python using the Omniverse Python API.

First, a vocabulary of possible labels is built as a three-layer tree: gpt-4o is shown multiple rendered images of the environment from different angles and proposes labels, with a general 'group' label and a specific 'name' label available for each mesh.

Then, each mesh is prepared: vertices are voxelized onto a 1 cm grid, duplicate points are deleted, and mesh faces are filled with points spaced 1 cm apart for distance calculations. Small meshes (below 1 cm³) are merged into the nearest larger mesh to keep the count manageable.

Each mesh becomes a node in the scene graph with geometric attributes (centroid and 3D bounding box). For semantic labeling, gpt-4o receives three pairs of 512x512 rendered RGB images per mesh — each pair shows one view with all meshes visible and one with only the target mesh visible — plus the mesh's bounding box dimensions and the pre-generated vocabulary. The model assigns a 'group' and a 'name' label, or proposes a new one.

DBSCAN (ϵ = 0.01, min_samples = 1) then clusters nodes by minimal mesh point proximity; the ground plane is excluded. Edges are added between nodes no more than 1 cm apart within the same cluster, and a parent node is created for each cluster.

Finally, for pipe-system clusters, Algorithm 1 infers functional relations. It uses a predefined list of "connector" labels (here 'Pipe assembly') to traverse the graph while ignoring non-relevant meshes such as structural supports, and it requires a list of functional units obtained by finding interconnected node clusters sharing the same 'group' label. The algorithm iteratively expands each functional unit into neighboring pipe nodes, then builds a functional graph whose nodes are the functional units and whose edges connect units that share a scene graph edge.

Why This Matters

This work addresses a gap the authors identify: no published scene graph generation approaches dealt with complex environments such as factories or power plants, and the lack of structure in such environments makes existing methods inapplicable because objects are neither named nor meaningfully sorted within the scene tree. Without structured, functional information, simulation of dynamic elements remains out of reach for current approaches.

Real-world applications:

  • Robot training and validation of perception, decision-making, task execution, and robot-environment interaction inside realistic industrial simulations.
  • Reasoning over structured scene data by LLM agents for task planning and scene analysis.
  • Generating synthetic simulation data for dynamic elements by giving an LLM the scene graph including functional relations.
  • Safety-oriented deployment of robots to take over complex tasks in industrial environments where human workers are continuously exposed to hazardous situations.

Industry relevance: The paper targets industrial facilities directly and demonstrates the pipeline on a real robot test field in Fukushima, Japan. The 75% mesh reduction shows a practical route to lowering computational load and LVLM prompting costs, which matters for scaling to larger or more detailed facilities. The authors note that CAD source data for technical components is already used for semantic segmentation in other work, and point to component catalogs as a possible improvement path.

Future Directions

  • Improving semantic labeling accuracy for specialized industrial components, for example by matching against a CAD model catalog or by leveraging component repetition within the environment itself to enforce labeling consistency.
  • Making preprocessing more geometry-aware to avoid the voxelization artifacts that caused a pipe to be split into two clusters and two unrelated structures to be joined.
  • Extending Algorithm 1 to resolve ordering when functional units connect only to one common mesh, and to capture directionality so inputs and outputs can be distinguished.
  • Testing the approach on environments with complex branching pipes and on larger or more detailed environments, where DBSCAN's need for a full distance matrix stresses scalability.

Target Audience

Robotics researchers working on scene understanding, semantic mapping, and simulation environments; engineers building industrial digital twins or robot test fields from CAD data; and researchers applying large vision-language models and LLM agents to structured spatial representations for task planning and dynamic simulation.

Authors’ abstract

Utilizing functional elements in an industrial environment, such as displays and interactive valves, provide effective possibilities for robot training. When preparing simulations for robots or applications that involve high-level scene understanding, the simulation environment must be equally detailed. Although CAD files for such environments deliver an exact description of the geometry and visuals, they usually lack semantic, relational and functional information, thus limiting the simulation and training possibilities. A 3D scene graph can organize semantic, spatial and functional information by enriching the environment through a Large Vision-Language Model (LVLM). In this paper we present an offline approach to creating detailed 3D scene graphs from CAD environments. This will serve as a foundation to include the relations of functional and actionable elements, which then can be used for dynamic simulation and reasoning. Key results of this research include both quantitative results of the generated semantic labels as well as qualitative results of the scene graph, especially in hindsight of pipe structures and identified functional relations. All code, results and the environment will be made available at https://cad-scenegraph.github.io

Read the original paper