Research
G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation
G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation Overview Research area: 3D computer vision — point cloud semantic and instance segmentation, 3D Gaussian Splatting (GS), b

- arXiv
- 2601.03510
- Published
- 2026-01-07
- Authors
- Hojun Song, Chae-yeong Song, Jeong-hun Hong, Chaewon Moon, Soo Ye Kim, Yiyi Liao, Jaehyup Lee, Sang-hyo Park
AI summary
G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D SegmentationOverview
- Research area: 3D computer vision — point cloud semantic and instance segmentation, 3D Gaussian Splatting (GS), boundary-aware perception.
- Technical level: Advanced. The paper assumes familiarity with point cloud backbones (MinkUNet, OctFormer, Point Transformer v3), 3D Gaussian Splatting parameterizations (covariance, opacity, scale), and feature distillation.
- Scope: The paper proposes G2P, a method that transfers attributes from optimized 3D Gaussians onto the original point cloud via covariance-aware correspondence, so that segmentation gains appearance/visibility cues while keeping point geometry and Gaussian-free inference.
Authors: Hojun Song, Chae-yeong Song, Jeong-hun Hong, Chaewon Moon, Soo Ye Kim (Adobe Research), Yiyi Liao (Zhejiang University), Jaehyup Lee, Sang-hyo Park — with affiliations at Kyungpook National University, Korea Electronics Technology Institute, Adobe Research, and Zhejiang University. Published 2026-01-07 (arXiv:2601.03510v3, cs.CV).
What This Paper Is About
Point clouds are sparse and irregularly sampled, so models tend to lean on coarse geometry and confuse objects that share a shape but differ in appearance — doors and windows coplanar with walls, or a reflective refrigerator flush against a wall. Existing fixes either stay geometry-only (boundary-aware methods) or borrow 2D image features that suffer from projection misalignment and occlusion loss. G2P's goal is to inject 3D-native appearance and visibility cues into point clouds by aligning 3D Gaussian attributes to point positions, without using pretrained 2D features or language supervision in the segmentation pipeline.
Key Contributions
- G2P alignment: a covariance-aware Gaussian-to-Point feature augmentation that transfers reliable 3D GS attributes (opacity and scale) to points while preserving the original point geometry, using Mahalanobis-distance matching over anisotropic Gaussian ellipsoids.
- Opacity distillation: a scheme that distills opacity-derived visibility cues from a GS-augmented appearance encoder into a point-only backbone, enabling Gaussian-free inference and avoiding cross-modal 2D–3D fusion.
- Scale-based boundary extraction: boundary pseudo-labels derived from anisotropic Gaussian scale magnitudes, combined with semantic boundary candidates, to improve boundary delineation.
- Evaluation: competitive results across ScanNet v2, ScanNet200, ScanNet++, and Matterport3D, with the strongest gains on geometrically challenging classes — and no external pre-training beyond the target dataset.
Main Findings
-
ScanNet v2 (20 classes, validation): G2P reaches 78.4 mIoU, exceeding PT v3 by +0.9 mIoU and BFANet by +0.4. Reference points: MinkUNet 72.2, OctFormer 75.7, SPG 76.0, BFANet 78.0, UniPre3D 77.6, ODIN (Swin-B) 77.8, PonderV2 77.0, VMVF 76.4, BPNet 69.7, PT v3 + PPT 78.6 (external pre-training). The paper notes G2P outperforms all geometric and 2D–3D fusion baselines that lack external pre-training and is competitive with PT v3 + PPT, which does use it.
-
ScanNet200 (200 classes, validation): G2P attains 36.6 mIoU and 83.8 OA, improving +1.4 mIoU over PT v3 (35.2 mIoU / 83.6 OA) and remaining competitive with UniPre3D (36.0 mIoU / 83.7 OA). Other baselines: MinkUNet 25.0/80.4, PointContrast 26.2, PT v2 30.2/82.7, OctFormer 32.6/83.0, SPG 31.5, PonderV2 32.3, PT v3 + PPT 36.0.
-
ScanNet200 hidden test set: G2P scores 35.7 mIoU, with Head 54.3, Common 28.7, and Tail 20.2 — the best Tail result among the compared methods (BFANet 36.0 mIoU / 55.3 / 29.3 / 19.3; CeCo 34.0/55.1/24.7/18.1; PonderV2 34.6/55.2/27.0/17.5; OctFormer 32.6/53.9/26.5/13.1; MinkUNet 25.3/46.3/15.4/10.6).
-
Instance segmentation (ScanNet200 validation, PointGroup framework): G2P reaches 41.8 mAP25, 33.3 mAP50, and 23.2 mAP, improving +1.7 mAP25 over PT v3 (40.1 / 33.2 / 23.1). PT v2 scores 39.6/31.9/21.4 and MinkUNet 32.2/24.5/15.8.
-
Geometrically challenging classes: G2P achieves an average IoU of 69.2 on the challenging subset versus 85.8 on the geometrically distinguishable subset. The largest class gains are refrigerator at 70.9 IoU (+6.0 versus PT v3's 64.9) and shower curtain at +7.2 IoU. For comparison, PT v3 averages 66.8 on challenging classes, BFANet† 67.3, and UniPre3D 68.6.
-
Additional benchmarks: On ScanNet++ and Matterport3D validation, G2P reaches 48.7 and 55.9 mIoU, versus PT v3 at 47.9 and 55.5, OctFormer at 44.5 and 55.0, and MinkUNet at 28.8 and 54.2.
-
Module ablation (ScanNet v2): Baseline PT v3† 77.0 mIoU; boundary guidance only 77.8; distillation only 77.8; both components together 78.4 — indicating complementary effects.
-
Boundary pseudo-label ablation: Semantic-only boundaries give 76.5 (r^s = 0.02), 76.9 (r^s = 0.04), 76.3 (r^s = 0.06). Scale-only boundaries give 77.3 (η = 0.3), 76.5 (η = 0.5), 76.9 (η = 0.7). The combined formulation peaks at 77.8 with r^s = 0.04 and η = 0.7. These runs fix the boundary loss at λ_b = 1.0 and exclude distillation.
-
Attribute learning ablation: Using raw 3D Gaussians (μ^g, c′, n′, α) yields only 73.1 mIoU, showing Gaussians alone are unreliable for point-level segmentation. Augmented points (μ^p, c, n, α′) reach 77.3. Fine-tuning the appearance encoder gives 74.9 (μ^p, c, α′) and 75.0 (μ^p, c, n, α′), while the paper's distillation approach reaches 77.8 (μ^p, c, α′) and removes the need for Gaussian features at inference.
-
Correspondence metric and neighborhood size: Euclidean matching gives 77.0 mIoU; Mahalanobis matching gives 77.4 at k = 10, 78.4 at k = 20, and 77.9 at k = 30. The paper attributes this to covariance-aware matching forming tighter matches near boundaries.
-
Efficiency: G2P uses 46.4M parameters (versus PT v3's 46.2M), training latency 220 ms versus 132 ms, training memory 7.3G versus 5.6G, inference latency 96 ms versus 79 ms, and inference memory 3.3G versus 1.9G, at 78.4 mIoU versus 77.0 for the reproduced PT v3†. Training overhead is reported as +67% latency and +30% memory; inference as +0.2M parameters (+0.4%), +21% latency, and +74% memory. MinkUNet is listed at 39.2M parameters, 71 ms training, 1.6G training memory, 29 ms inference, 1.4G inference memory, 72.3 mIoU; OctFormer at 44.0M, 259 ms, 4.2G, 94 ms, 4.4G, 74.3 mIoU.
-
Qualitative behavior: Baselines exhibit two failure modes — category confusion (e.g., door labeled as wall, window as door, refrigerator as cabinet) and incomplete segmentation of thin or coplanar structures. G2P produces label-consistent and more complete masks in these cases.
Methodology in Plain English
The pipeline has a preparation stage and a training stage.
Preparation — aligning Gaussians to points. The method keeps the original point cloud, where each point is a coordinate, color, and normal in R^9, as the geometric source of truth, because 3D Gaussians drift during optimization and blur structure. For each point, candidate Gaussians are first gathered within a Euclidean radius r^g = 0.06 m; the k = 20 closest are then selected by Mahalanobis distance, which accounts for each Gaussian's anisotropic shape through its covariance Σ = R S S^T R^T. Neighbors are weighted by the inverse of their Mahalanobis distance, normalized across the k neighbors, and used to aggregate each Gaussian's scale S and opacity α into per-point attributes. The augmented point becomes (μ^p, c, n, S′, α′) in R^13.
Preparation — boundary pseudo-labels. Points with small aggregated Gaussian scale magnitudes cluster near object boundaries, while large scales dominate smooth planar regions. After removing background classes (floor, wall), the method prunes the top 100η% of points by scale magnitude with η = 0.7, treating remaining low-scale points as scale-based boundary candidates. Because small scales can also arise from texture or photometric variation, these are unioned with semantic boundary candidates — points having a neighbor within radius r^s = 0.04 m with a different semantic label.
Preparation — appearance encoder. A Sonata architecture is repurposed as the teacher and trained from scratch on the augmented representation (μ^p, c, α′) in R^7, replacing geometric normals with view-consistent opacity cues. It is trained 400 epochs per dataset with batch size 1 and learning rate 0.002 on an NVIDIA A6000.
Training. A PT v3 backbone takes (μ^p, c, n), jointly predicts semantics and boundaries, and incorporates the boundary-semantic (B-S) block from BFANet (64-channel PT v3 decoder features producing 128-dimensional boundary-aware features with 8 attention heads; a two-layer MLP head with 64-dimensional hidden layer and 1-dimensional output). A mapping MLP projects backbone features and a cosine-similarity distillation loss pulls them toward the appearance encoder's features. Total loss is semantic (cross-entropy + Lovász-softmax) + λ_b × boundary (binary cross-entropy + Dice) + λ_d × distillation, with λ_b = 0.9 and λ_d = 0.4. Main training runs for 800 epochs on a single NVIDIA RTX 3090 with batch size 4 and AdamW at learning rate 0.003.
Gaussian reconstructions come from SceneSplat-7K, where each scene has approximately 1.5M Gaussian primitives. Results are reported on validation splits for all four datasets.
Why This Matters
Impact on research. The paper reframes Gaussian Splatting as a source of point-level attributes rather than an intermediate rendering representation. It offers an alternative to 2D–3D fusion, which struggles with point-to-pixel alignment errors in occluded regions, projection-induced information loss, and spatial misalignment between sparse 3D distributions and dense 2D grids. It also separates the two roles of GS attributes — opacity as a view-consistent confidence cue and scale as a boundary cue — and shows empirically that raw Gaussian coordinates are unreliable for point-level tasks (73.1 mIoU) while aggregated attributes help (77.3 mIoU).
Real-world applications.
- Indoor scene understanding and digital twins built from scanned rooms, where walls, doors, windows, and appliances must be separated despite shared planar geometry.
- Robotics and embodied agents that need to distinguish a refrigerator from a cabinet or a curtain from a wall using appearance rather than shape alone.
- AR/VR content pipelines that need labeled 3D geometry for editing, collision, or occlusions in reconstructed spaces.
- Large-scale 3D asset annotation where boundary-accurate masks reduce manual cleanup of thin or coplanar structures.
Industry relevance. The method needs no 2D foundation-model features or language supervision inside the segmentation pipeline, and distillation means inference uses only points — no Gaussian features at test time. The reported inference cost increase over PT v3 is 0.2M parameters (+0.4%), 21% latency, and 74% memory, which is the deployment-relevant number; the heavier +67% training latency and +30% training memory apply to the offline pipeline.
Future Directions
- Removing the offline dependency: the authors state as a limitation that G2P relies on an offline preparation stage and therefore depends on the availability of input Gaussians; making the pipeline less dependent on this stage is an open problem.
- Outdoor scenes: extending to outdoor environments "may require more robust GS construction." The SceneSplat-7K dataset provides indoor GS reconstructions, and outdoor environments are explicitly not available in it.
- Reducing training overhead: the +67% training latency and +30% training memory from appearance distillation are unaddressed by the current design.
- Correspondence quality at extreme boundaries: the ablation shows sensitivity to the neighborhood size k (78.4 at k = 20 versus 77.4 at k = 10 and 77.9 at k = 30) and to the semantic radius r^s and pruning ratio η, leaving room to make correspondence and boundary thresholds more adaptive.
Target Audience
- Researchers working on 3D point cloud segmentation, boundary-aware perception, and 3D Gaussian Splatting.
- Practitioners building indoor scene understanding, robotics, or AR/VR systems who need appearance-aware 3D labels without a 2D fusion stack.
- Readers interested in representation transfer between neural 3D representations (Gaussians) and discrete geometry (points).
- Prerequisite background: familiarity with point cloud backbones such as MinkUNet, OctFormer, and PT v3, plus the 3D GS parameterization (centroid, opacity, spherical harmonics, covariance). Beginners will find the concepts accessible at a high level but the implementation details demanding.
Authors’ abstract
Point cloud segmentation is critical for 3D scene understanding. However, sparse and irregular point distributions provide limited appearance evidence, making geometry-only features insufficient to distinguish objects with similar shapes but distinct appearances e.g., color, texture, and material. We propose Gaussian-to-Point (G2P), which transfers Gaussian attributes from 3D Gaussian Splatting to point clouds for more discriminative and appearance-consistent segmentation. Our G2P addresses the misalignment between optimized Gaussians and original point geometry by establishing point-wise correspondences. By distilling opacity-derived visibility cues, we mitigate the geometric ambiguity that limits existing models. Additionally, Gaussian scale attributes enable precise boundary localization in complex 3D scenes. Extensive experiments demonstrate that our approach achieves competitive performance on standard benchmarks and shows notable improvements on geometrically challenging classes, without pretrained 2D features or language supervision in our segmentation pipeline.