Skip to content
AI.info

Research

UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning

UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning Overview Research area: Computer vision, specifically 3D scene understanding — lifting 2D instance/semantic segmentation labels int

UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning
arXiv
2512.24763
Published
2025-12-31
Authors
Ankit Dhiman, Srinath R, Jaswanth Reddy, Lokesh R Boregowda, Venkatesh Babu Radhakrishnan

AI summary

UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning

Overview

  • Research area: Computer vision, specifically 3D scene understanding — lifting 2D instance/semantic segmentation labels into 3D representations (3D Gaussian Splatting) for consistent 3D instance segmentation.
  • Technical level: Advanced. The paper assumes familiarity with 3D Gaussian Splatting, neural radiance fields, contrastive learning, and triplet loss.
  • Scope: A single-stage framework that embeds a learnable per-Gaussian vector, optimizes it with contrastive, triplet (boundary-mined) and 3D neighborhood losses, and decodes labels directly from the embedding without post-processing clustering — evaluated on ScanNet, Replica3D and Messy-Rooms.

What This Paper Is About

2D segmentation models label objects accurately in individual images but give inconsistent labels across different views of the same scene. The goal is to "lift" those inconsistent 2D masks into a 3D Gaussian Splatting representation so that every object gets one consistent 3D instance label. UniC-Lift does this in a single unified training stage, instead of the two-stage pipelines used by prior work that either preprocess labels for consistency or run a separate clustering algorithm after optimization.

Key Contributions

  1. A single-stage 3D segmentation method that directly decodes a learned 3D embedding into a consistent segmentation label from inconsistent 2D segmentation labels, eliminating separate clustering or label-association steps.
  2. A contrastive triplet loss that reduces inter-class variance and aids accurate 3D segmentation, applied on boundary-mined samples after passing the rasterized embeddings through a linear layer to stabilize training.
  3. An "Embedding-to-Label" procedure: sigmoid-restrict the rasterized embedding to [0,1], threshold at 0.5 into a binary vector, and decode to an integer label as l = Σ_k Ṽ_k · 2^(k−1).
  4. Demonstration of high-quality downstream applications — object manipulation and object extraction — enabled by the resulting segmentation.

Main Findings

  • ScanNet and Replica3D results (PQ_scene): On ScanNet, UniC-Lift reaches 63.0 versus Contrastive-Lift 62.3, Gaussian-Grouping 61.83, Panoptic-Lifting 58.9, PNF+GT Boxes 54.3, PNF 48.3, and DM-NeRF 41.7. On Replica3D, UniC-Lift reaches 88.7 versus Gaussian-Grouping 66.52, Contrastive-Lift 59.1, Panoptic-Lifting 57.9, PNF+GT Boxes 52.5, DM-NeRF 44.1, and PNF 41.1 — the paper describes this as a 1.3× higher PQ_scene (88.7 vs. 66.52) than Gaussian Grouping and a gain of nearly 10 points on Replica3D.
  • Messy-Rooms scalability (PQ_scene, mean over 8 scenes): UniC-Lift 71.5, Contrastive-Lift 69.0, Unified-Lift 69.0, OmniSeg3D-GS 66.0, Panoptic-Lifting 63.2. The paper states UniC-Lift outperforms the baselines in 6 out of 8 scenes. Per-scene, UniC-Lift leads in Old Room with 25 objects (86.0 vs. OmniSeg3D-GS 80.1), 50 objects (79.1 vs. Contrastive-Lift 75.8), 100 objects (70.8 vs. Contrastive-Lift 69.1), 500 objects (57.4 vs. Contrastive-Lift 55.0), and in Large Corridor at 100 objects (73.0 vs. Unified-Lift 70.7) and 500 objects (56.7 vs. Unified-Lift 54.1).
  • Training time: On a scene from Replica on an NVIDIA A6000, UniC-Lift trains in under 40 mins, versus over 15 hrs for Contrastive-Lift and over 20 hrs for Panoptic Lifting.
  • Head-to-head against a 3DGS-based Contrastive-Lift (Replica "room_1", single NVIDIA 3090): CL+3DGS takes 42.2 m for 3DGS optimization plus 43.21 m for clustering (85.41 m total, PQ_scene 94.6, IoU 97.4); UniC-Lift takes 42 m optimization, 0 m clustering (42 m total, PQ_scene 94, IoU 96.2).
  • Ablation on loss terms (PQ_scene / mIoU): Contrastive loss alone gives 83.7 / 91.8; Contrastive plus 3D regularization gives 88.0 / 94.4; Contrastive plus triplet loss with MLP gives 89.0 / 95.2; without MLP projection 88.0 / 94.0; with MLP projection 89.0 / 95.4. The combination of all three losses is reported as best.
  • Timing of the 3D neighborhood loss (PQ_scene / mIoU): enabling at 7k iterations gives 60.5 / 71.3; 12k gives 65.6 / 77.8; 15k (chosen) gives 65.2 / 74.5; 20k gives 64.6 / 75.0.
  • Boundary mining efficiency: with boundary triplets the method reached PQ_scene of 94 in 25k iterations; with random triplet pairs it required 50k iterations to attain the same quality.
  • Robustness to limited masks: using only 5%, 10% and 20% of segmentation masks on Messy-Rooms produced accuracy comparable to training with all masks, while still using all RGB views for novel-view synthesis.
  • Robustness to mask resolution: on four Messy-Rooms scenes, results at 0.5× resolution showed no significant difference from full resolution (0.25×, 0.5× and 1× were tested).
  • Comparison to Feature-3DGS: rendered feature fields from Feature-3DGS show significant multi-view inconsistencies, so it is not suitable for 3D instance/panoptic segmentation.
  • Qualitative failures of the baseline: Contrastive-Lift consistently fails to predict the leg of a highlighted chair in ScanNet, and misses objects such as pillows on a sofa and objects inside a mini-cupboard in Replica3D.

Methodology in Plain English

Start with multi-view RGB images and their camera poses, plus a set of 2D instance/semantic masks produced by an off-the-shelf 2D segmenter — masks that disagree with each other across views. Model the scene with 3D Gaussian Splatting, but attach one extra learnable vector to every Gaussian. That vector is rendered to each camera view in exactly the same way as color, using the same alpha-blending formula, and is view-independent.

Inside each 2D mask, compute the mean embedding of all pixels in that segment and pull pixels toward their own segment's mean while pushing segment means apart — a clustering loss. Because this distance metric does not guarantee a consistent penalty, the rasterized embedding is passed through a sigmoid (restricting it to [0,1]) and then a linear layer, and a triplet loss with margin δ is applied on the transformed vectors. Anchors, positives and negatives are deliberately sampled along segment boundaries, where triplets are informative rather than "easy." To keep neighboring Gaussians consistent, a 3D neighborhood regularization loss penalizes embedding differences between Gaussian primitives whose centers are within a threshold τ; this is switched on after 15,000 iterations, once adaptive density control has stabilized. The triplet and 3D losses are also applied after the adaptive density control step for stability.

At inference there is no clustering: the rendered embedding is sigmoided, thresholded at 0.5, and the resulting binary vector is read as an integer label. This yields O(n) complexity for novel-view prediction instead of the O(n log c) needed when clustering over n pixels and c clusters. The toy experiment in the paper shows embeddings constrained to a [0,1]² square converging to the corners; a point near the top-right such as [0.9, 0.8] thresholds at 0.5 to [1, 1], which decodes to label 3, while the other corners decode to 0, 1 and 2.

Implementation: built on 3DGS, trained for 30k iterations on a single RTX A6000 with the ADAM optimizer and a learning rate of 1e−4. Embedding dimension d is 12, at most 3,000 triplets are used, and the loss weights λ_cluster, λ_triplet, λ_3D, δ and τ are 1e−1, 1e−1, 1e−1, 1 and 1e−2 respectively. Gradients from the segmentation losses are excluded from the adaptive density control step. A separate semantic embedding per primitive is also learned and rasterized for the PQ_scene computation. Evaluation uses mIoU for label accuracy and scene-level Panoptic Quality (PQ_scene) for label consistency, matching segments of the same class across views and treating IoU > 0.5 as valid matches.

Why This Matters

  • Impact on research: The paper argues that the dominant pipeline for 3D segmentation from 2D masks is structurally two-stage — either preprocessing masks for consistency (Panoptic-Lifting, DM-NeRF, Gaussian Grouping) or clustering learned embeddings afterwards (Contrastive-Lift with HDBSCAN, which the paper calls hyperparameter-sensitive). UniC-Lift removes both stages and reports better quality at a fraction of the training cost, which reframes embedding decoding as a plain thresholding operation rather than a separate algorithmic problem.
  • Real-world applications:
    • AR/VR scene content creation, where objects in a captured room must be individually selectable.
    • Autonomous driving, where consistent 3D object instances are needed in a driving scene.
    • Path planning for robots, where the geometry of distinct objects must be known.
    • Interactive scene editing and mesh extraction, demonstrated in the paper through extracting 3D objects of different shapes and sizes and translating a blue bag to a different location.
  • Industry relevance: The measured cost profile is the selling point for deployment. Prior lifting took more than 15–20 hours per scene; UniC-Lift reports under 40 minutes on an RTX A6000, and its inference is O(n) rather than O(n log c) because no clustering is applied on novel views. Combined with robustness to 5–20% of segmentation masks and to 0.5× mask resolution, this makes the method plausible for pipelines where per-scene 3D annotation is currently too slow.

Future Directions

  1. Extending UniC-Lift to dynamic scenes, which the conclusion explicitly names as future work.
  2. Scaling to large-scale unbounded scenes rather than the room-scale and corridor-scale scenes used here.
  3. Hierarchical segmentation for a 3D representation, so that objects can be decomposed into parts.
  4. Open questions the paper leaves implicit: how the fixed 0.5 thresholding and the embedding dimension d = 12 behave as the number of distinct objects grows far beyond the 500-object Messy-Rooms setting, and whether the hand-set loss weights (all 1e−1) and the fixed 15000-iteration switch for the 3D loss transfer to other scene types. The paper reports an ablation showing earlier (7k) or later (20k) activation degrades accuracy, but only on one additional Replica3D scene.

Target Audience

Researchers and graduate students working on 3D scene understanding, Gaussian Splatting, neural radiance fields, or panoptic/instance segmentation, particularly those who follow the "lifting 2D labels to 3D" line of work and need to know where Contrastive-Lift, Panoptic-Lifting and Gaussian Grouping now stand. It also suits engineers who need practical 3D instance segmentation for AR/VR, robotics or scene-editing products and care about training time per scene. Readers should already be comfortable with 3DGS rendering and contrastive/triplet losses; the paper's core idea is described concretely, but the surrounding literature is assumed.

Authors’ abstract

3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) have advanced novel-view synthesis. Recent methods extend multi-view 2D segmentation to 3D, enabling instance/semantic segmentation for better scene understanding. A key challenge is the inconsistency of 2D instance labels across views, leading to poor 3D predictions. Existing methods use a two-stage approach in which some rely on contrastive learning with hyperparameter-sensitive clustering, while others preprocess labels for consistency. We propose a unified framework that merges these steps, reducing training time and improving performance by introducing a learnable feature embedding for segmentation in Gaussian primitives. This embedding is then efficiently decoded into instance labels through a novel "Embedding-to-Label" process, effectively integrating the optimization. While this unified framework offers substantial benefits, we observed artifacts at the object boundaries. To address the object boundary issues, we propose hard-mining samples along these boundaries. However, directly applying hard mining to the feature embeddings proved unstable. Therefore, we apply a linear layer to the rasterized feature embeddings before calculating the triplet loss, which stabilizes training and significantly improves performance. Our method outperforms baselines qualitatively and quantitatively on the ScanNet, Replica3D, and Messy-Rooms datasets.

Read the original paper