Research
Clutt3R-Seg: Sparse-view 3D Instance Segmentation for Language-grounded Grasping in Cluttered Scenes
Overview Research area: Computer vision and robotics — zero-shot 3D instance segmentation from sparse multi-view RGB input, combined with open-vocabulary language grounding for robotic grasping. Techn

- arXiv
- 2602.11660
- Published
- 2026-02-12
- Authors
- Jeongho Noh, Tai Hyoung Rhee, Eunho Lee, Jeongyun Kim, Sunwoo Lee, Ayoung Kim
AI summary
Overview
Research area: Computer vision and robotics — zero-shot 3D instance segmentation from sparse multi-view RGB input, combined with open-vocabulary language grounding for robotic grasping.
Technical level: Advanced. The paper assumes familiarity with multi-view geometry, 3D point clouds and super-voxels, graph-based mask merging, vision-language embeddings, and 6-DoF grasp pose estimation.
Scope: The paper presents Clutt3R-Seg, a zero-shot pipeline that turns noisy 2D masks from sparse views into view-consistent 3D instances, enriches them with open-vocabulary semantics for text-based target selection, and updates the scene from a single post-interaction image to support multi-stage grasping in cluttered scenes.
What This Paper Is About
Reliable 3D instance segmentation is required for robots that must pick a specific object named in natural language out of a cluttered pile, where occlusion, few camera views, and imperfect 2D masks degrade perception. Existing pipelines assume high-quality per-view masks and therefore inherit over- and under-segmentation errors into the 3D representation; methods that handle dense input also fail to update after objects are removed or displaced. Clutt3R-Seg instead treats noisy masks as informative cues rather than something to be refined away, organizing them into a hierarchical instance tree so that cross-view grouping yields consistent 3D instances and stable language grounding.
Key Contributions
- Hierarchy-based 3D instance segmentation: Noisy mask candidates are organized into per-view hierarchical instance trees (object-cluster, object, sub-object) that enforce a one-proper-segment-per-root-to-leaf-path constraint; bottom-up cross-view grouping of leaf nodes plus residual-node parent substitution produces view-consistent masks and robust 3D instances under sparse multi-view input.
- Language-grounded target identification: Each instance is enriched with open-vocabulary embeddings (Duoduo CLIP) aggregated over multi-view evidence, allowing a robot to identify and localize a specific target object directly from a text prompt.
- Efficient update method for multi-stage grasping: After each grasp, a single post-interaction image is associated with prior instances to preserve segmentation consistency; displaced objects are detected via IoU and their rigid transform optimized with a differentiable chamfer, photometric, and regularization loss.
- Extensive robotic validation and public release: Experiments on the real-world GraspClutter6D dataset, a custom synthetic dataset (NVIDIA Isaac Sim), and on a real robot with a Franka Research 3 arm and Intel RealSense D435i, with source code, simulation environment, and synthetic datasets to be publicly released (code at https://github.com/jeonghonoh/clutt3rseg).
Main Findings
- Class-agnostic segmentation on GraspClutter6D: Clutt3R-Seg reaches AP@25 / AP@50 / AP of 83.32 / 28.25 / 4.58 (easy), 78.04 / 21.99 / 3.71 (intermediate), and 61.66 / 12.15 / 1.97 (difficult), with 57.2 s runtime. MaskClustering (tuned) reaches 59.59 / 5.21 / 0.70, 54.12 / 5.68 / 0.83, and 27.95 / 0.95 / 0.16 at 45.7 s; GraphSeg (tuned) reaches 36.83 / 10.72 / 1.87, 34.60 / 5.64 / 0.97, and 20.18 / 2.42 / 0.39 at 71.9 s.
- Heavy-clutter performance: On the most challenging difficult sequences, 61.66 AP@25 is over 2.2× higher than the baselines, and the advantage grows as difficulty increases while baseline scores drop sharply.
- Synthetic dataset results: Clutt3R-Seg reaches 94.40 / 64.78 / 16.06 AP@25 / AP@50 / AP at 21.5 s, versus GraphSeg at 61.32 / 9.89 / 3.28 (51.9 s) and MaskClustering at 33.77 / 0.08 / 0.01 (40.7 s).
- Language-grounded 3D semantic segmentation: On GraspClutter6D, Clutt3R-Seg reaches 52.54% 3D IoU in 47.2 s, versus ConceptGraphs 32.59% (148.5 s) and GraspSplats 17.38% (81.1 s). On the synthetic dataset it reaches 55.10% (24.8 s) versus ConceptGraphs 34.90% (91.2 s) and GraspSplats 20.70% (67.5 s). The CLIP-backbone variant, Clutt3R-Seg-, reaches 50.61% (47.6 s) on GraspClutter6D and 41.69% (23.7 s) on synthetic, still about 20% above the baselines.
- Ablations: Removing spatial similarity lowers results to 66.77 / 16.33 / 2.61 (55.2 s); removing substitution lowers them to 74.10 / 20.60 / 3.19 (57.0 s); the full method reaches 74.34 / 20.79 / 3.41 (57.2 s). Substitution adds less than 0.5 s of overhead.
- Viewpoint sparsity: Clutt3R-Seg with 4 views achieves 69.00 / 16.54 / 2.67 at 23.4 s, with 6 views 72.46 / 19.46 / 3.15 at 37.6 s, and with 8 views 74.34 / 20.79 / 3.41 at 57.2 s — outperforming GraphSeg (8 views, 30.54 / 6.20 / 1.08, 71.9 s) and MaskClustering (8 views, 47.22 / 3.95 / 0.56, 45.7 s) at every metric despite fewer views; the paper states that with four input views it surpasses MaskClustering with eight views by more than 2×.
- Why baselines fail: MaskClustering's view-consensus weighting becomes unreliable when the observing set is small under sparse, occluded views; GraphSeg under-segments occluded regions due to fixed 2D thresholds and over-segments textured objects where small noisy 2D fragments do not merge by chamfer distance. SAI3D was excluded from all comparisons because its cross-view affinity requires co-visibility and discards invalid views, yielding an unreliable affinity matrix in sparse, cluttered scenes.
- Qualitative behavior: Baselines show severe under-segmentation across easy, intermediate, and difficult sequences in both real and synthetic data, while Clutt3R-Seg segments distinct instances correctly.
Methodology in Plain English
The pipeline has three stages.
Preprocessing. From posed sparse RGB images, the method estimates depth (MVSAnywhere) to reconstruct a point cloud, and generates 2D masks by prompting Grounded SAM with the single generic token "object" — chosen because the present categories are unknown, avoiding per-scene prompt engineering. The point cloud is voxel-downsampled (retaining averaged appearance and geometry), then k-NN clustering that combines Euclidean distance and surface-normal cosine similarity merges geometrically homogeneous regions into compact "super-voxels."
Hierarchy-based 3D instance segmentation. Because Grounded SAM produces a mix of proper, under-segmented (too large) and over-segmented (too small) masks, and because these errors appear irregularly across views, each view's masks are organized into a tree where a child node is a mask contained within its parent. Along any root-to-leaf path at most one mask is the correct segment, so leaf nodes serve as the proper-segment candidates. Leaves from different views are pooled into a graph where every cross-frame pair carries two scores: a spatial similarity (a weighted Jaccard index over super-voxel occupancy) and a semantic similarity (cosine similarity of Duoduo CLIP embeddings). Grouping runs in two stages — first merging the highest-scoring spatial pairs above a threshold, then the highest-scoring semantic pairs — and each merge rewires the graph. Leftover "residual" nodes, typically fragments from over-segmentation, are conditionally substituted by their parent node (only when all the parent's other descendants are also residual) and re-grouped; the paper reports this converges in a single iteration. Finally, the grouped 2D masks are projected onto super-voxels by point-coverage-weighted majority voting, and group size defines instance confidence.
Language grounding and update. Each instance group is embedded with the multi-view Duoduo CLIP encoder (background outside the mask replaced with white), and a text query is matched by similarity to select the target; a grasp pose detector [16] estimates a 6-DoF grasp. After an interaction, one post-interaction RGB image is processed, new noisy masks are matched to existing instances using the same hierarchy-based grouping and stored embeddings, and instances whose reprojected point cloud has IoU below τ_IoU are flagged as displaced. Their rigid body transform is optimized with a combined chamfer (2D contour), photometric, and vertical-motion regularization loss in a coarse-to-fine schedule that terminates early if IoU already exceeds τ_IoU. Only displaced instances are updated; static ones are left untouched.
Experimental setup. GraspClutter6D provides real multi-view RGB-D sequences at 1920×1080; the custom Isaac Sim synthetic dataset is at 640×480 and uses objects from HouseCat6D across 10 household categories. The 99 GraspClutter6D sequences are split by average visibility into easy, intermediate, and difficult subsets with mean overlap ratios of 11.62%, 19.83%, and 31.71%. Eight posed images per sequence are used, selecting approximately uniformly spaced frames; evaluation uses 10 mm voxel grids and five visible object categories. Key parameters: 5 mm voxel aggregation with minimum occupancy m=3, α=0.5, β=1.0, τ_merge=0.01, τ_spat=0.5, τ_sem=0.65, τ_IoU=0.75, and loss weights of λ_chamfer=50.0, λ_photo=0.5, λ_reg,z=10.0 in stage 1 and 10.0, 2.0, 1.0 in stage 2. All methods run on a single NVIDIA RTX 3090.
Why This Matters
Impact on research. The paper reframes noisy 2D masks as useful structural cues rather than errors to be cleaned up, showing that a hierarchy over masks — with a one-proper-segment-per-path constraint — yields better cross-view consistency than affinity or graph-merging approaches built on the assumption of high-quality masks. It also demonstrates that a scene representation can be updated from a single post-interaction image rather than dense consecutive frames, which decouples multi-stage manipulation from costly rescanning.
Real-world applications.
- Warehouse and logistics picking, where robots must retrieve a specific item named in an instruction from a bin of overlapping objects.
- Manufacturing and assembly, where parts are jumbled and a text command selects which component to grasp.
- Household service robotics, where a user names an object on a cluttered shelf or table.
- Multi-stage tabletop clearing or rearrangement, where objects move between grasps and the scene representation must stay synchronized.
Industry relevance. The pipeline is zero-shot (no per-scene training or prompt engineering), runs on a single RTX 3090, and reports the shortest runtime among the compared language-grounded methods (47.2 s versus 148.5 s and 81.1 s on GraspClutter6D) while achieving higher 3D IoU. That combination of training-free deployment, sparse-view input, and demonstrated on-hardware execution with a Franka Research 3 and Intel RealSense D435i is directly relevant to industrial manipulation deployments with constrained sensing and compute budgets. The work was supported by Hyundai Motor Company and Kia.
Future Directions
- Improving depth estimation fidelity, since the paper attributes lower reconstruction quality at stricter thresholds to predicted depth.
- Handling severe occlusion that can prevent recovery of fully hidden objects or limit the completeness of single-view updates.
- Extending the framework to handle newly emerging objects in dynamic scenes.
- Incorporating multi-modal sensing to improve reconstruction quality.
- The paper also notes the runtime/accuracy trade-off from reducing the number of views, suggesting a lightweight configuration mode as a direction for deployment (e.g., 4 views at 23.4 s versus 8 views at 57.2 s on GraspClutter6D).
Target Audience
Researchers and engineers working on 3D perception for robotics — particularly those combining multi-view segmentation, open-vocabulary vision-language grounding, and manipulation. It is most useful to readers already comfortable with point clouds, graph-based mask merging, and 6-DoF grasping; readers new to the field will need to consult the cited baselines (Grounded SAM, MVSAnywhere, Duoduo CLIP, GraphSeg, MaskClustering, SAI3D, ConceptGraphs, GraspSplats) for background.
Authors’ abstract
Reliable 3D instance segmentation is fundamental to language-grounded robotic manipulation. Its critical application lies in cluttered environments, where occlusions, limited viewpoints, and noisy masks degrade perception. To address these challenges, we present Clutt3R-Seg, a zero-shot pipeline for robust 3D instance segmentation for language-grounded grasping in cluttered scenes. Our key idea is to introduce a hierarchical instance tree of semantic cues. Unlike prior approaches that attempt to refine noisy masks, our method leverages them as informative cues: through cross-view grouping and conditional substitution, the tree suppresses over- and under-segmentation, yielding view-consistent masks and robust 3D instances. Each instance is enriched with open-vocabulary semantic embeddings, enabling accurate target selection from natural language instructions. To handle scene changes during multi-stage tasks, we further introduce a consistency-aware update that preserves instance correspondences from only a single post-interaction image, allowing efficient adaptation without rescanning. Clutt3R-Seg is evaluated on both synthetic and real-world datasets, and validated on a real robot. Across all settings, it consistently outperforms state-of-the-art baselines in cluttered and sparse-view scenarios. Even on the most challenging heavy-clutter sequences, Clutt3R-Seg achieves an AP@25 of 61.66, over 2.2x higher than baselines, and with only four input views it surpasses MaskClustering with eight views by more than 2x. The code is available at: https://github.com/jeonghonoh/clutt3r-seg.