Skip to content
AI.info

Research

Finding NeMO: A Geometry-Aware Representation of Template Views for Few-Shot Perception

Overview Research area: Computer vision — few-shot object perception (detection, segmentation, and 6DoF pose estimation) for objects never seen during training, using only RGB template images. Technic

Finding NeMO: A Geometry-Aware Representation of Template Views for Few-Shot Perception
arXiv
2602.04343
Published
2026-02-04
Authors
Sebastian Jung, Leonard Klüpfel, Rudolph Triebel, Maximilian Durner

AI summary

Overview

  • Research area: Computer vision — few-shot object perception (detection, segmentation, and 6DoF pose estimation) for objects never seen during training, using only RGB template images.
  • Technical level: Intermediate. The paper assumes familiarity with Vision Transformers, attention/transformer decoders, unsigned distance fields, dense prediction heads (DPT), and standard pose-estimation tooling such as PnP and RANSAC.
  • Scope: The paper introduces a geometry-aware, object-centric representation called the Neural Memory Object (NeMO) built from a few RGB views of an object, and evaluates a single encoder-decoder network on multiple perception tasks across the BOP benchmark, in both model-free and model-based settings.

What This Paper Is About

Most strong object perception systems either require a 3D CAD model of every target object or require fine-tuning a network for each new object. The authors target the harder case: given only a small, unordered set of RGB images of an unseen object (with no camera calibration or pose annotations), can a single network detect it, segment it, and estimate its 6DoF pose without any retraining? Their answer is NeMO, a compact point-based "memory" of the object that is produced by an encoder offline and consumed by a decoder at inference time.

Key Contributions

  1. The NeMO representation. A geometry-sensitive, compact, object-centric encoding of template views, expressed as a sparse, continuous point cloud with semantic and geometric features, built without intrinsic or extrinsic camera parameters.
  2. A joint evaluation across perception tasks and settings. The encoder-decoder network is evaluated on model-based and model-free few-shot detection, segmentation, and pose estimation of unseen objects, without training on test objects, reporting competitive and partially better results than state-of-the-art methods.
  3. A new synthetic dataset. An object-centric, class-balanced synthetic dataset that mimics realistic cluttered scenes, built from 11,077 different objects drawn from a subset of Objaverse, GSO, and OmniObject3D, rendered with BlenderProc.
  4. Practical properties of the representation. Constant decoder inference time and memory regardless of the number of template images, support for dynamic point counts (memory/precision trade-off), transformability of the NeMO point cloud, and extendability by merging NeMOs.

Main Findings

  • Model-free detection: On the BOP test splits of HOPEv2 and HANDAL, the method reports AP of 0.411 and 0.273 respectively, which the paper describes as the previous best outperformed by 2.7pp on HOPEv2 and 1.9pp on HANDAL. Comparison points shown include GFreeDet-SAM (0.384, 0.264), GFreeDet-FastSAM (0.364, 0.255), dounseen-SAM-CTL (0.380, not reported on HANDAL), and CNOS (SAM) with static onboarding (0.345, not reported on HANDAL).
  • The paper claims a first: All other compared methods rely on SAM segmentations of the scene for bounding boxes; the authors state they are, to their knowledge, the first to use a single network to predict amodal segmentations/detections in a model-free setting, which they convert into amodal bounding boxes.
  • Model-free 6DoF pose estimation: Using NeMO detections, the method reports AP of 0.302 on HOPEv2 and 0.235 on HANDAL, versus OPFormer with CNOS detections at 0.335 (HOPEv2) and 0.204 (HANDAL). With CNOS detections the method reports 0.307 on HOPEv2 (2.8pp behind OPFormer); with GFreeDet-FastSAM detections it reports 0.329 and 0.213. On HOPEv2, ICP refinement against depth is used; on HANDAL no refinement is used because no depth data is available.
  • Model-based detection: AP of 0.623 on TUD-L is reported as state of the art, outperforming the previous best by 3pp. On YCB-V the method reports 0.602 and the paper states it would rank 6th out of the 13 methods reported on the BOP leaderboard. On T-LESS it reports 0.183.
  • Model-based segmentation: AP of 0.488 on TUD-L and 0.579 on YCB-V, with 0.169 on T-LESS. The paper attributes lower pixel-level precision than competing networks to border artifacts from patch-scaling, and the low T-LESS numbers to object similarity, lack of texture, and cluttered same-instance scenes that are absent from the training data.
  • Model-based, refiner-free 6DoF pose (AR with CNOS detections): 0.190 on T-LESS, 0.476 on TUD-L, and 0.504 on YCB-V; the paper notes only Co-op (0.592 / 0.642 / 0.626) achieves better results. With ground truth detections the method reports 0.295 / 0.538 / 0.566, and with its own NeMO detections 0.082 / 0.466 / 0.493 — lower than with CNOS detections despite NeMO's higher detection precision, which the paper attributes to differences in how AP and AR are evaluated.
  • Number of template images: More template images increase 6DoF pose AP on YCB-V with ground truth detections, but the paper states acceptable performance is already achieved with just three views. Offline NeMO generation memory grows with the number of templates, while decoder inference memory stays constant.
  • Number of NeMO points: Pose AP on YCB-V rises as points increase (10 points: 0.004 AP; 50: 0.112; 100: 0.214; 200: 0.290; 500: 0.380; 1000: 0.383; 1500: 0.378), stagnating around 0.38 from 500 points, indicating the model adapts to point-cloud sizes it was not trained on.
  • Rotation equivariance: Rotating the NeMO point cloud around the z-axis in 10-degree steps produces pointmaps that follow the rotation, with Chamfer distance to the correspondingly rotated ground truth CAD model staying low.
  • Extending NeMO: Combining two NeMOs of the same object with the same anchor image yields a lower Chamfer distance to the ground truth CAD model than the original NeMO, without disturbing the decoder.
  • Limitations reported: Difficulties with symmetric objects (seen on T-LESS), poor performance on highly textureless objects, and the fact that the encoder does not directly predict bounding boxes — segmentation masks are used as a proxy, which can merge bounding boxes when multiple instances of the same object are present.

Methodology in Plain English

The system is split into an encoder and a decoder.

The encoder takes a set of RGB template images of an object (at least two, resized to 224 × 224 crops). A ViT — specifically a DINOv2 backbone — extracts patch features from each image, and a multi-view encoder using cross- and self-attention combines them while anchoring the object's orientation to one randomly chosen "anchor image." Separately, a random set of 3D points in the cube [-1, 1]³ is passed through an MLP to produce point features. In a "Geometric Mapping" block, these point features act as queries that attend to the image features, letting 3D points absorb object-specific 2D appearance. A jointly learned MLP-based unsigned distance field (UDF) then predicts how far each point is from the object surface, and the points are pushed onto that surface. The result is the NeMO: a set of surface points, each paired with a feature vector, written as χ = {(s_i, f_i^Q)} for i = 1 to M. Notably, no loss is applied directly to the NeMO features — they are learned only through gradients flowing back from the decoder.

The decoder takes a query image and the NeMO. Query image features from the ViT are updated by attending to the NeMO features via cross- and self-attention, and a DPT head with multiple output heads upscales them into: a modal segmentation mask, an amodal segmentation mask, a dense pointmap linking each 2D pixel to a 3D surface point in the NeMO coordinate system, and a confidence map for that mapping. After filtering by confidence, RANSAC and Perspective-n-Point (PnP) recover the object's pose.

Training is end-to-end on synthetically rendered RGB-D images with ground truth poses, masks, and intrinsics. Losses include a regression loss pulling predicted surface points toward ground truth surface points, dice plus binary cross-entropy for the two segmentation masks, a confidence-weighted L1 loss for the pointmap, and two terms that push confidence high on object pixels and low on background pixels. During training, the NeMO points are randomly rotated, translated, and scaled before decoding, which teaches the decoder a coordinate system independent of the anchor image.

Practical design choices: Because NeMO is precomputed offline, inference cost does not grow with the number of template views. The representation also supports varying the number of points for a memory/precision trade-off, applying transformations to the point cloud, and appending new points from another NeMO of the same object.

Why This Matters

The work argues for decoupling object knowledge from network weights: instead of retraining or fine-tuning for each new object, the object lives in an external, interpretable point-based structure that can be generated once and reused. This is a different scaling story from template-matching pipelines that rely on pairwise local comparisons and scale poorly with the number of templates. The authors also release a large, class-balanced synthetic object dataset intended to support broader research, and the code is available at the project's GitHub repository.

Real-world applications implied by the paper:

  • Robotics manipulation: Quickly onboarding a new object for grasping without a CAD model, using detection, segmentation, and pose estimation from a single pipeline.
  • Augmented reality: Placing and tracking virtual content on novel physical objects captured with ordinary cameras.
  • Autonomous systems: Recognizing previously unseen objects in cluttered scenes where 3D models are unavailable.
  • Multi-stage perception pipelines: Using amodal detections from NeMO to crop regions of interest for downstream modules, which the paper demonstrates qualitatively.

Industry relevance: the work comes from the German Aerospace Center (DLR) and targets the practical bottleneck of object onboarding in real deployments — the paper emphasizes no camera-specific parameters, no retraining on target data, constant inference cost regardless of template count, and applicability to images captured with a normal smartphone.

Future Directions

  • Addressing symmetric objects. The low T-LESS results are attributed not to the pointmap representation itself but to limitations in the current training procedure, which the authors say requires further research.
  • Scaling and diversifying training data. Highly textureless objects perform poorly, likely due to underrepresentation in the training set; the authors plan to include a broader, more diverse set of objects.
  • Direct bounding box prediction. Because the encoder currently uses segmentation masks as a proxy, multiple instances of the same object can yield merged bounding boxes — an issue slated for future work.
  • Articulated objects and online transformation. The authors propose combining multiple NeMOs for articulated objects and note the rotation-equivariance of the pointmap as an interesting property for transforming parts of the NeMO online during inference.

Target Audience

Researchers and practitioners working on few-shot or unseen-object perception, 6DoF pose estimation, and the BOP benchmark; robotics and augmented-reality engineers who need fast object onboarding without CAD models or retraining; and readers interested in object-centric representations that separate object information from model weights, including those working with neural fields, distance fields, and transformer-based dense prediction.

Authors’ abstract

We present Neural Memory Object (NeMO), a novel object-centric representation that can be used to detect, segment and estimate the 6DoF pose of objects unseen during training using RGB images. Our method consists of an encoder that requires only a few RGB template views depicting an object to generate a sparse object-like point cloud using a learned UDF containing semantic and geometric information. Next, a decoder takes the object encoding together with a query image to generate a variety of dense predictions. Through extensive experiments, we show that our method can be used for few-shot object perception without requiring any camera-specific parameters or retraining on target data. Our proposed concept of outsourcing object information in a NeMO and using a single network for multiple perception tasks enhances interaction with novel objects, improving scalability and efficiency by enabling quick object onboarding without retraining or extensive pre-processing. We report competitive and state-of-the-art results on various datasets and perception tasks of the BOP benchmark, demonstrating the versatility of our approach. https://github.com/DLR-RM/nemo

Read the original paper