Skip to content
AI.info

Research

Beyond the Current Scene: Event-Referential Grasping with Active View Selection

Beyond the Current Scene: Event-Referential Grasping with Active View Selection Overview Research area: Robotics — robotic grasping, language grounding, episodic memory, and active perception. Technic

Beyond the Current Scene: Event-Referential Grasping with Active View Selection
arXiv
2609.39375
Published
2026-09-30
Authors
Hyunjoon Lee, Haebeom Jung, Eunsung Cha, Daeun Lee, Yu-Chiang Frank Wang, Jaesung Choe, Jaesik Park

AI summary

Beyond the Current Scene: Event-Referential Grasping with Active View Selection

Overview

Research area: Robotics — robotic grasping, language grounding, episodic memory, and active perception.

Technical level: Intermediate. The paper is readable by someone familiar with robot manipulation and vision-language models, but the method section assumes comfort with probabilistic spatial beliefs, truncated signed distance function (TSDF) maps, and Bayesian likelihood updates.

Scope: The paper presents BeyondCSe, a zero-shot system that resolves grasping instructions referring to a past event and actively moves a wrist-mounted camera to find the target when it is no longer visible.

Authors and affiliations: Hyunjoon Lee, Haebeom Jung, Eunsung Cha, Daeun Lee, and Jaesik Park (Seoul National University); Yu-Chiang Frank Wang and Jaesung Choe (NVIDIA). The paper is listed as arXiv:2609.39375v1 [cs.RO], 30 Sep 2026, under CC BY-NC-SA 4.0.

What This Paper Is About

A robot that has watched a person handle objects should later be able to act on instructions that refer back to that interaction, such as "pick up the object I just used." These event-referential instructions identify a target by the role it played in a past event rather than by its name, category, or appearance, and the target may already be hidden or occluded when the request is made. The paper's goal is a system that recovers both the identity of the requested object or part and the spatial evidence of where it went, then actively chooses camera viewpoints that reveal it so it can be grasped.

Key Contributions

  1. A zero-shot event-referential grasping pipeline. The system grounds objects or parts referred to by past events and recovers an event prior using an off-the-shelf multimodal large language model (MLLM) and point tracker, without additional task-specific training.

  2. A probabilistic volumetric belief formulation. The method integrates an event-conditioned spatial prior with transmittance-aware visibility for Bayesian belief updates and viewpoint selection, discounting candidate views whose sight lines pass through unobserved space rather than treating that space as free.

  3. A real-robot evaluation with a single wrist-mounted RGB-D camera. The study reports grasping and active-perception outcomes, including how event-history evidence affects search cost and grasp success.

  4. Ablations isolating each belief component. Variants without belief weights, without transmittance, and without the event prior are compared under heavy occlusion while keeping candidate views, feasibility checks, target confirmation, and the grasp pipeline fixed.

Main Findings

  • Visible-target grasping: With a single wrist-mounted RGB-D camera, the system grasps successfully in 38/50 visible trials (76%) in real-robot experiments, versus 20/50 (40%) for the strongest baselines in that condition.

  • Occluded-target grasping: The system succeeds in 77/100 occluded trials (77%), versus 55/100 (55%) for the strongest baselines in that condition.

  • Pooled result across 150 trials: The method reaches 76.7% grasp success, compared with 50.0% for the strongest baseline (Point2Act † using reasoned instructions, per Table I). Under original instructions, baselines score 0.0 (LERF-TOGO), 9.3 (Point2Act), 0.0 (GraspSplats), and 22.7 (Point2Act †); under reasoned instructions they score 18.7, 19.3, 34.7, and 50.0 respectively.

  • Localization advantage: The method localizes 43/50 visible and 88/100 occluded targets, exceeding the strongest baselines by 32 and 20 percentage points, respectively.

  • Heavy-occlusion search: On four additional heavily occluded scenes, the full method achieves 95% grasp success (19/20) with 2.20 mean views, versus 75% (15/20) and 3.35 mean views for the Breyer et al. active-perception baseline — even though that baseline is given the target's ground-truth 3D bounding box.

  • Belief-weight ablation: Removing posterior-weighted view ranking raises the mean view count to 4.65 and lowers grasp success to 80% (16/20), suggesting views are wasted on unlikely regions.

  • Transmittance ablation: Setting σ_t = 0 raises the mean view count to 3.10, with scene S3 requiring 5.2 views, consistent with optimistic scoring of sight lines through unknown space.

  • Event-prior ablation: Without the history-derived prior, grasp success matches the full method at 95% (19/20), but the mean view count rises from 2.20 to 4.00, showing the search-cost benefit of event history.

  • Qualitative generalization: Applied without dataset-specific prompt changes to egocentric clips from EgoDex and EPIC-KITCHENS, the pipeline selects targets referred to by order of plate placement or washing and produces 2D points on the requested objects in the final frames. In the EPIC-KITCHENS example, which includes large camera viewpoint changes, Locate proposes an inaccurate region but Point identifies the carrot within the expanded crop.

Methodology in Plain English

The system ingests a video of a person's tabletop activity, recorded by a wrist-mounted camera that stays stationary, plus a language instruction issued afterward. The final video frame becomes the robot's initial observation, and the robot may move the camera afterward.

Video reasoning. Four stages handle the instruction. Record watches the video without the instruction and builds a time-ordered event record, where each event is described by an action, its object, and another involved object if any. Select reads that record and the instruction and outputs an appearance description of the requested object or part (the target) and an event description that distinguishes it (the cue). Locate uses the video and cue to propose a box in the initial image; Point receives the target and the crop and returns an action pixel. Cropping narrows the search region and enlarges small parts; if crop-based pointing fails, the region is expanded, and if no usable box exists, the full image is used. The pixel is mapped to the full image and lifted to a 3D action point using valid depth and camera pose.

Recovering historical locations. When pointing in the initial image fails, the system searches sampled past frames from recent to earlier ones, tracks candidate pixels together, and rejects trajectories that remain observed at the end of the video. Remaining candidates are checked for appearance, most recent first, and the first to pass is selected. Its track samples with valid depth are lifted into the robot base frame to form a 3D event track.

Event-conditioned belief. The initial map is a TSDF built from the initial RGB-D observation. The target's location is represented as a probability density over the workspace, including regions unresolved by occlusion or missing depth. The belief is initialized from the 3D event track: it is centered on the last observation and shaped by the terminal segment of the path (a tail covariance), without extrapolating the path. A small uniform mixture keeps nonzero probability everywhere so the search can recover if the historical estimate is wrong; with no valid 3D history, the belief is uniform.

View selection. Candidate views are generated around high-probability regions and filtered for kinematic feasibility, self-collision, and scene collision. For each candidate, the system renders a predicted depth map from the current map and evaluates how likely a negative observation is at each target hypothesis. A transmittance term attenuates visibility confidence according to how far a ray travels through unobserved space, inspired by accumulated volumetric transmittance in NeRF but without learning a radiance field. Each view is scored by the target-belief mass that a negative observation is expected to downweight, and the highest-scoring view with a valid motion plan is executed.

Update and refinement. After each new keyframe, the depth is integrated into the map and the belief is updated with a depth-consistency likelihood and a target-miss likelihood, both taking values in (0, 1] so negative observations downweight rather than rule out locations. The loop ends when pointing returns a confirmed 3D point, or when no feasible view remains or the active-view budget is exhausted. Grasp candidates are then generated and validated; incomplete target geometry can trigger one target-centered refinement view before grasp regeneration, following Breyer et al., though unlike that setting the target region only becomes available after the event-conditioned search confirms it.

Experimental setup. All real-world experiments use a ROBOTIS OMY-F3M robot with a wrist-mounted Intel RealSense D435i RGB-D camera; perception models run on a workstation with a single NVIDIA GeForce RTX 4090 GPU. The evaluation set comprises 10 visible scene-query pairs across four scenes and 20 occluded scene-query pairs from 10 scenes, with five trials each, giving 50 visible and 100 occluded trials (150 total). Metrics are 3D localization, planning, and grasp success, with human-annotated 3D oriented bounding boxes withheld from all methods. The pipeline uses Qwen3-VL-8B for all MLLM stages, 96 uniformly sampled frames, deterministic decoding at temperature zero, and settings of ρ = 2, δ₀ = 0.05 m, ε = 0.1, κ = 0.33, and σ_t estimated from newly resolved voxels (2.0 m⁻¹ fallback). AnyGrasp candidates are generated from virtual views following LERF-TOGO's procedure.

Why This Matters

Impact on research. The work separates two problems that prior systems often conflate: identifying a target from history and obtaining the visual evidence needed to grasp it. It shows that an active-perception baseline given a privileged ground-truth target box still achieves only 75% grasp success with 3.35 views, while the event-conditioned method reaches 95% with 2.20 views. It also contributes a concrete evaluation protocol for event-referential grasping, including visible and occluded conditions and heavily occluded scenes.

Real-world applications:

  • Assistive and service robots in homes or kitchens, responding to requests like "hand me the utensil I was just using" after the item has been put away or covered.
  • Warehouse and logistics picking, where a supervisor's earlier action (for example, placing an item into a bin) defines the pick target rather than a barcode or label.
  • Laboratory and clinical support, where a tool's identity is defined by the procedure step it was used in rather than by its appearance among similar instruments.
  • Collaborative manufacturing cells, where a robot must fetch a part that a human moved earlier in an assembly sequence and that is now behind other objects.

Industry relevance. The system is zero-shot, relies on off-the-shelf models (an MLLM and a point tracker) without task-specific training, and uses a single wrist-mounted RGB-D camera plus a robot arm. That combination lowers the barrier to deploying event-referential grasping on existing hardware, and the reported reduction in camera moves from 3.35 to 2.20 views directly translates into faster task completion.

Future Directions

  • Targets that keep moving. The formulation assumes the target remains stationary during search; the authors state that extending the system to interactions in which objects continue to move remains future work.

  • Robust region estimation in the final frame. The EPIC-KITCHENS example shows Locate proposing an inaccurate region under large camera viewpoint changes, with Point recovering only after the crop is expanded — indicating final-frame region estimation can remain imprecise.

  • Retention of the event prior under weaker history. Without the event prior, grasp success stayed at 95% but view count rose from 2.20 to 4.00, raising the question of how the method degrades when only sparse or noisy 3D history is available.

  • Broader deployment beyond the capture setup. The qualitative results on EgoDex and EPIC-KITCHENS were obtained without dataset-specific prompt changes, suggesting further testing outside a wrist-mounted, stationary-capture tabletop setting.

Target Audience

Researchers and practitioners in robot manipulation, embodied AI, and vision-language robotics who are interested in episodic memory, reference resolution beyond the current scene, or active perception under occlusion. It also suits engineers building grasping systems on arm-plus-wrist-camera platforms, and readers evaluating how far zero-shot, off-the-shelf multimodal models can be pushed on real-robot tasks without task-specific training.

Authors’ abstract

A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests may specify a grasp target by the role it played in a past event rather than by its name or appearance. Moreover, the target may no longer be visible when the robot is asked to act. We present BeyondSCe, a zero-shot robotic grasping system for this event-referential setting. Given the event history and the current scene, the system identifies the requested object or part and localizes it for grasping. If the target is occluded, it combines an event prior recovered from the history with current scene geometry to select camera viewpoints likely to reveal the target. The system uses pretrained models without additional task-specific training. In real-robot experiments with a single wrist-mounted RGB-D camera, it achieves grasp success rates of 76% and 77% for initially visible and occluded targets, respectively, compared with 40% and 55% for the strongest baseline in each condition. On four additional scenes with heavy occlusion, it increases grasp success rates from 75% to 95% while reducing the mean number of views from 3.35 to 2.20, compared with an active-perception baseline given the target's ground-truth 3D bounding box.

Read the original paper