Skip to content
AI.info

Research

Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation

Overview Research area: Robot learning — memory-augmented manipulation policies and benchmarks for partially observable (POMDP-style) manipulation tasks. Technical level: Advanced. The paper assumes f

Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
arXiv
2609.38886
Published
2026-09-30
Authors
Yansong Shi, Jiange Yang, Xijie Yang, Shaowei Zhang, Yuhan Zhu, Tao Lu, Limin Wang

AI summary

Overview

Research area: Robot learning — memory-augmented manipulation policies and benchmarks for partially observable (POMDP-style) manipulation tasks.

Technical level: Advanced. The paper assumes familiarity with vision-language-action (VLA) models, behavior cloning, attention-based memory, and POMDP formulations.

Scope (one sentence): The paper introduces HIDE, a 15-task RLBench-based benchmark that isolates decision-relevant "hidden task states" in robotic manipulation, and SEEK, a three-mechanism memory framework that improves success on that benchmark in simulation and on four real-robot tasks.

What This Paper Is About

Most manipulation policies act on the current observation, but many tasks depend on information that is no longer visible — how many repetitions have been completed, what object was seen earlier, or which substeps are done. The paper formalizes these as "hidden task states" that cannot be inferred from the current visual and proprioceptive observation but can be recovered from interaction history, and argues that retaining history alone does not guarantee recovering the right hidden state. To study this, the authors build HIDE, a benchmark with decision points where visually similar observations require different actions, and SEEK, a policy framework combining recent context, retrieved historical evidence, and an explicit progress counter.

Key Contributions

  1. A hidden-state benchmark (HIDE). A 15-task benchmark built on RLBench, organized into three categories — repetition counting, historical-state recall, and execution-progress tracking — with randomized initializations, appearance variations, automated demonstration generation, and structured hidden-state (stage) annotations. Tasks are defined as hidden-state-dependent only when two histories produce visually equivalent observations but require different optimal actions.

  2. A memory framework (SEEK). Three complementary mechanisms combined into a coarse-to-fine, language-conditioned multi-view policy: Windowed Context Memory (WCM) for recent interactions, Persistent Anchor Memory (PAM) for retrieving older evidence after it leaves the window, and Stage-Counter Memory (SCM) for tracking execution progress via a discrete counter plus a learned stage embedding.

  3. Systematic memory analysis. Ablations over what to store (proprioception vs. visual features), how much history to retain (memory length 0, 2, 4, 6), and which components contribute, evaluated both individually and in combination across categories.

  4. Evaluation beyond the benchmark. Testing on standard manipulation (18-task RLBench), perturbation robustness (The Colosseum), and four real-robot tasks, to check that memory does not trade off general performance.

Main Findings

  • Existing policies struggle on HIDE. Memory-based methods help but leave category-specific gaps: SAM2Act+ improves the average from 42.7% (SAM2Act) to 51.2%, and μVLA improves from 13.9% (OpenVLA-OFT) to 18.1%. SAM2Act+ reaches 60.8% on execution-progress tracking but only 46.4% on repetition counting and historical-state recall, while MME achieves 35.2% on historical-state recall but only 2.4% on repetition counting. The strongest baseline differs by category: SAM2Act+ leads repetition counting and execution-progress tracking, while GR00T-N1.7 leads historical-state recall at 49.6%.

  • SEEK leads overall and in every category. SEEK reaches 62.9% average success, 11.7 percentage points above SAM2Act+ (51.2%). Category scores are 61.6% on repetition counting, 59.2% on historical-state recall, and 68.0% on execution-progress tracking — gains of 15.2, 9.6, and 7.2 points over the strongest category-wise baselines.

  • What to store matters more than architecture. With encoding dimensionality and memory architecture fixed, storing visual features raises average success from 42.5% to 62.9% versus proprioceptive histories. Gains span all categories and are largest on repetition counting (+30.4 points), followed by historical-state recall (+17.6) and execution-progress tracking (+13.4).

  • Longer memory is better, but not monotonically at first. Average success rises from 46.7% at memory length 0 to 53.7% at length 2, 57.2% at length 4, and 62.9% at length 6, with length 6 best in every category. Repetition counting initially drops from 46.4% to 44.2% at length 2 before rising to 61.6% at length 6; execution-progress tracking saturates gradually, from 67.2% at length 4 to 68.0% at length 6.

  • Memory components have distinct strengths. Adding SCM, WCM, and PAM progressively improves all three categories, with the largest gain at each step on repetition counting from SCM (+14.0 points), on execution-progress tracking from WCM (+12.0), and on historical-state recall from PAM (+11.6). Individually, SCM leads on repetition counting, PAM on historical-state recall, and WCM on execution-progress tracking — but their combination leads all three categories.

  • Memory does not cost standard-task performance. SEEK scores 84.7% average success on 18-task RLBench, slightly above the authors' SAM2Act reproduction (84.1%) and above RVT-2 (81.4%) and RVT (62.9%).

  • Robustness is preserved. On The Colosseum, SEEK achieves 68.6% on Clean and 61.9% averaged over perturbations (a 9.8% relative drop), compared with 68.4% and 61.5% (10.1% drop) for the SAM2Act reproduction.

  • Large real-world gains. Across four real-robot tasks (repeated button pressing, cup stacking, desk cleaning, and searching for a hidden white piece), SEEK reaches 89% average success versus 47% for SAM2Act+ and 13% for π0.5. SEEK scores 76% on button pressing, 100% on cup stacking, 80% on desk cleaning, and 100% on the search task.

  • Per-task highlights. On individual HIDE tasks, SEEK reaches 100% on wipe desk, 100% on reopen drawer, 92% on push button, 92% on swap pegs, and 88% on lift and check.

Methodology in Plain English

The authors first define the problem precisely: a task counts as hidden-state-dependent only if two different interaction histories can lead to near-identical current observations but require different optimal actions. That definition separates genuine memory dependence from mere task length or occlusion.

They then build HIDE on top of RLBench, reusing its robot, object, and scene assets. Each task specifies object initialization ranges, randomized scenes, and a sequence of manipulation keypoints, so the simulator can generate many variations with different colors, materials, and counts. Demonstrations come from RLBench's scripted expert pipeline; afterward, trajectories are automatically segmented into task stages using the predefined keypoints and success conditions, which supplies execution-progress labels without manual annotation. Difficulty is varied along factors such as number of repetitions, number and similarity of distractors, delay between the informative observation and the later decision, and number of stages.

SEEK itself adds three memories to a coarse-to-fine policy. Multi-camera RGB-D observations are rendered into three virtual views, and a shared encoder turns each view into an "interaction memory" tied to the predicted manipulation target. WCM keeps the latest K entries in order, dropping the oldest. PAM keeps everything that has fallen out of the window in an episode archive and retrieves the single most similar past entry by comparing pooled visual descriptors against the current one (excluding entries already in WCM). SCM holds a discrete counter that advances only when the policy predicts a stage boundary, plus a learned stage embedding. Current visual features attend to these memory slots, so the read budget stays fixed at at most K recent entries, one anchor, and one stage representation per view regardless of episode length. Training is behavior cloning on expert demonstrations with an action loss plus a weighted stage-boundary loss; at inference, memory and the counter update online from the policy's own predictions, and all memory resets between episodes.

Evaluation uses 100 training demonstrations and 25 held-out test episodes per HIDE task, with a single policy trained jointly across all 15 tasks. Baselines span VLA models (OpenVLA, OpenVLA-OFT, π0, π0.5, GR00T-N1.7), 3D policies (RVT, RVT2, SAM2Act), and memory-based methods (μVLA, MME, SAM2Act+), with SAM2Act as the backbone baseline.

Note: the provided paper content is truncated partway through Appendix A.1's task descriptions, so the remaining simulation task specifications and Appendices A.1 details beyond "search cup from cabinet" are not available here.

Why This Matters

Impact on research. The paper reframes memory in manipulation as recovering a specific decision-relevant latent state rather than simply retaining more history, and it supplies a benchmark that isolates that property by construction. The finding that individual memory mechanisms help some categories and can degrade others, while their combination wins, suggests memory design should be matched to the hidden-state requirement rather than treated as a single general-purpose capability.

Real-world applications.

  • Industrial assembly or packaging where the same motion must be repeated exactly N times and a missed or extra repetition is a defect.
  • Household and service robots performing multi-stage chores (loading a dishwasher, cleaning a surface, stacking and storing objects) where progress must be tracked across visually repetitive steps.
  • Warehouse and logistics picking, where the robot must remember which bins, drawers, or containers it has already checked during a search.
  • Inspection and maintenance, where an observation made early (which part was faulty, where an item was seen) is only needed much later after the scene has changed.

Industry relevance. Robot foundation models are widely deployed with limited temporal context; this work quantifies that gap with concrete baseline numbers and shows a lightweight memory layer can be added on top of an existing backbone (SAM2Act) while preserving standard-task and perturbation performance (84.7% on RLBench, 61.9% on The Colosseum perturbations). The 89% versus 47% real-world gap against SAM2Act+ is directly relevant to companies shipping manipulation policies for repeat-count and search-heavy workflows.

Future Directions

  • Extending HIDE beyond controlled, interpretable hidden states to richer real-world settings and broader sources of partial observability, as the authors explicitly propose.
  • Developing more general memory mechanisms than the WCM/PAM/SCM triple, and testing whether the combination rule generalizes to hidden-state types outside the three benchmark categories.
  • Investigating memory content and horizon more deeply, given that visual features and length 6 performed best here but repetition counting regressed at short memory length — the conditions governing these trade-offs are unresolved.
  • Scaling real-world evaluation beyond the four reported tasks (repeated button pressing, cup stacking, desk cleaning, hidden-piece search) to more environments and longer multi-stage procedures.

Target Audience

Robotics and embodied-AI researchers working on manipulation policies, memory architectures, or VLA models; benchmark designers interested in partial observability and controlled task construction; graduate students entering robot learning who need a clear formulation of hidden-state dependence; and applied engineers building repeat-count, search, or multi-stage manipulation systems who want evidence on which memory mechanisms to combine.

Authors’ abstract

Recent advances in robot learning have enabled manipulation policies to perform increasingly diverse tasks and generalize across environments. However, reliable execution often depends on hidden task states that cannot be determined from current observations alone, making interaction history essential. We introduce $HIDE$, a benchmark for evaluating manipulation memory under partial observability. HIDE comprises 15 tasks covering repetition counting, historical-state recall, and execution-progress tracking, with randomized initial configurations and decision points where similar observations require different actions depending on prior events. We further propose $SEEK$, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state. Evaluations reveal substantial limitations in existing policies on HIDE, while memory augmentation improves task success in both simulation and real-world experiments. Individual mechanisms benefit some tasks but can degrade others; their combination achieves the highest average success rate on HIDE among the evaluated configurations. These findings highlight the importance of maintaining internal representations of hidden task states and matching memory design to task-specific information requirements.

Read the original paper