Skip to content
AI.info

Research

StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

Overview Research area: Robot learning, specifically online reinforcement learning (RL) post-training for vision-language-action (VLA) models on long-horizon manipulation tasks. Technical level: Advan

StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
arXiv
2609.36352
Published
2026-09-28
Authors
Ziyi Yin, Sangmin Woo, Kang Zhou, Sungyeon Kim, Aosong Feng, Haibo Ding, Jun Huan

AI summary

Overview

Research area: Robot learning, specifically online reinforcement learning (RL) post-training for vision-language-action (VLA) models on long-horizon manipulation tasks.

Technical level: Advanced. The paper assumes familiarity with VLA architectures (flow-matching policies such as GR00T-N1.5 and π0.5), PPO and GRPO, and action-chunk rollouts.

Scope: The paper proposes StructRL, an online RL framework that replaces terminal-only rewards with verifiable, dependency-structured intermediate rewards, and evaluates it on RoboCasa365 and LIBERO-Long with two VLA backbones.

What This Paper Is About

VLA models handle short manipulation skills but struggle when a single command requires a sequence of dependent manipulations, such as opening a box, adding several objects, and closing it. Online RL can improve such policies, yet most existing methods issue reward only when the entire task succeeds, which is sparse and cannot distinguish an early failure from a rollout that completed nearly everything. StructRL addresses this by decomposing each command into verifiable subtasks, arranging them into a dependency structure, and granting paced intermediate rewards only when prerequisite subtasks have been completed.

Key Contributions

  1. A structured intermediate reward framework. StructRL decomposes each long-horizon command into subtasks that pass a verifiability check (a binary environment-state criterion) and a progress check (the completion indicates progress toward the goal), then organizes the retained subtasks into ordered dependency groups with prerequisite sets X(v_i).

  2. Two reward-shaping mechanisms. Structure-aware reward gating decides whether a detected completion is eligible for credit (only after all prerequisites in X(v_i) are satisfied, and at most once per subtask), while dynamic reward pacing scales the reward by completion pace relative to demonstration durations, λ_i = β · 1/(1 + T(v_i)/T_d(v_i)).

  3. An automatic decomposition and grounding pipeline with no learned reward model at rollout time. Claude Opus 4.8 generates the subtask decomposition and dependency structure once per task; each subtask is grounded to a binary simulator predicate, and decompositions are fixed and reused during training, so no LLM or reward-model inference is needed during rollout collection.

  4. Broad empirical validation. Results across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and π0.5, plus a reward-source comparison against the Robometer learned progress model, a reward-component ablation, a decomposition-density sweep, a GRPO variant, zero-shot evaluations, and shorter-horizon LIBERO suites.

Main Findings

  • StructRL outperforms evaluated online RL baselines on both benchmarks and backbones. On RoboCasa365 with GR00T-N1.5, StructRL reaches 49.1% success rate (SR), compared with 38.6% for SFT and 41.5% for the strongest evaluated online RL baseline, a gain of 7.6 percentage points. On LIBERO-Long with GR00T-N1.5, it reaches 96.6% versus 92.4% for that baseline, a gain of 4.2 points. With π0.5, improvements are 3.9 points on RoboCasa365 (45.8% vs 41.9%) and 2.2 points on LIBERO-Long (96.2% vs 94.0%).

  • Gains appear in every horizon bucket. StructRL exceeds the strongest online RL baseline in every column of the bucket-level table. The largest positive gains are 7.0 percentage points on RoboCasa365 (1400–2900 bucket) and 6.6 points on LIBERO-Long (340–400 bucket), both with GR00T-N1.5.

  • Simulator-verified structured rewards beat a learned progress model under matched training. Against the Robometer vision-language reward model (built on Qwen3-VL-4B and trained on 1M trajectories), StructRL is higher by 2.4 points with GR00T-N1.5 and 3.2 points with π0.5 on LIBERO-Long under the same SFT initialization, PPO optimizer, interaction budget, and reward scale.

  • Verified subtask rewards provide most of the ablation gain. Adding components one at a time to a terminal binary reward with GR00T-N1.5 raises overall SR from 41.3% to 47.4% on RoboCasa365 and from 91.2% to 94.7% on LIBERO-Long (subtask rewards); dynamic pacing adds 0.5 and 1.2 points; structure-aware gating adds a further 1.3 and 0.5 points, reaching 49.2% and 96.4%. The largest end-to-end gain is on the longest RoboCasa365 tasks, from 26.2% to 36.5% (10.3 points).

  • Decomposition density has a sweet spot. In a controlled sweep on the 16 RoboCasa365 tasks with GR00T-N1.5, raising the average number of completion signals per task from 0 to 1 lifts SR from 40.2% to 45.9%, performance peaks at 50.4% at 5 signals, and then declines; at 52 signals SR falls to 37.6%, below the terminal-only configuration.

  • The structured reward is compatible with GRPO but PPO remains stronger. StructRL-GRPO exceeds SimpleVLA-RL in every horizon bucket, improving overall SR by 3.4 points with GR00T-N1.5 (44.9%) and 1.3 points with π0.5 (43.2%). The PPO variant is stronger overall by 4.2 and 2.6 points on the two backbones respectively.

  • Zero-shot transfer to unseen compositions is weak; forgetting is not observed on atomic tasks. On 16 held-out Composite-Unseen tasks, all policies remain low, with StructRL at 4.8% versus 4.3% for the strongest non-StructRL comparison. On 65 Atomic-Seen tasks, StructRL reaches 20.5%, versus 17.0% for the SFT policy and 16.9% for SimpleVLA-RL.

  • Modest but consistent gains on shorter-horizon LIBERO suites. With GR00T-N1.5, StructRL improves over SimpleVLA-RL from 91.2% to 92.1% on Spatial, 98.5% to 99.4% on Object, and 94.8% to 95.6% on Goal, averaging 95.7% versus 94.7%.

Methodology in Plain English

StructRL starts after supervised fine-tuning. For each task command, an LLM (Claude Opus 4.8) proposes candidate subtasks and keeps only those that satisfy two checks: completion can be detected by a binary environment-state criterion, and completing them actually indicates progress toward the goal. Candidates like "Approach Box" fail the first check, and "Open Gripper" fails the second. The LLM then groups the retained subtasks into stages that must occur in order, while subtasks inside a stage may be done in any order, producing a prerequisite set for each subtask. This structure is generated once, then frozen.

Before training, each retained subtask is deterministically mapped to a binary predicate over the simulator state using a rule-based procedure (the verb phrase determines the predicate type, its arguments identify the objects or fixtures). Reference durations for each subtask are precomputed from the SFT demonstrations, measured from when all prerequisites first became complete to when the subtask is completed, averaged over demonstrations where both events are observed.

During online RL, the policy rolls out action chunks. When a subtask completion is detected and all prerequisites have been completed and the subtask has not already been rewarded, the chunk that triggered it receives a reward of β · 1/(1 + T(v_i)/T_d(v_i)), where T(v_i) is the number of chunks elapsed since the previous rewarded completion. A fixed terminal reward λ_c is added when the whole task first succeeds. These chunk-level rewards drive PPO (the default), or GRPO, where returns are summed per rollout and group-normalized. Defaults are β = 0.6 and λ_c = 2.0, with chunk length L = 16 for GR00T-N1.5 and L = 10 for π0.5. All online RL methods train for 100 iterations under the same environment-interaction budget, with 512 rollouts per iteration on RoboCasa365 and 768 on LIBERO-Long, using 256 parallel environments. Training uses four nodes with eight NVIDIA A100 GPUs per node and takes approximately 48 hours on RoboCasa365 or 24 hours on LIBERO-Long.

Why This Matters

Impact on research. The paper reframes dense reward design for VLA RL as a structure problem rather than purely an estimation problem: rather than learning a scalar progress model from pixels, it derives rewards from simulator-verifiable events and uses task dependencies to decide when those events count. It also provides a matched-protocol comparison against a learned reward model and shows that excessive decomposition density can hurt, which is a caution for anyone assuming more intermediate signals are always better.

Real-world applications (from the task domains the paper evaluates):

  • Household and kitchen manipulation, including composing several pick-and-place, container, and fixture operations into one instruction.
  • Mobile manipulation where a robot must coordinate arm manipulation and base navigation across a kitchen.
  • Multi-step food packing and storage style tasks, such as sorting items into a container and moving them to a target location.
  • Tabletop manipulation suites where existing success rates are already high but small consistent gains are still available.

Industry relevance. The method is implemented in the RLinf-VLA framework and applied to released checkpoints (GR00T-N1.5, π0.5), with decompositions released as static configuration files so experiments can be reproduced without querying an LLM. Because intermediate credit is derived from simulator state rather than a large reward model, no reward-model inference is required during rollout collection, which matters for training cost. The paper also notes the compute footprint (about 1,536 and 768 GPU-hours respectively) and warns that reward shaping can induce unintended behavior, such as policies collecting intermediate rewards and then idling until timeout, so deployment on physical robots would require separate safety validation.

Future Directions

  • Improving zero-shot transfer to unseen task compositions. Composite-Unseen SR remains low for all evaluated policies (StructRL at 4.8%), so generalizing the structure learned on training tasks to new compositions is unresolved.

  • Choosing decomposition density automatically. The density sweep shows a peak at 5 signals and degradation at 52; the paper reports that the decline reflects earlier, less task-aligned completion signals and growth in total intermediate reward, but does not provide a principled selection rule.

  • Reducing or eliminating reward-induced failure modes. The ethics statement notes that policies can collect intermediate rewards and then remain idle until timeout, and that dynamic pacing discourages but does not eliminate this.

  • Isolating why composite RL does not degrade Atomic-Seen performance. The paper reports that StructRL improves Atomic-Seen SR over SFT and offers the hypothesis that intermediate completion rewards reinforce skills shared across atomic and composite tasks, but states that the table does not isolate that mechanism.

Target Audience

Robotics and embodied-AI researchers working on VLA post-training and reinforcement learning; practitioners who need reward design for long-horizon manipulation without training a separate reward model; and engineers building robot training pipelines who care about simulator-verifiable supervision, PPO versus GRPO tradeoffs, and the compute cost of large-scale online RL. Readers unfamiliar with flow-matching policies, action chunks, or PPO clipping will find parts of the methodology and tables demanding.

Authors’ abstract

Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.

Read the original paper