Skip to content
AI.info

Research

Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies

Summary: Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies Overview Research area: Robotics and embodied AI — the evaluation of a large reasoning model (GPT-6 Astra, arXiv:

Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies
arXiv
2609.38537
Published
2026-09-29
Authors
Galbot Team, Xuchuan Chen, Xiaoqian Cheng, Yu Deng, Lihe Ding, Shaocong Dong, Xiangjun Gao, Haozhe Jia, Zekai Li, Zhoujian Li, Yunrui Lian, Sikai Liang, Chenghuai Lin, Dairu Liu, Jiahang Liu, Qingtao Liu, Yuxuan Ma, Zekun Qi, Jiayi Su, He Wang, Ruochen Xu, Tianyu Xu, Xudong Xu, Zhe Xu, Mi Yan, Siming Yan, Li Yi, Ruixi Yu, Jinlu Zhang, Yintianrun Zhang, Zhikai Zhang, Zhizheng Zhang, Yixin Zheng, Weiyi Zhu

AI summary

Summary: Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies

Overview

  • Research area: Robotics and embodied AI — the evaluation of a large reasoning model (GPT-6 Astra, arXiv:2609.38537v1 [cs.RO], 29 Sep 2026, CC BY 4.0) as a robot control policy across manipulation, navigation, locomotion, and humanoid loco-manipulation.
  • Technical level: Advanced. The paper assumes familiarity with vision-language-action models (VLAs such as π0.5), inverse kinematics, operational-space and proportional–derivative control, whole-body controllers, and navigation metrics such as SR, SPL, nDTW, and sDTW.
  • Scope: A single-page (one-sentence) description: the paper systematically asks which tasks GPT-6 Astra can control directly, where it works better cooperating with learned policies or frozen whole-body controllers, where it fails, and what inference resources its decisions consume.

What This Paper Is About

GPT-6 Astra can emit numerical robot actions, not just high-level plans, so the authors set out to measure how far that ability actually goes as a general-purpose embodied policy. They evaluate it across six domains — gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation — separating Direct control (Astra alone over analytic control) from Hybrid control (Astra reviewing or supplying commands to a learned policy or frozen whole-body controller). The goal is to characterize not just successes but capability boundaries, failure modes, and the token and latency costs of each decision.

Key Contributions

  1. A six-domain evaluation framework with a shared control taxonomy. The paper defines an action as a command submitted to the robot's control interface, distinguishes Direct from Hybrid control, and specifies for each study what Astra decides, how commands become motion, and what feedback informs the next action (Figure 1, Table 1).
  2. Paired, matched-start comparisons against learned policies and published baselines. Comparisons are run on shared task instances, scene configurations, seeds, and initial-state fingerprints (RoboDojo, RoboLab, RoboCasa365, four navigation subsets, HumanoidBench, SIMPLE, PASSAGE, DexJoCo).
  3. Failure analysis and intervention accounting. The paper quantifies what fraction of executed control steps Astra actually generates or corrects — for example, 85.6% of Hybrid RoboDojo steps follow π0.5 and 14.4% are Astra-generated or corrected; Astra corrections are 11.98% of executed Hybrid steps in the ten-task dexterous benchmark; 73.6% of DexJoCo steps are unmodified policy actions and 26.4% are Astra corrections.
  4. Resource measurement separated from task performance. The paper reports model calls, token counts, and inference latency independently of robot motion — including 624.8 million tokens for policy-assisted control and 1.132 billion for direct control across 50 RoboDojo instances per condition, and a 30-second locomotion run requiring 250 model calls averaging 39.86 seconds each with physics paused during inference.

Main Findings

  • Gripper manipulation favors Hybrid on long-horizon bimanual tasks. On RoboDojo, Hybrid succeeds on 24 of 50 instances (48%) versus 13 of 50 (26%) for Direct, with mean Score 62.60 versus 37.81. Gains concentrate in sequence imitation (53 vs 0), tower construction (64 vs 12), clothes folding (100 vs 40), and bottle disposal (100 vs 36). Direct is higher on language-based classification (60 vs 38) and object classification (100 vs 71), so the Hybrid advantage is not uniform.
  • Reweighted published RoboDojo references sit below Hybrid. Reweighting published results to the evaluated ten-task and scene mixture gives π0.5 a success rate of 15.67% and a Score of 24.43; aggregate Scores are 31.82 (DM0.5), 38.26 (Galaxea), 33.97 (Xiaomi R1), and 32.73 (OpenWAM) versus 62.60 (Hybrid) and 37.81 (Direct).
  • RoboLab, a single-arm zero-shot benchmark, favors Direct. Direct succeeds on 49/50 trials (98%), with five successes on nine of the ten tasks; Hybrid succeeds on 46/50 (92%); π0.5, Cosmos3-Nano-Policy, and DreamZero succeed on 18/50 (36%), 18/50 (36%), and 17/50 (34%). Hybrid improves over π0.5 by 56 percentage points but is three successful trajectories below Direct.
  • Policy assistance is worth less when the motion prior fits the task poorly. The paper argues this explains the reversed benchmark rankings: RoboDojo needs bimanual, contact-sensitive coordination, while RoboLab is mostly single-arm semantic pick-and-place that Direct can already plan.
  • Dexterous manipulation improves with Hybrid but Direct lags badly. Across ten simulated Sharpa-hand tasks, Hybrid reaches mean Score 61.6, π0.5 44.2, and Direct 16.6. Hybrid beats Direct on all ten tasks and beats or ties π0.5 on all ten, with the largest gains at 40 points (mahjong tile storage), 32 (upright egg placement), and 28 (bottles/cans sorting); bread-slot insertion and nesting-doll ordering stay at 28 and 24.
  • Experience-guided DexJoCo bimanual disk stacking reaches 5/10. Astra reviewing a fixed π0.5 policy completes 5/10 trials in ten trials run as five successive pairs; the trials and five reviews consume 63.60 million recorded tokens including cached input, of which only 0.136 million are used for review.
  • Direct in-hand control struggles to sustain object motion. RL tracks the rotation target within tolerance for 76.90% of cylinder steps and 63.50% of cuboid steps, versus 0.51% and 4.40% for Astra. On paired translation, Astra succeeds in 1/5 cases versus 4/5 for RL, with mean terminal errors of 59.05 mm and 17.29 mm; on combined translation and rotation Astra succeeds in 0/5 cases versus 4/5 for RL at matched steps and 5/5 after 15 s, against thresholds of 22.4 mm position error and 10° rotation error.
  • The in-hand gap is not just about dropping objects. Both controllers retain the cuboid through all five horizons, yet Astra rotates too slowly. The paper concludes that preserving support is easier for Astra than reorganizing finger contacts to sustain motion.
  • Mobile manipulation shows uneven gains and lost policy successes. On RoboCasa365, Hybrid completes 29/75 (38.7%), Direct 25/75 (33.3%), and standalone π0.5 17/75 (22.7%). Hybrid exceeds Direct on atomic seen tasks (52.0% vs 28.0%) and composite seen tasks (28.0% vs 16.0%), but trails on composite unseen tasks (36.0% vs 56.0%). An exploratory exact McNemar test gives p = 0.5235; Hybrid alone succeeds on 13 paired starts and Direct alone on nine, with 16 shared successes.
  • Rewriting instructions is a large part of Hybrid mobile manipulation. 66 of 75 episodes use rewritten instructions, which condition 72.6% of proposals; 55.2% of executed steps are attributed to accepted policy actions and 44.8% to Astra. Conversely, only standalone π0.5 succeeds on CoffeeSetupMug/03, and both Astra conditions fail CoffeeSetupMug and WashLettuce while the policy completes three and two cases.
  • Navigation is Astra's clearest strength in local comparisons. Astra reaches 39/50 (SR 78, SPL 65.27, nDTW 72.20, sDTW 59.35) on R2R, 46/50 (92, 77.25, 84.73, 80.42) on English RxR, 28/50 (56, SPL 22.43) on MP3D ObjectNav, and 41/50 (82, SPL 43.69) on HM3D ObjectNav. Its SR and SPL exceed every locally evaluated released system on all subsets with reported results, with SR gains over LightNav-0 of 20, 20, 36, and 16 percentage points; LightNav-0 has slightly higher R2R nDTW (72.77% vs 72.20%) while Astra has higher sDTW (59.35% vs 48.89%).
  • Object search succeeds but is inefficient. SPL of 22.43% and 43.69% on the two ObjectNav subsets reveals substantial detours that success rate alone conceals; ten MP3D episodes and five HM3D episodes are budget-limited.
  • Dense motion-reference generation for locomotion is unreliable. None of five sequential Astra attempts on a single obstacle course reaches the goal, despite improvements in stability and forward progress. Each proposal must specify 0.5 s at 50 Hz — 1,625 consistent values covering root height, projected gravity, planar velocity, yaw rate, and positions and velocities for 29 joints.
  • Humanoid loco-manipulation shows positive but partial results. Astra exceeds baseline methods on 13 of 30 HumanoidBench tasks when using pretrained whole-body controllers.
  • Latency constrains practical control. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each, with physics paused during inference; policy-assisted and direct control consume 624.8 million and 1.132 billion tokens across 50 RoboDojo instances per condition.
  • The headline interpretation: there is a gap between useful task decisions and reliable physical control. Astra can set up conditions for a policy's next action without generating the full trajectory, but it cannot yet be trusted to coordinate contact changes or construct dense body motion.

Methodology in Plain English

The team treats GPT-6 Astra as a controller inside a closed loop: the model observes the scene, chooses a command, the robot moves, and the new observation feeds the

Authors’ abstract

GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities as general-purpose embodied policies, we conduct comprehensive evaluations across six domains, examining direct control, cooperation with learned policies, and feedback-driven adaptation. In gripper manipulation, Astra can correct task targets and prepare contact conditions for subsequent policy execution; hybrid control with π0.5 achieves 48% success on the evaluated RoboDojo subset. In dexterous manipulation, hybrid control achieves 50% success in ten experience-guided DexJoCo trials, while direct in-hand control struggles to coordinate finger contacts. In mobile manipulation, hybrid control reaches 38.7% success on the evaluated RoboCasa365. In navigation, Astra leads our local comparisons, reaching 92% success on RxR instruction following and 82% on HM3D object search, although search incurs substantial detours. In locomotion, dense motion-reference generation remains unreliable: none of five sequential attempts on a single obstacle course reaches the goal, despite improvements in stability and forward progress. In humanoid loco-manipulation, Astra exceeds baseline methods on 13 of 30 HumanoidBench tasks with pretrained whole-body controllers. These findings reveal a gap between useful task decisions and reliable physical control. Inference latency further constrains practical control: across 50 RoboDojo instances per condition, policy-assisted and direct control consume 624.8 million and 1.132 billion tokens. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each, with physics paused during inference.

Read the original paper