Skip to content
AI.info

Research

PhysEvo: Astra Can Act, Let It

Overview Research area: Robotics and embodied AI — specifically LLM/VLM-driven robot manipulation, agent harness design, and self-improving agents. Technical level: Advanced. The paper assumes familia

PhysEvo: Astra Can Act, Let It
arXiv
2610.08995
Published
2026-10-06
Authors
Wenqing Tian, Zeyu Zhang, Zhaocheng Liu, Fengwei Liu, Qiang Liu, Liang Wang

AI summary

Overview

Research area: Robotics and embodied AI — specifically LLM/VLM-driven robot manipulation, agent harness design, and self-improving agents.

Technical level: Advanced. The paper assumes familiarity with vision–language–action (VLA) policies, inverse kinematics, control interfaces, and the self-improvement literature (Reflexion, STOP, AutoHarness, and similar).

Scope: The paper introduces PhysEvo, a framework in which one frozen model (Astra) plays both a robot-executing task role and a self-revising meta role, accumulating persistent edits to its tools, skills, and diagnostic resources through physical execution feedback.

What This Paper Is About

Astra can already control a robot, but this does not mean it executes tasks reliably: the paper reports that direct Astra reaches 100% success on object classification yet succeeds on only 26% of attempts across a ten-task RoboDojo evaluation, rising to 48% when paired with the specialized VLA policy π_0.5. The paper argues that the bottleneck is not the model but the harness — the tools, skills, and execution logic sitting between the model and the robot — citing failures such as a held bottle blocking the wrist view during pouring and local inverse kinematics folding the arm into an obstructive posture under a tool-center-point interface. The goal is to let the same frozen model diagnose those harness failures after the fact and revise the harness itself, so that corrections persist into later attempts and later rounds of improvement.

Key Contributions

  1. A physical recursive self-improvement (RSI) framework. One frozen Astra serves as both task agent and meta-agent. Accepted revisions to the task harness change later execution; accepted revisions to the meta harness change how later failures are diagnosed and corrected, so the retained system includes both a way to act and a way to improve.

  2. An evolved interaction architecture. PhysEvo develops joint-level control (twelve absolute arm-joint targets and two continuous gripper commands), evidence-seeking and retrospective observation (free-arm active viewing plus review_frames retrieval of earlier images and aligned robot states), geometry tools, and reusable manipulation skills — with no model-weight updates and no separately trained action policy.

  3. A 42-task RoboDojo evaluation with mechanism-focused cases. Retained task-specific deployment versions score 68.14/100 and 62.00% success under equal five-dimension weighting, compared with 47.17% for RoboDawn's one-shot Astra agent, plus mechanism cases tracing execution feedback to specific control, observation, and diagnosis changes.

  4. A five-task real-world evaluation of continued adaptation. The simulation-evolved harness is deployed on an AgileX PiPER and revised further, reaching 90.60/100 mean score and 84.00% success over 25 trials.

Main Findings

  • Five-dimension RoboDojo result: PhysEvo reaches 68.14/100 score and 62.00% SR under equal dimension weighting, exceeding RoboDawn one-shot's 47.17% SR by 14.83 points. Its per-dimension scores/SR are: generalization 51.08/45.00%, precision 50.38/40.00%, long-horizon 56.38/45.00%, memory 100.00/100.00%, and open-ended 82.88/80.00%. Equal weighting of all 42 tasks instead gives 65.00 score and 58.57% SR.

  • Dimension-specific gains over RoboDawn: SR improves by 46.67 percentage points in memory, 17.50 in open-ended tasks, and 15.00 in precision.

  • Ten-task GPT-as-Policy comparison: PhysEvo reaches 68.00% SR versus 26.00% for direct Astra and 48.00% for the Astra–π_0.5 hybrid. PhysEvo improves SR over Direct on eight of the ten tasks and ties on two.

  • Eight tasks that challenge direct Astra: On the subset defined by VLA SR ≥ 20% and direct-Astra SR < 5%, PhysEvo achieves 55.00% SR and 62.88 score, versus official Astra's 1.25% and 5.96 — a 53.75-point SR gain. Its mean SR also exceeds Simate-beta's 50.83%, Liber-0 Preview's 46.44%, and RoboDawn's 45.00%. On the two pouring tasks, SR improves over RoboDawn from 0% to 80% (pouring balls) and from 80% to 100% (pouring liquid).

  • Combined subset: The union of the ten-task GPT-as-Policy set and the eight-task manipulation subset contains 16 distinct tasks and 80 attempts, giving 75.06 score and 65.00% SR.

  • Joint-level control beats TCP in a controlled comparison: Holding model, task skill, initialization per layout, and budget fixed across two layouts and six episodes, mean scores are 40 for TCP, 100 for Joint, and 55 for Both. Joint outperforming Both indicates that a larger tool menu alone is not sufficient; the benefit comes from direct authority over arm configuration.

  • Observation tooling accuracy: A synthetic-marker reconstruction test accepted 8 of 12 points, with median error 0.615 mm and 95th-percentile error 1.437 mm; four points were rejected by validity checks. A separate wrist-camera translation of approximately 57.525 mm yielded a marker error of about 2.14 mm.

  • Real-world AgileX PiPER results (five trials per task, 25 trials total): overall 90.60/100 score and 84.00% SR. Block stacking, bowl stacking, and pen placement each reach 100% SR; pouring water reaches 60% SR and writing "PhysEvo" reaches 60% SR with a score of 93.00. Two documented revisions compensate for liquid-stream forward momentum overshooting the cup and reverse the geometry-first localization sequence, using visual feedback for coarse approach and close-range geometry for fine positioning.

  • Mechanism cases: Retained experience changed the system in three identifiable ways — exposing posture as a policy decision (joint reconfiguration recovering from a rejected TCP motion), maintaining useful views during manipulation (a revised pouring skill keeps a free-arm side view through pouring and upright recovery), and inheriting diagnostic tools (a review tool built in an earlier round traced box displacement to wrist or gripper contact with raised flaps, leading to a packing-skill revision requiring clearance above the flaps before turning).

Methodology in Plain English

PhysEvo splits work between two roles that share one frozen model but use separate contexts and tool access. The task agent receives the instruction, selected skills, camera images, and robot proprioception, then picks a tool — either a motion tool that advances the environment or an observation tool that retrieves evidence or computes geometry without advancing physics. The harness links every requested action to its dispatched command and observed outcome, so a failed grasp can be examined through approach, gripper closure, and subsequent object motion rather than reduced to a final score.

After execution, the meta-agent reviews the trajectory alongside the current implementation, forms a failure hypothesis, and tests it through further inspection or diagnostic experiments. It can propose a revision to action tools, observation tools, execution logic, skills, or its own diagnostic utilities. When a tool changes, its skill guidance is revised with it, so the agent inherits both a capability and instructions for using it. Candidate revisions are tested during development — first checking that edited code runs and returns intended outputs, then comparing the candidate against the current harness on the same task and initial scene. A score gain, or progress beyond the failure without a score decrease, supports adoption; a score decrease leads to rejection. Diagnostic-tool revisions are checked against the specific error or measurement they target.

Accepted changes take effect between episodes, never mid-episode. Development resources are shared across tasks with an important restriction: at the task-agent level, only tools and task-agnostic skills are shared, while task-specific skills stay task-local. Each task then retains a deployment version that is fixed across its five scored layouts. Before each scored episode, session context, historical-frame indices, and writable memory are reset; online correction within an episode is allowed but persistent tools and skills do not change. Model weights, robot dynamics, and the task evaluator are never editable. On hardware, a separate adaptation protocol continues skill revision after deployment, with the meta-agent analyzing each attempt and revising skills for the next.

Why This Matters

The paper shifts attention from the model to the harness: the claim is that experience can improve not only the next action but the interface and diagnostic resources available to future actions. This matters for research because it suggests a route to higher aggregate manipulation performance without training a separate action policy or updating model weights, and because the meta-agent retains tools for improving itself, not just for acting.

Real-world applications:

  • Warehouse and logistics manipulation, where tasks such as packing objects and placing items must be executed reliably across varying layouts rather than succeeding only on familiar scenes.
  • Liquid handling in lab or kitchen settings, where the pouring revisions — maintaining a side view of the bottle mouth and cup rim, and compensating for stream momentum by shifting the bottle opening backward — directly address visibility and flow effects.
  • Tabletop assembly and sorting tasks, such as block stacking, bowl stacking, and tower building, where arm posture control determines whether a target pose is reachable without obstruction.
  • Post-deployment adaptation of existing robot fleets, where a simulation-evolved harness is used as a starting point and revised on the physical platform to account for hardware characteristics such as lower positioning accuracy.

Industry relevance: The reported 68.14/100 and 62.00% SR against 47.17% for RoboDawn one-shot, and 55.00% versus 1.25% on tasks that challenge direct Astra, indicate that interface-level revision can compete with and exceed specialist robot policies on selected task sets. Since no model weights are updated, the approach is compatible with keeping a single frozen foundation model in production while the surrounding tools and skills evolve.

Future Directions

  • Whether the same frozen-model, two-role division scales beyond 42 RoboDojo tasks and the five evaluated hardware tasks — the paper's real-world evidence covers one AgileX PiPER platform and five tasks.
  • How far retained diagnostic resources can compound. The packing case shows a review tool built in one round supporting a correction in a later round; the paper does not report how this inheritance behaves over longer improvement sequences or whether revisions can conflict.
  • Robustness of candidate-revision acceptance, which relies on comparing a candidate against the current harness on the same task and initial scene, accepting on score gain or progress without a score decrease. The paper does not report failure modes where this criterion adopts a revision that generalizes poorly.
  • Transfer of evolved harnesses across embodiments, since the paper demonstrates continued adaptation after deploying a simulation-evolved harness on physical hardware, but does not report transferring an evolved harness between different robot platforms.

Target Audience

Robotics and embodied-AI researchers working on LLM/VLM-driven manipulation, agent harness design, and self-improving agent systems; engineers deploying foundation-model controllers on real hardware who need reliability without retraining; and readers already familiar with VLA policies, control interfaces, and the self-improvement literature (Reflexion, STOP, AutoHarness, SHAPER, RHO, and similar) who want to see harness evolution applied end to end from simulation to a physical arm.

Authors’ abstract

Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.

Read the original paper