Research
Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation
Overview Research area: Robotics — zero-shot sim-to-real transfer for robot manipulation, combining agentic LLM/VLM (vision-language model) code generation with hierarchical skill learning. Technical

- arXiv
- 2610.02788
- Published
- 2026-10-02
- Authors
- Xincheng He, Siyu Ma, Chang Yu, Yunuo Chen, Yanjia Huang, Ying Nian Wu, Yin Yang, Chenfanfu Jiang
AI summary
Overview
- Research area: Robotics — zero-shot sim-to-real transfer for robot manipulation, combining agentic LLM/VLM (vision-language model) code generation with hierarchical skill learning.
- Technical level: Advanced. The paper assumes familiarity with sim-to-real transfer, domain randomization, hierarchical reinforcement learning, vision-language-action models, and multi-agent LLM pipelines.
- Scope: The paper presents Skill2Real, an agentic framework that learns two levels of reusable, code-based robot skills entirely in simulation through an asymmetric Proposer–Verifier–Governor loop, then freezes and deploys that hierarchy on a real UR5e without real-world task demonstrations or fine-tuning.
What This Paper Is About
Sim-to-real transfer for robots usually spends its learning budget on adapting low-level representations (lighting, texture, dynamics) rather than on acquiring task-level knowledge that a robot can reuse. Skill2Real instead treats learned robot knowledge as executable programs written over a shared code-based API that exists in both simulation and reality, so that privileged simulator information can supervise learning without ever entering the deployed policy. The goal is zero-shot transfer: learn local manipulation skills and long-horizon task compositions in simulation, freeze them, and run them on a physical robot on unseen tasks.
Key Contributions
- An agentic skill-learning framework with a shared code-based policy interface. Skill2Real learns executable robot programs in simulation and transfers frozen skill knowledge across domains using only public observations and API semantics at deployment, with no real-world task-policy adaptation.
- An asymmetric Proposer–Verifier–Governor (PVG) scheme. The Proposer acts on public observations and API returns, the Verifier converts privileged simulation evidence into publicly grounded diagnoses, and the Governor admits only updates supported by validation rollouts.
- A two-level hierarchical policy learning strategy. A Cerebellum learns reusable local manipulation skills (Stage I), and a Brain learns long-horizon compositions conditioned on the frozen Cerebellum library, including task-level planning and recovery (Stage II).
- Cross-model and cross-domain empirical validation. The framework is tested with multiple foundation VLM Proposers (GPT-6 Astra, GPT-5.6 Sol, Claude Opus 5, Opus 4.6) on LIBERO-90 to LIBERO-Pro Long transfer, seven Robosuite tasks, and four real UR5e tasks.
Main Findings
- Cerebellum learning transfers to unseen long-horizon tasks. On LIBERO-Pro Long, Cerebellum learning raises Overall success from 2.0% to 35.3% with Astra and from 0.5% to 29.3% with Opus 5, starting from frozen LIBERO-90 checkpoints.
- The Brain adds large gains on top of frozen local skills. With the C3 Cerebellum fixed, Brain learning raises success to 56.3% (Astra) and 49.0% (Opus 5), adding 21.0 and 19.8 percentage points respectively. At B4, instruction/position success is 76.0%/36.5% for Astra and 66.5%/31.5% for Opus, indicating that object placement remains the harder half of the task.
- Independent Robosuite libraries show the same two-stage pattern. Across seven single-arm and bimanual tasks, Cerebellum learning raises mean success from 16.6% to 45.6% for Sol and from 19.4% to 49.6% for Opus; Brain learning then reaches 85.1% and 89.4%, adding 39.6 and 39.9 percentage points and improving all seven tasks.
- Real-robot zero-shot transfer improves substantially over baselines. On a UR5e with a fixed Astra Proposer and API, full Skill2Real reaches 78.75% mean completion across R1–R4, versus 56.25% for Cerebellum-only, 37.50% for Brain-only, 27.50% with no learned skills, and 11.25% for CaP-Agent0. Adding the Brain improves R1–R3 by 15, 5, and 20 percentage points and is the only configuration to complete drawer manipulation (50%).
- Skill2Real leads on long-horizon generalization benchmarks. With Astra using the fixed Sol-trained C3/B3 hierarchy, Skill2Real attains the best Overall (55.3%) and Task (75.0%) success on LIBERO-Pro Long, exceeding Zetta's 51.5% and 63.0%. The same hierarchy reaches 35.3% Overall with Opus 4.6, above ASPIRE's 30.5%.
- Position perturbations favor an end-to-end VLA baseline. Zetta leads under position perturbations (40.0% versus 35.5%), driven by cases such as bowl-in-drawer and book-in-caddy, where an end-to-end vision-language-action policy provides tighter reactive control for narrow-space pick-and-place than the code-based skill interface.
- Both PVG roles matter. Ablating either role lowers Pro Long success throughout hierarchical learning: full Skill2Real reaches 56.3% Overall, versus 39.0% without the Verifier and 43.0% without the Governor, losses of 17.3 and 13.3 percentage points.
- Source-suite learning improves but is uneven across families. On LIBERO-90, Sol improves from 14.2% at C0 to 38.2% at C3 and 62.0% at B4; Opus improves from 18.7% to 43.8% and 73.3%, with Brain learning adding 23.8 and 29.6 percentage points. Gains are larger in caddy storage, basket/tray storage, and surface placement than in stacking. Opus reaches 100.0% in caddy storage at B4, up from 56.7% at C3, while its stacking score remains at 40.0%.
Methodology in Plain English
The central idea is to make learned robot knowledge executable and reusable by representing it as programs written over an API contract shared by the simulator and the real robot. Both environments expose the same set of documented operations and observations (𝒜 = {a_k}, with K operations), so a program written in simulation can run on the robot. Each backend provides its own perception, calibration, and low-level control; the policy only reasons through the shared interface.
Learning runs as a three-role loop, in which all three roles are large language models using the same underlying model with reasoning effort set to "medium":
- Proposer generates executable programs from the language goal, the history of public observations and API returns, the API documentation, and the current skill memories.
- Verifier inspects privileged simulator state after a rollout (a deterministic execution evaluator summarizes the physical outcome) and converts it into feedback that is constrained to be expressible through public observations and API semantics. Simulator-only quantities never reach the Proposer.
- Governor looks across repeated validation rollouts and admits a candidate skill update only when the evidence supports it. The paper notes that Governor-based runs evaluate each candidate over 25 validation rollouts using seeds disjoint from training.
Training is sequential and two-level. Stage I runs the shared PVG procedure from an empty memory on primitive-level tasks to produce a Cerebellum library of local manipulation skills, each combining natural-language guidance, an API-level code template, and a lightweight structured state description. Stage II starts a fresh memory, receives the frozen Cerebellum library as fixed context, and learns Brain skills that compose it into long-horizon task programs with planning and recovery. At deployment, only the Proposer, the shared interface, and the frozen memories remain; the Proposer grounds both memories in real observations.
Experimental setup: each benchmark uses three Cerebellum iterations (C1–C3) then four Brain iterations (B1–B4) with C3 frozen, with C0 as the pre-training baseline. Every iteration covers every task with five seeded training rollouts per task, giving 450 episodes per iteration and 3,150 across C1–B4 on LIBERO-90, and 35 per iteration and 245 total across the seven Robosuite tasks. LIBERO-90 uses a 7-DoF Franka Panda with a gripper and joint-position control at 20 Hz, a maximum episode horizon of 4,000 environment steps, and 800×512 RGB images from a fixed scene camera and a wrist-mounted camera, with top-down views requestable through the API; Robosuite images are rendered at 512×512. Pro Long evaluation uses 20 tasks and 10 seeds. Real-world tasks use a UR5e with a Pika gripper and scene and wrist RGB-D cameras, with 20 trials per method for each of four tasks (R1 pick-and-place, R2 attribute-based sorting, R3 equation assembly, R4 drawer manipulation), and no real-world demonstrations, task-policy fine-tuning, or online skill-memory updates.
Why This Matters
- Impact on research: The paper reframes sim-to-real transfer from adapting policies or representations to transferring validated executable task knowledge. It also demonstrates that the learning gains are orthogonal to the underlying foundation model: the same two-stage trend appears with GPT-6 Astra, GPT-5.6 Sol, Opus 5, and Opus 4.6, and the fixed Sol-trained hierarchy transfers to test-time Proposers it was not trained with.
- Real-world applications:
- Tabletop pick-and-place of household or lab objects into specified containers (R1).
- Attribute-based sorting of objects by color or semantic category into labeled destinations (R2).
- Educational or instructional assembly, such as arranging cube digits to form an equation (R3, demonstrated with 23 + 45 = 68).
- Multi-step articulated-container tasks such as opening a drawer, inserting an object, and closing it (R4, demonstrated with citrus retrieval).
- Industry relevance: The design keeps simulator-only state out of the deployed policy and freezes skill memories before deployment, which is attractive for settings where retraining on the physical robot is expensive or risky — warehouse logistics, manufacturing kitting and assembly, laboratory automation, and service robotics. The shared code-based interface also means the robot backend can be swapped without rewriting the task knowledge.
Future Directions
- Expose end-to-end policies as callable tools. The paper states that future work will extend the current pure code-as-policy interface by exposing end-to-end policies as callable tools, which would directly address cases such as bowl-in-drawer and book-in-caddy where an end-to-end VLA provided tighter reactive control.
- Reduce VLM inference latency for faster closed-loop execution.
- Isolate hierarchy effects from training budget. The paper acknowledges that stage comparisons track the combined effect of additional training and the second memory level rather than isolating hierarchy at a matched training budget.
- Measure variability across independent training runs. Reported success rates are point estimates for fixed libraries; repeated episodes assess a fixed library's performance but do not measure variability across independent training runs.
- Explain uneven performance across manipulation families. The aggregate gain does not imply equal improvement across manipulation types, with stacking improving far less than caddy or basket/tray storage.
Target Audience
Robotics and embodied-AI researchers working on sim-to-real transfer, hierarchical skill learning, and LLM-driven code-as-policy agents will benefit most, particularly those interested in using privileged simulation supervision without contaminating deployment. It is also relevant to engineers building generalist manipulation stacks who need to know where code-based skill interfaces beat end-to-end VLA policies (instruction-perturbed long-horizon tasks) and where they do not (tight position-perturbed pick-and-place). Readers without a background in robot learning, VLMs, or simulation benchmarking will find the method sections dense.
Authors’ abstract
Transferring robotic skills from simulation to reality requires task knowledge that remains usable across differences in perception, dynamics, and embodiment. We introduce Skill2Real, an agentic policy framework that learns executable skills through a shared application programming interface (API). A Proposer-Verifier-Governor (PVG) loop uses privileged simulation evidence to diagnose outcomes and validate updates, while keeping learned skills grounded in public observations and API semantics. The Cerebellum first acquires local manipulation skills; the Brain then learns task-level composition with the Cerebellum frozen. Both memories transfer to the real robot without task-policy fine-tuning or skill-memory updates. As GPT-5.6 Sol learns skills on LIBERO-90, evaluating each frozen checkpoint with GPT-6 Astra raises LIBERO-Pro Long success from 2.0% to 56.3%, without training on Pro Long. Independent Robosuite training reaches 85.1% and 89.4% mean success with Sol and Opus 5 across seven tasks, respectively. Frozen Sol-trained LIBERO-90 skills achieve 78.75% mean completion across four real-world manipulation tasks with Astra. Removing the Verifier or Governor during LIBERO-90 training lowers final Pro Long success by 17.3 and 13.3 percentage points, respectively. These results support learning and transferring a hierarchy of executable skills through a common robot interface.