Skip to content
AI.info

Research

Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

Overview Research area: Computer Vision and robot learning, at the intersection of Real2Sim2Real scene reconstruction, physics simulation, and agentic manipulation (LLM/VLM-driven robot control). Tech

Real2Gym: Building Gyms from Videos, Bringing Skills to Robots
arXiv
2609.37089
Published
2026-09-29
Authors
Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu, Bo Dai, Mulin Yu, Tao Lu

AI summary

Overview

  • Research area: Computer Vision and robot learning, at the intersection of Real2Sim2Real scene reconstruction, physics simulation, and agentic manipulation (LLM/VLM-driven robot control).
  • Technical level: Advanced. The paper assumes familiarity with 3D reconstruction (Gaussian Splatting, point-cloud backends such as MoGe-3 and π³), physics engines (MuJoCo, Blender), and coding-agent frameworks for robot control.
  • One-sentence scope: Real2Gym is a framework that turns human or robot demonstration videos into physically validated simulation "gyms" and then accumulates reusable manipulation skills inside them for transfer back to a real Franka robot.

What This Paper Is About

Building robot skills from real-world videos requires three things at once: a simulated scene that visually matches the video, physics that actually executes the demonstrated motion, and a way to learn from repeated attempts. Existing pipelines require heavy manual work (scene scanning, mesh repair, articulation modeling, physics parameterization), and the authors show that even strong multimodal baselines such as GPT-6 Astra produce reconstructions with inaccurate object dimensions, relative spatial layouts, and camera extrinsics, which cause interpenetration, missed contacts, and failed grasps. Real2Gym addresses this by coupling a physics-verified Real2Sim pipeline with an agent that distills execution feedback into reusable skills, without updating the underlying model weights.

Key Contributions

  1. A Real2Sim pipeline that builds executable digital twins from human and robot demonstrations through event-centered correction, native-physics validation, and task-conditioned augmentation with action-feasibility checks.
  2. An experience-driven manipulation agent that combines perception tools, executable operation stages, and feedback-driven skill extraction to support efficient interaction and skill reuse in both simulation and on physical robots.
  3. Evaluation across 24 reconstructed environments (12 from DROID, 12 from EgoDex) showing consistent gains in reconstruction fidelity, task success, and token efficiency.
  4. Real-robot validation of closed-loop Real2Sim2Real self-improvement on a 7-DoF Franka robot, including a task that failed zero-shot and succeeded after simulation-refined skill acquisition.

Main Findings

  • Reconstruction quality: On DROID, Real2Gym scores 66.25 content alignment, 70.83 viewpoint alignment, 83.17 action fidelity, and 80.88 simulation success score, versus 55.00 / 30.00 / 76.08 / 71.08 for GPT-6 Astra Medium and 50.00 / 24.17 / 63.83 / 48.96 for GPT-5.6 Sol xhigh. Viewpoint alignment improves by 40.83 points over GPT-6 Astra.
  • EgoDex reconstruction: Real2Gym reaches 75.08 / 73.75 / 86.25 / 85.76, versus 57.58 / 38.67 / 65.75 / 56.59 for GPT-6 Astra Medium and 55.42 / 34.17 / 54.83 / 46.77 for GPT-5.6 Sol xhigh. The simulation success score of 85.76 is reported as a 51.5% relative improvement over the same baseline.
  • Agent success rate: Averaged across the two datasets, the agent with accumulated skills achieves 87.5% task success versus 70.8% for GPT-6 Astra Direct Mode, with approximately 74.9% fewer policy-execution tokens.
  • Per-dataset policy results (DROID): Ours (w/ skills) reaches 91.67% success rate with 13.33 responses, 0.51M tokens, and 6.32 min, versus 75.00% / 28.33 / 1.42M / 7.81 min for GPT-6 Astra Medium and 58.33% / 71.42 / 8.57M / 25.72 min for GPT-5.6 Sol xhigh.
  • Per-dataset policy results (EgoDex): Ours (w/ skills) reaches 83.33% success rate with 14.25 responses, 0.49M tokens, and 7.05 min, versus 66.67% / 33.75 / 2.56M / 10.31 min for GPT-6 Astra Medium and 41.67% / 105.50 / 13.63M / 28.73 min for GPT-5.6 Sol xhigh.
  • Before skill accumulation: Success averages 79.2%, with 67.3% fewer tokens than GPT-6 Astra.
  • Decision granularity drives efficiency: The GPT-6 Astra baseline makes individual-action decisions, whereas Real2Gym generates code for a complete operation stage, so each response can coordinate multiple actions and verification checks.
  • Real-robot zero-shot results on a Franka Emika Research 3 with a Robotiq 2F-85 gripper and two Intel RealSense RGB-D cameras: Pick Block into Plate reaches 100% SR (3/3) versus 33% (1/3); Hang the Mug reaches 100% (3/3) for both, but faster for Real2Gym at 1,302.7 s versus 1,522.7 s total execution time; Place Cup in Microwave and Close Door reaches 100% (3/3) versus 33% (1/3); Place Plate on Shelf reaches 0% for both methods zero-shot.
  • Failure-driven self-evolution: For Place Plate on Shelf, a digital twin was reconstructed from an uncalibrated third-person human video; the agent failed its first two attempts and acquired a stable skill by the third evolution iteration, reaching 100% simulation success rate and 100% success when deployed back to the physical Franka.
  • Emphasis on physical validity: Validated seed environments are expanded across object geometry, pose, support height, material properties, distractors, background, and lighting, with each candidate executed under native physics to verify task completion, contact dynamics, support stability, and Blender–MuJoCo cross-renderer consistency.

Methodology in Plain English

Building the gym. Given a demonstration video and a robot description file (URDF), the system reconstructs an editable 3D scene. Geometry is initialized from the first frame with MoGe-3 for single-view inputs or the Pi3X implementation of π³ across multiple views, with calibrated cameras and metric cues fixing global scale. The agent uses semantic reasoning plus SAM2 segmentation to identify manipulated objects, supporting surfaces, and background entities, completing occluded surfaces using RGB silhouettes and up to ten sparse multi-view frames. The robot is imported and aligned to the observed base pose, initial joint configuration, and camera mount, with iterative reprojection checks refining object geometry, poses, and cameras.

Validating the action. Interactions are recovered via keyframes tied to topological changes in contact, grasp, support, and containment. Robot joint and gripper states are used directly where recorded; human hand–object interactions are retargeted onto the target robot. At each keyframe the agent self-inspects and corrects across five diagnostic dimensions: primary discrepancy, camera alignment, relative object placement, contact/penetration, and appearance fidelity. The scene is then instantiated in MuJoCo with collision geometry, container cavities, material properties, and actuators, and physical execution is calibrated progressively from approach trajectories and contact orientation through gripper closure, grasp retention, support stability, and object release. The native physics trajectory is reimported into Blender for frame-matched comparison across Real RGB, Blender RGB, and MuJoCo RGB.

Scaling the gym. From each validated seed environment, the system synthesizes variations along geometry, pose, support height, materials, distractors, background, and lighting, propagating geometric changes across visual models, collision meshes, supporting surfaces, and robot mounting constraints, and adapting manipulation stages to new grasp regions and clearances. Failed candidates route back for refinement; validated ones join the environment pool.

Executing and learning. At each decision step the policy receives the task instruction, current observations, within-episode history, and selected skills, and outputs an executable Python program for a manipulation subtask, with entry conditions, intermediate checks, and an observable completion condition. Programs combine perception calls (SAM3 for segmentation, GraspNet for grasp proposals), simple numerical computations such as coordinate transformations and target-pose offsets, and robot commands such as goto_pose(), open_gripper(), and close_gripper(). After each episode, a separate extraction process labels every code round as having a positive, negative, or unknown local effect and distills skills containing applicability conditions, task procedures, effect checks, object-relative motion rules, and recovery guidance. Motion rules specify an acting entity, an object anchor, a relative position or orientation, and a stopping condition, and are instantiated at run time from currently observed object poses. The skill library is updated while the model parameters stay fixed.

Experimental protocol. Twelve DROID and twelve EgoDex scenes (four easy, four medium, four hard each) were reconstructed, giving 24 MuJoCo evaluation environments. Both baselines and Real2Gym were denied privileged simulator information such as ground-truth object poses or trajectories. The budget was 50 model decisions per task and 600 control steps per code execution. All scene construction, policy execution, and skill extraction used GPT-6 Astra with medium reasoning effort. Tools: Blender 4.5.3 LTS, MuJoCo 3.3.7, SAM3 0.1.0, and the PyTorch implementation of Contact-GraspNet.

Why This Matters

The work targets the cost bottleneck of robot self-improvement: physical resets are time-consuming, failed attempts risk damaging hardware, and parallel exploration is limited by the number of physical setups. By making simulation environments that are both visually aligned and physically executable, and by letting the agent accumulate skills without weight updates, the paper argues that human manipulation experience captured on video can be converted into reusable robot capability.

Real-world applications suggested by the tasks and datasets:

  • Tabletop pick-and-place and kitchen manipulation, including bowl stacking and cabinet manipulation shown in the qualitative comparisons.
  • Articulated-object and long-horizon household tasks, such as placing a cup in a microwave and closing the door.
  • Precision insertion and narrow-clearance placement, such as fitting a plate into a shelf frame.
  • Bimanual assembly from egocentric human video, such as placing wheels and a square nut onto a bolt, tightening the nut, and inserting the assembly into a base (used in the Appendix D ablation).

Industry relevance: the 74.9% reduction in policy-execution tokens and lower response counts point to cheaper agent operation; the use of off-the-shelf tools (Blender, MuJoCo, SAM2/SAM3, GraspNet) lowers the barrier for teams building simulation-backed robot learning pipelines; the Real2Sim2Real loop is directly relevant to companies generating training data and evaluation suites from existing demonstration datasets such as DROID and EgoDex.

Future Directions

  • Cross-view registration and appearance refinement. Multi-view misalignment between wrist-mounted and static cameras can produce ghosting or layered surfaces in noisy Pi3X point clouds, leaving residual geometric and textural discrepancies from the source video.
  • Adaptive control granularity. Coarse stage-level code generation is efficient but lacks the local fine-grained reactivity needed for highly constrained manipulation; the authors propose switching adaptively between stage-level and fine action-level control.
  • Broader embodiment coverage. Real-world experiments are confined to physical Franka arms with parallel-jaw grippers; extending to dexterous hands and diverse humanoid platforms is described as critical for assessing practical generality.
  • Remaining task failures. Even with skills, several tasks in Table 4 remain unsolved, and the DROID and EgoDex zero-shot success rates leave room for improvement, particularly on the hardest scenes.

Target Audience

Robotics and embodied-AI researchers working on Real2Sim2Real transfer, manipulation policy learning, and agentic code-generation for robot control; computer vision researchers interested in 3D reconstruction quality metrics tied to physical execution; and robot learning engineers who need to build interactive simulation environments and skill libraries from existing demonstration datasets without large-scale physical trial and error.

Authors’ abstract

Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.

Read the original paper