Skip to content
AI.info

Research

InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation

Overview Research area: Robotics — humanoid whole-body loco-manipulation, motion retargeting, physics-based motion imitation, and iterative data generation. Technical level: Advanced. The paper assume

InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation
arXiv
2610.06850
Published
2026-10-05
Authors
Yucheng Zhang, Sirui Xu, Jinhong Li, Liuyu Bian, Anatulya Nandi, Derek Zhang, Xiangchen Liu, Xueting Li, Umar Iqbal, Yu-Xiong Wang, Liang-Yan Gui

AI summary

Overview

Research area: Robotics — humanoid whole-body loco-manipulation, motion retargeting, physics-based motion imitation, and iterative data generation.

Technical level: Advanced. The paper assumes familiarity with reinforcement learning for character/humanoid control, motion retargeting, contact modeling, and sim-to-real transfer, though its central idea (growing training data around verified demonstrations) is conceptually simple.

Scope: InterMimicGen is a self-evolving motion-imitation framework that consolidates human-object interaction captures, retargets them to six robot configurations, trains a single generalist tracking policy, and then iteratively edits and validates new motions so that the reference data and the tracker improve together.

What This Paper Is About

Captured human-object interactions are a rich source of supervision for humanoid robots, but they are sparse, heterogeneous, and not directly executable by a robot. The paper's goal is to turn a finite set of human motion-capture clips into a continually expanding collection of robot-executable motions, while simultaneously improving the tracking policy that runs them. The authors argue that a fixed reference set can never cover the feasible variations of a loco-manipulation task, because a small shift in posture or object placement can break a grasp or make the robot fall — so data must grow together with the policy.

Key Contributions

  1. A consolidated and retargeted interaction reference collection. The authors merge InterAct and the object-interaction portion of HiPHI into one collection of 16,059 motions, 140.68 reference hours, and 157 object entries, then retarget it into what they describe as, to their knowledge, the most diverse humanoid robot reference collection for dexterous whole-body loco-manipulation.

  2. A physics-based generalist tracker for dexterous loco-manipulation. A single policy trained on the whole collection executes the references in simulation on a humanoid with dexterous hands, at a scale and diversity beyond prior humanoid tracking systems for loco-manipulation.

  3. A self-evolving augmentation-learning loop. Each round makes small task-preserving edits to where an interaction takes place (object edits) and how the body performs it (body edits), fine-tunes the tracker, executes the candidates in physics, and keeps only the successful rollouts as seeds for the next round.

  4. Cross-embodiment and real-robot validation. The same formulation retargets to six robot configurations, and the resulting trajectories are executed on real robots, including box lifting and placement, suitcase pulling, tripod carrying, and chair relocation.

Main Findings

  • Retargeting beats a prior dexterous retargeter on motion quality. On G1 with Inspire hands over 3,059 OMOMO clips, the method achieves 98.2% frame-level and 98.8% hand-level contact versus 96.1% and 97.1% for Weave, reduces object penetration to 37.2% (from 67.4%) and penetration depth to 2.39 cm (from 2.62 cm), eliminates ground penetration (0.0% versus 31.9%), cuts foot sliding to 12.3% (from 62.0%), and lowers MPJPE to 5.74 cm (from 9.54 cm), at the cost of slightly higher in-hand slip (0.239 versus 0.225). The text reports that joint acceleration and joint-limit violations are also greatly reduced.

  • One generalist policy nearly matches per-object specialists on body tracking, but trails on object control. Over 15 objects grouped into bimanual, sitting, and grasping interactions, the generalist on G1 with Inspire hands reaches 73.2%, 64.9%, and 58.3% success on the three groups, versus 93.8%, 75.4%, and 84.9% for specialists. Body MPJPE stays close (5.86 versus 5.19 on bimanual), while object rotation error falls behind (23.4 versus 10.4 on bimanual; 35.5 versus 14.1 on grasping). Sharing one policy costs mainly object control rather than body tracking.

  • Task type drives difficulty. Bimanual interactions are tracked most reliably on every G1 configuration; sitting and grasping are harder because sitting shifts body weight onto the object and grasping requires finger contact on small object regions. The Booster K1 specialist (75.7% on bimanual) trails the G1 specialists.

  • The verified reference set grows more than a hundredfold over five rounds. Growth relative to the original references goes from 21.8× at round 1 to 150.5× at round 5 for Inspire bimanual, 21.3× to 142.0× for Dex3 bimanual, and 19.6× to 146.4× for Inspire grasping. The text describes roughly twentyfold growth in the first round and more than a hundredfold by the fifth.

  • Self-evolution makes the tracker better, not just the dataset bigger. The original tracker succeeds on only 52.1% (Inspire bimanual), 59.3% (Dex3 bimanual), and 64.2% (Inspire grasping) of accepted references — failing on more than a third of them — while the evolved tracker reaches 98.4%, 98.7%, and 98.9%. Both trackers execute 100% of original references, so no prior capability is lost.

  • Augmented motions stay close to the originals in quality. Averaged over all rounds, body joint acceleration and hand jitter stay near original levels, while foot sliding rises from about 1% to about 2% of contact frames. For Inspire bimanual, body acceleration goes from 11.4 to 12.4, foot sliding from 1.18 to 1.87, and hand jitter from 1.00 to 1.17.

  • Growth requires learning. A frozen tracker that only filters candidates without fine-tuning accepts almost none of the new candidates after the first round, because those candidates start from first-round successes that already lie beyond its reach.

  • Body edits widen coverage beyond a comparable augmentation baseline. The posture variation covered by body edits spreads over a region of body motion several times wider than the scene-scaling variants of ULTRA.

  • Trajectories transfer to real hardware without object feedback. The low-level controllers on each real platform are blind to the object and track only robot motion, yet the trajectories execute well on every platform tested.

Methodology in Plain English

Step 1 — Gather and clean human interaction data. The authors combine two large sources of motion-captured human-object interaction into one collection, converting everything to a common body-object-hand format. Missing hand annotations are treated as unknown rather than as evidence of no contact. Interactions that a given robot cannot physically support — for example, a bag hanging from an unmodeled strap, or finger-dependent grasps on a robot without dexterous hands — are removed for that platform.

Step 2 — Retarget human motion to robots in two stages. Retargeting has to satisfy constraints at two very different scales: the whole body must preserve support and body-object geometry over long horizons, while the hand must resolve millimeter-scale contact around small objects. So the authors split the problem. A coarse stage solves the whole body against an "interaction mesh" linking body landmarks to object geometry, with a finger-opposition term that pushes each finger to approach the object from the side opposite the thumb, so the hand actually wraps around the object instead of resting on one side. Several constrained solvers are tried in turn and the first feasible solution is kept. A cleanup stage then runs two passes of Adam optimization to remove penetrations, jitter, and contact drift while explicitly preserving hand-object contact. Each pass is kept only if it passes deviation and quality checks.

Step 3 — Train a generalist tracker. A single policy observes robot state, upcoming reference frames, and object-relative geometry, and outputs joint targets for PD control. It is trained with proximal policy optimization across parallel environments, with a reward balancing body motion, object motion, interaction geometry, and contact fidelity. Reference-state initialization exposes the policy to different phases of long sequences, and adaptive sampling emphasizes difficult motions.

Step 4 — Close the data flywheel. Each round takes verified references as parents and edits them in two ways: object edits move or turn the object's trajectory, and body edits change posture (deeper crouch, shifted pelvis, wider stance) while keeping the hands on their original targets. Every proposal must preserve four invariants of its parent: interaction type, intended contacts, object, and task outcome. The tracker is then fine-tuned on the verified references plus the proposals, and each proposal is rolled out in simulation. A proposal is accepted only if its rollout reaches the end of the reference without early termination or a fall, passes smoothness checks, and preserves the task outcome. Accepted rollouts seed the next round; failures are never edited again, so errors cannot accumulate.

Why This Matters

The paper reframes a tracking policy as both a consumer and a producer of data. Instead of collecting more demonstrations or hand-crafting variations, it uses imitation itself to repair and accept physically valid variations, which sidesteps the brittleness of separate kinematic motion synthesizers for object interaction. It also shows that a single generalist policy can serve as an interaction-data validator at a scale beyond prior humanoid tracking systems.

Real-world applications:

  • Warehouse and logistics robots that need to lift, carry, and place boxes of varying shapes and placements.
  • Home and service robots performing everyday interactions such as pulling a suitcase, carrying a tripod, or relocating furniture.
  • Humanoid platforms in industrial settings where a robot must hold contact with a tool or object while keeping balance.
  • Data-efficient training pipelines for robot vendors that want to reuse a single motion library across multiple hardware configurations.

Industry relevance: The framework is explicitly cross-embodiment — a new robot only needs to declare its own landmarks and hand mappings. That makes the pipeline attractive for companies with mixed fleets of humanoids and wheeled manipulators, and for teams that want to turn existing motion-capture assets into training data rather than collecting new robot demonstrations.

Future Directions

  • Scaling beyond five rounds. The paper reports five rounds of self-evolution; it does not report what happens at longer horizons or whether coverage eventually saturates.
  • Closing the gap to per-object specialists. The generalist trails specialists most on object rotation and grasping success, so recovering fine object control within a shared policy remains open.
  • Extending to interactions currently filtered out. Sequences simulation cannot represent, such as bag hanging with an unmodeled strap, are removed; the paper does not report how such deformable or

Authors’ abstract

Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.

Read the original paper