Research
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Overview Research area: Robot learning for dexterous manipulation — specifically learning manipulation skills from human hand-object interaction (HOI) data and transferring them zero-shot to real robo

- arXiv
- 2609.28660
- Published
- 2026-09-23
- Authors
- Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
AI summary
Overview
Research area: Robot learning for dexterous manipulation — specifically learning manipulation skills from human hand-object interaction (HOI) data and transferring them zero-shot to real robot hands.
Technical level: Advanced (requires familiarity with hand modeling, optimization, reinforcement learning, and sim-to-real transfer).
Scope: The paper presents Morphometric Imitation, a three-stage pipeline that retargets reconstructed human hand-object motion onto three-, four-, and five-fingered robot hands, refines it into dynamically feasible demonstrations with residual RL, and distills it into zero-shot sim-to-real visuomotor policies.
What This Paper Is About
Human hand-object interaction data is an attractive, low-cost source of manipulation demonstrations, but human and robot hands differ in morphology and size, natural human motion is not always dynamically feasible for a robot, and policies trained in simulation must survive the sim-to-real gap. The paper's goal is a pipeline that carries a reconstructed human trajectory all the way from kinematic retargeting through physics-based refinement to a real-world visuomotor controller, while preserving the demonstrated hand-object contacts at every stage.
Key Contributions
- Morphometric Optimization (MMO) — a kinematic retargeting method that builds a Scaled MANO model (scaling vector in R^6, scaling the palm and each finger independently), optimizes it once per robot hand to match the robot's morphology, then re-optimizes it per frame to recover the demonstrated hand-object contacts, and finally transfers it to the robot via linear blend retargeting plus inverse kinematics. It also handles coupled fingers (multiple human fingers mapped to one robot finger) and table penetration.
- A residual RL dynamic retargeting formulation in which the policy observes object pose and contact information (observed and reference object poses, current and future reference joint configurations, observed and reference contact states, robot-table clearance, object properties, table friction, previous action) and is rewarded by a product of an object-tracking term and a contact term. This makes object motion and the contact behavior that produces it complementary signals rather than a single tracking objective.
- A demonstration that the resulting teacher policies distill into visuomotor policies that deploy zero-shot, reaching 89.3% success in 300 real-world trials on 30 objects across 10 categories on the Sharpa hand.
- A comparison against prior retargeting methods and prior learning-from-human-motion pipelines, showing MMO improves contact F1 over five baselines by 8 to 28 points and raises downstream dynamic retargeting success by as much as 35 points, while doing so with the same hyperparameters across all three hands (except the correspondence mapping, which depends on finger count).
Main Findings
- Contact preservation: Across three robot hands (Dex3, Allegro, Sharpa) and ten GRAB hand-object trajectories, MMO outperforms five baselines in contact F1, improving on the strongest ones by 8 to 28 points, with consistently lower mean patch distance.
- Better kinematic references help downstream RL: MMO's kinematic retargeting improves the success rate of downstream dynamic retargeting by as much as 35 points over the strongest baseline, with more accurate object trajectory tracking and final grasps closer to the demonstrated contacts on average across the three hands.
- Object pose and contact information are complementary: Ablations show both play complementary roles in dynamic retargeting.
- Residual RL transfers across embodiments without tuning: One residual RL policy is trained per object category using the same hyperparameters across all hands and trajectories, and rollouts are shown for Dex3, Allegro, and Sharpa across hammer, apple, and cube categories.
- Zero-shot sim-to-real: On the Sharpa hand, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects spanning 10 categories, across diverse initial poses.
- Qualitative policy behaviors: The paper shows zero-shot rollouts for six objects — flashlight, cube, apple, hammer, lightbulb, and bowl — where the flashlight trajectory performs in-hand reorientation to point it forward and the hammer trajectory reorients the hammer head toward the table before lifting.
- Speed: MMO retargets 180 frames per second on an NVIDIA RTX 4090 GPU, including the preprocessing that computes hand-object contact indices and excluding the IK step.
- Training cost: Residual RL training takes 60–90 minutes per policy on an NVIDIA RTX 4090.
Methodology in Plain English
The pipeline has three sequential stages.
Stage 1 — Kinematic retargeting with Morphometric Optimization. The MANO hand model is limited to natural human hand-shape variation and cannot represent the large, part-specific proportion changes between a human hand and a robot hand. The authors extend it to a Scaled MANO model that scales the palm and each finger independently, initializing those scales from the robot's URDF (palm scale from wrist-to-MCP distances, per-finger scale from MCP-to-fingertip distances). Morphology matching then aligns corresponding joints and fingertips of the scaled MANO hand to the robot hand by minimizing a weighted joint and fingertip loss with Levenberg–Marquardt, producing a morphology-aligned hand once per robot hand. Because that morphological change can shift contacts, contact matching re-optimizes the hand's pose at every frame: it detects contact vertices (hand vertices within 4.5 mm of the object, the threshold recommended for GRAB), re-poses the morphology-aligned hand to hit those target contact positions, preserves the relative configuration of coupled fingers, penalizes vertices below the table, and regularizes toward the reference pose. Non-contact frames inherit the contact indices of the most-contacted frame so the pre-grasp is preserved continuously. Finally, robot pose recovery uses linear blend retargeting — computing blend weights on the 1D skeleton with a heat-diffusion method adapted via a 1D cotangent Laplacian — to produce dense Cartesian targets for every robot joint and fingertip, which directly define full 6-DoF link pose targets. Inverse kinematics (PyRoki) then solves for arm and hand joints jointly, with costs for position, orientation, self-collision, joint limits, joint velocities, and hand-table collisions; a JAX-optimized collision check samples link surfaces because capsule approximations with fixed-weight penalties were insufficient for hand-table interactions.
Stage 2 — Dynamic retargeting with residual RL. In ManiSkill, a PPO policy takes the kinematic reference as a warm start and outputs a residual action so the commanded configuration is the reference plus the residual. The reward is the product of an object-tracking term using the average point distance (ADD) metric — capturing both translational and rotational deviation — and a contact term that compares the number of contacting robot fingers to a goal of the minimum of the human reference's contacting fingers and the robot's finger count, normalized by finger count to account for embodiment. Episodes terminate early on sustained object-tracking error or contact mismatch (measured with exponential moving averages) and immediately on robot-table collision or excessive hand-object contact force. Domain randomization covers PD gains, hand and object friction, object mass, center of mass, anisotropic scale, table friction, initial object position within a 10 cm × 10 cm region, yaw within ±15 degrees, and the starting timestep of the reference trajectory, plus random external wrenches on the object.
Stage 3 — Visuomotor distillation. Teacher rollouts (collected under the same randomizations but starting at the beginning of the reference trajectory) are recorded and distilled into a visuomotor student policy for zero-shot sim-to-real deployment.
Why This Matters
The work argues that morphology- and contact-aware retargeting is what makes both kinematic and dynamic retargeting produce physically feasible, safe robot demonstrations suitable for sim-to-real visuomotor training — a departure from approaches that constrain human demonstrations to robot-friendly workspaces or rely on teleoperated fine-tuning. It also treats robot-table collision avoidance explicitly across both retargeting stages, which matters when transferring natural tabletop human motion. Table I positions Morphometric Imitation as the only listed method that combines kinematic retargeting, dynamic retargeting, zero-shot sim-to-real, a visuomotor policy, and the use of natural (unconstrained) human motion data.
Real-world applications:
- Household and service robots that need to manipulate everyday objects such as flashlights, hammers, bowls, apples, lightbulbs, and cubes across 10 object categories and varied initial poses.
- Multi-embodiment robot platforms, since the method is demonstrated on three hands (Dex3, Allegro, Sharpa) with three, four, and five fingers under shared hyperparameters.
- Scalable data pipelines that convert existing human interaction datasets (here GRAB) into robot training data instead of relying on costly teleoperation.
- Dexterous in-hand manipulation, illustrated by the flashlight reorientation and hammer reorientation behaviors.
Industry relevance: the approach targets reduced reliance on expensive teleoperated demonstration collection, reuse of human motion data at scale, and zero-shot deployment, all relevant to robot manufacturers and automation groups working with different hand embodiments.
Future Directions
- Extending the method beyond the rigid object categories used here; the authors note that concurrent work DemoMimic targets articulated box trajectories instead, and that REGRIND and TopoRetarget had not released code at the time of writing, leaving room for standardized comparison.
- Scaling beyond the ten GRAB trajectories and three hands to a wider range of natural human motion, since the paper emphasizes unconstrained motion rather than demonstrations collected under robot-friendly constraints.
- Investigating the boundaries of zero-shot transfer: the paper reports 89.3% success but does not report the failure modes or the remaining gap in the truncated content, so characterizing them is an open question.
- The visuomotor policy section is truncated in the provided content, so details of the student policy architecture, observation space, and demonstration volume are not reported here.
Target Audience
Robotics researchers and graduate students working on dexterous manipulation, learning from human motion data, motion retargeting, reinforcement learning for manipulation, and sim-to-real transfer. It is also relevant to engineers building multi-embodiment manipulation systems who need retargeting that preserves contact, and to readers already familiar with MANO, inverse kinematics, PPO, and domain randomization, given the paper's advanced formulation.
Authors’ abstract
Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. Second, residual reinforcement learning (RL) refines the kinematic reference using object pose and contact information from the human motion to produce dynamically feasible robot demonstrations. Third, these demonstrations are distilled into visuomotor policies. Across three robot hands and ten HOIs, MMO improves contact F1 over the strongest of five baselines by at least 8 points for every hand, while also improving the success rate of downstream dynamic retargeting by as much as 35 points. Ablations on the residual RL show complementary benefits from using object pose and contact information. Finally, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects. Project page: $\href{https://morphometricimitation.github.io}{\text{this https URL}}$