Skip to content
AI.info

Research

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning Overview Research area: Computer Vision / Robot Learning — 3D hand motion reconstruction and vision-based tactile

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning
arXiv
2610.09455
Published
2026-10-07
Authors
Seungjun Moon, Subin Jeon, Sangwoo Kim, Hanbyul Joo, Jinwoo Shin

AI summary

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

Overview

Research area: Computer Vision / Robot Learning — 3D hand motion reconstruction and vision-based tactile (contact and force) estimation from monocular egocentric video.

Technical level: Advanced. The paper assumes familiarity with MANO hand parameterization, video diffusion backbones, LoRA adaptation, linear blend skinning, inverse kinematics, and dexterous-hand retargeting.

One-sentence scope: The paper turns a pre-trained video diffusion model (Cosmos 3 Nano) into a clip-level hand tracker that jointly outputs metrically accurate bimanual 3D hand motion and dense per-vertex contact and force, and shows the output is usable as robot action labels.

What This Paper Is About

Robot policies need 3D hand action labels to be trained from human egocentric video, but existing hand trackers regress pose from cropped frames with weak priors on hand motion and object interaction, producing jittery, physically inconsistent 3D trajectories — the paper reports roughly 8 mm of within-video variation in estimated hand size, plus depth oscillation and drift, as shown in Figure 1 over 850 frames of an ARCTIC test clip. A second gap is that human video datasets carry no tactile information, even though touch distinguishes a successful grasp from hovering over an object.

RLHND's goal is to close both gaps with one model: a single video foundation model-based tracker that outputs both MANO hand motion and dense contact and force over the 778 MANO vertices, so human videos can be used directly for robot policy training.

Key Contributions

  1. Joint pose-and-tactile tracking from one video backbone. RLHND routes clips through the clean-latent conditioning interface of Cosmos 3 Nano to extract spatiotemporal features, and from those features estimates metrically accurate bimanual motion together with dense per-vertex contact and force over the hand surface.

  2. Three physically grounded design choices. A cached hand shape (one shape parameter per video, optionally injected from a pre-calibrated value), an anatomically constrained pose parameterization that regresses 29 DoF instead of the full 45 rotational DoF of MANO, and a decoupled tactile expert with LBS-based feature spreading that spreads 16 bone-level features to 778 vertices without per-vertex attention.

  3. State-of-the-art results on both tasks. RLHND outperforms baselines on nearly every metric on HOT3D, ARCTIC ego, and zero-shot EgoDex for motion reconstruction, and achieves the best contact scores on OpenTouch, DexYCB, HOT3D, and ARCTIC plus competitive force estimation on OpenTouch and PVDB.

  4. Demonstrated robot-learning utility. RLHND's outputs retarget to five dexterous hands with the lowest mean Q-err and among the smoothest commands, and let Dexterous Point Policy improve real-robot success rates using automatically labeled fingertip contact instead of manual annotation.

Main Findings

  • Motion reconstruction improves substantially on all three benchmarks. On HOT3D, RLHND reaches F_Acc 0.996, MPJPE-p 13.01, PA-p 6.72, MPJPE+OOS 13.08, and EPE2D-p 6.84, compared with ACE-Ego-Hand's 0.928, 24.41, 8.60, 20.30, and 38.09. On ARCTIC ego RLHND records 0.993, 13.36, 6.42, 15.40, and 12.19. On the held-out, zero-shot EgoDex it records F_Acc 1.000, MPJPE-p 19.90, PA-p 10.12, MPJPE+OOS 19.89, and EPE2D-p 20.72.

  • Shape caching eliminates hand-size variation but not jitter. The β-cache variant drives σ_shape to 0.00 on all three test sets (versus 1.30 for the per-window RLHND on HOT3D and 1.22 for ACE-Ego-Hand), yet jitter stays essentially unchanged (4.39 vs 4.38 on HOT3D), which the authors attribute to per-clip analysis in Appendix B.2.

  • Contact estimation leads on all four datasets. RLHND scores F1 0.696 / AUROC 0.980 on OpenTouch, 0.572 / 0.915 on DexYCB, 0.589 / 0.959 on HOT3D, and 0.602 / 0.935 on ARCTIC, ahead of HACO (0.373 / 0.541, 0.543 / 0.883, 0.244 / 0.749, 0.580 / 0.907) and HOPE (0.663 / 0.894, 0.506 / 0.868, 0.197 / 0.762, 0.591 / 0.914).

  • Force estimation is strong on OpenTouch but not best everywhere. RLHND achieves MAE 0.489 and RMSE 2.508 kPa on OpenTouch, beating HOPE (1.781 / 5.388), PressureVision (1.930 / 6.190), and PressureVision++ (1.920 / 6.200). On PVDB, RLHND's RMSE 2.066 is the best reported, but its MAE 0.274 is above HOPE's 0.236; the paper describes its force estimation as "competitive."

  • Retargeting quality is best across all five hands. With the DexPilot solver and wrist-relative targets, RLHND achieves the lowest mean Q-err on Sharpa Wave (22 DoF, 12.7), WUJI v2 (20 DoF, 16.1), Shadow (24 DoF, 9.6), Inspire RH56 (12 DoF, 6.0), and ALLEX (15 DoF, 10.2), with jerk at 0.18, 0.16, 0.19, 0.07, and 0.17. ACE-Ego-Hand records 14.0, 17.4, 10.8, 7.4, and 12.0 on Q-err.

  • Real-robot success rates rise with RLHND labels. Over 100 human : 100 robot episodes, DPP + RLHND improves the All average from 80.8 to 87.5, the Pick and Place average from 89.1 to 93.8, and the Bimanual average from 64.1 to 75.0, with the largest single-task gain on Assemble tissue (75.0 to 90.6).

  • The video backbone is the dominant component. In leave-one-out ablations under a 20k-step schedule, removing Cosmos 3 (variant A0) degrades every pose metric on ARCTIC ego and EgoDex (ARCTIC MPJPE-p 16.62 vs 13.36 for the full model), more than removing constrained pose (14.26) or β-conditioning (14.87).

  • Contact-only mesh-derived data matters greatly. Removing it (variant B2) collapses contact F1 to 0.047 on DexYCB, 0.145 on HOT3D, and 0.117 on ARCTIC, while force changes little (MAE 0.500, RMSE 2.524).

  • Ablations support both tactile design choices. Dropping Cosmos 3 (B0) or replacing the fixed MANO skinning weights with a learnable 778×16 spread matrix (B1) both reduce performance relative to the full tactile expert (B3).

  • Ground-truth shape helps retargeting, not joints. Variant A4 (full model + GT β) gains nothing in joint-level metrics on ARCTIC ego but lowers Q-err from 10.12 to 9.69; the paper reports no EgoDex numbers for A4.

Methodology in Plain English

RLHND processes a video as a sequence of clips, each holding a fixed number of consecutive frames, and produces MANO pose parameters, metric camera-space wrist translation, per-frame hand presence/visibility, and per-vertex contact and force.

Instead of feeding cropped hand images to a per-frame regressor, the model feeds whole clips into the pre-trained Cosmos 3 Nano backbone through its conditioning-frame interface — the path designed to receive clean context frames during pre-training. This differs from prior work that pushes clean video through the denoising interface at zero noise level, and the paper argues the conditioning path keeps fine-tuning aligned with pre-training. The video tokenizer (Wan2.2 VAE) turns each clip into a clean latent, and only the patch embedding and generation pathway are fine-tuned with LoRA, yielding a deterministic spatiotemporal feature grid. The authors also feed an empty prompt because the video stream is already conditioned on the clean frames.

A pose expert stream tokenizes those features with spatial positional and Fourier-encoded ray embeddings and decodes them with alternating spatial cross-attention and bidirectional temporal self-attention. Two changes distinguish it from the base tracker. First, hand tokens are built from a cached shape parameter (plus a zero-initialized MLP), so one hand geometry is used for the whole video, and a pre-measured shape can be injected directly; in training the ground-truth shape is used with probability 0.5 and the shape-head loss is disabled for those samples. At inference, clips are decoded until the first sufficiently visible clip is found, whose shape becomes the cache, after which earlier held-back clips are decoded again — if no clip qualifies, each clip keeps its own shape. Second, the pose output is anatomically constrained: MANO allows 3-axis rotation at 15 joints (45 DoF), but the human hand has only 21 DoF, so rotation is factorized about twist, spread, and bend axes and twist and spread are frozen for the PIP and DIP joints, leaving 29 DoF. The thumb is left unconstrained because its carpometacarpal joint is a saddle joint. Because existing labels are 45-DoF, the authors convert them into the feasible 29-DoF set using inverse kinematics with damped Gauss-Newton steps.

A separate tactile expert stream takes the same video features but has its own weights and is trained with the pose stream frozen, so tactile supervision never perturbs pose. Rather than attending over all 778 vertices per frame, it attaches one token per each of the 16 MANO kinematic-tree joints and expands to vertices only at the output, using the fixed MANO linear-blend-skinning weights plus a learned vertex embedding. Two small MLP heads read a contact logit and a force-distribution logit per vertex; total per-hand force is read from the hand token through a softplus and distributed over vertices with a softmax. Training proceeds in two stages: stage 1 trains only the pose stream with a combination of rotation, shape, 3D, 2D, translation, presence, temporal, and ray losses, and stage 2 trains only the tactile stream with a positively re-weighted binary cross-entropy for contact, a per-hand total force term, and a distribution term.

Why This Matters

Impact on research. The paper argues that the reliability of 3D hand labels — not just 2D image alignment — is what limits the use of human video for robot learning, and shows that a video foundation model's learned priors on hand-object interaction can be redirected into a deterministic tracker. It also treats tactile estimation as a first-class output rather than an add-on, using a frozen-pose decoupled expert so scarce force labels cannot degrade pose. The retargeting and real-robot results position the tracker output as an action space usable across heterogeneous dexterous hands.

Real-world applications.

  • Training dexterous robot policies from ordinary human demonstrations, with fingertip contact labeled automatically rather than by hand.
  • Retargeting human hand motion into joint commands for many different dexterous hands, such as the five evaluated (Sharpa Wave, WUJI v2, Shadow, Inspire RH56, ALLEX).
  • Annotating existing egocentric human-video corpora with dense contact and force, enabling tactile-aware behavior cloning without wearable sensors.
  • Providing contact supervision where vision-language-action models already consume touch, distinguishing a real grasp from hovering.

Industry relevance. The work targets robot foundation model training pipelines that depend on human video, where data collection cost — not model size — is the bottleneck. Automatic contact labeling removes a manual per-demonstration annotation step that the paper notes was previously required for Dexterous Point Policy, and a single model providing both keypoints and contact simplifies the labeling stack.

Future Directions

  • Reduce jitter further. Shape caching removes hand-size variation but leaves frame-to-frame jitter essentially unchanged (4.39 with cache vs 4.38 without on HOT3D), so the source of residual jitter remains an open problem.
  • Pool tactile datasets more broadly. The paper notes that differences in sensor calibration and hand-surface correspondence make it difficult to pool pressure-board, glove, and mesh-derived label sources without a common representation; RLHND uses HOPE's unified format and could extend further.
  • Close the force gap. RLHND's PVDB MAE (0.274) is above HOPE's (0.236), leaving headroom in force accuracy even as RMSE is best reported.
  • Address dataset diversity. The ethics statement flags that the limited diversity of public datasets may cause performance variation across subjects, hand sizes, and skin tones that the evaluation sets do not capture.
  • Validate inferred tactile signals in deployment. Contact and force are inferred, not measured, so the authors recommend joint-torque or current limits and human supervision when acting on retargeted commands.

Target Audience

Robotics and embodied-AI researchers building policy learning pipelines from human video; computer vision researchers working on hand pose reconstruction, MANO parameterization, and video diffusion backbones; tactile sensing and contact-estimation researchers; and engineers working on dexterous-hand retargeting and teleoperation who need metrically stable, temporally smooth joint commands. Readers without background in hand parameterization or diffusion model interfaces will find the method sections demanding, though the motivation, results tables, and robot-learning demonstrations are accessible.

Authors’ abstract

Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.

Read the original paper