Research
InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
Overview Research area: Computer vision — 3D hand motion estimation, egocentric (first-person) video, streaming 3D reconstruction. Technical level: Advanced. The paper assumes familiarity with MANO ha

- arXiv
- 2609.35743
- Published
- 2026-09-28
- Authors
- Kerui Ren, Kaiwen Song, Weiguang Zhao, Yuxi Wang, Yufei Liu, Bo Dai, Haoyu Guo, Chunhua Shen, Mulin Yu, Tao Lu, Junting Dong
AI summary
Overview
- Research area: Computer vision — 3D hand motion estimation, egocentric (first-person) video, streaming 3D reconstruction.
- Technical level: Advanced. The paper assumes familiarity with MANO hand models, SLAM/bundle adjustment, transformer backbones, and 3D reconstruction metrics.
- Scope: The paper introduces InfiniHand, an end-to-end streaming feed-forward model that jointly predicts hand locations, MANO hand parameters, and camera trajectories from uncalibrated egocentric video, and evaluates it against state-of-the-art baselines on four camera-space benchmarks and three world-space benchmarks.
What This Paper Is About
Estimating how hands move in a shared three-dimensional world from head-mounted camera video normally requires stitching together several separate systems: a hand detector, a hand pose network, and a SLAM pipeline for camera motion. The paper argues this cascade causes error accumulation, engineering complexity, and heavy computation.
InfiniHand instead trains a single streaming network that simultaneously outputs hand masks, MANO hand parameters, and camera poses, then places the hands into world coordinates using the estimated camera trajectory, supported by a lightweight sparse bundle adjustment step to limit long-term drift.
Key Contributions
- A unified streaming feed-forward framework that combines hand localization, MANO parameter prediction, and camera trajectory estimation in one architecture, enabling fast world-space hand motion reconstruction from egocentric video.
- A curated large-scale training corpus assembled by aggregating and cleaning existing public egocentric datasets into a standardized, high-quality format, paired with a dedicated two-stage training scheme.
- State-of-the-art accuracy claims with high efficiency, reporting the lowest PA-p on all four camera-space benchmarks and the lowest W-MPJPE on all three world-space benchmarks, at 11.19 FPS.
- Demonstrated generalization to in-the-wild video, with qualitative results on Ego4D sequences that lack ground-truth MANO annotations.
Main Findings
- Camera-space accuracy: InfiniHand reports the lowest PA-p on all four evaluation benchmarks, reducing ARCTIC PA-p from 9.82 mm (ViDiHand) to 7.72 mm, a 21.4% reduction, and EgoDex PA-p from 17.22 mm to 7.29 mm, a 57.7% reduction.
- World-space accuracy: InfiniHand reports the lowest W-MPJPE on all three evaluated datasets, improving ARCTIC from 65.86 mm (WiLoR-SLAM) to 59.21 mm (10.1% reduction) and EgoDex from 78.41 mm (Dyn-HaMR) to 28.77 mm (63.3% reduction).
- Throughput: InfiniHand runs at 11.19 FPS, compared with WiLoR-SLAM at 8.53 FPS, HaWoR at 5.48 FPS, and Dyn-HaMR at 0.81 FPS, yielding a 2.04× speedup over HaWoR.
- Training data scale: The authors aggregate approximately 5,000 hours of egocentric data from nine public datasets (ARCTIC, HOT3D, EgoDex, DexYCB, HO3D, H2O-3D, EgoVerse, EgoLive, Xperience-10M), sampling frames at 10 FPS.
- Data cleaning impact: A Sapiens 2-based consistency check against projected MANO joints removes roughly 30–40% of the candidate training data, most prominently from EgoDex.
- Ablation — appearance features: Removing WiLoR encoded crop features nearly doubles ARCTIC MP-p, from 17.09 mm to 33.31 mm.
- Ablation — bundle adjustment: Disabling sparse BA raises ARCTIC W-MPJPE from 59.21 mm to 80.48 mm while leaving camera-space metrics unchanged.
- Ablation — joint training: Omitting Stage II raises W-MPJPE from 59.21 mm to 185.76 mm and WA-MPJPE from 43.59 mm to 111.05 mm, with only minor changes in camera-space errors.
- Ablation — localization: Replacing the learned mask head with HaWoR masks degrades detection and raises ARCTIC PA-p from 7.72 mm to 25.50 mm.
- Ablation — projection constraint: Replacing the differentiable least-squares projection (LSP) step with direct translation regression degrades translation accuracy and image-space alignment (MP-p 17.09 mm to 18.63 mm, PA-p 7.72 mm to 8.02 mm).
- Qualitative generalization: On in-the-wild sequences, the paper reports that InfiniHand preserves plausible hand geometry and image alignment where ViDiHand shows visible degradation.
- HOI4D caveat: On the HOI4D benchmark, InfiniHand reports PA-p of 12.11 versus ViDiHand's 13.96, but MP-p of 33.64 versus ViDiHand's 30.09.
Methodology in Plain English
The system is built on a streaming 3D foundation model (LingBot-Map) that maintains a persistent memory of past frames, including anchor features that define a shared reference frame, a dense recent-feature window, and compressed trajectory tokens. This memory is managed through a Geometric Context Attention mechanism.
For each frame, the model first uses full-image geometric features and a Dense Prediction Transformer decoder (initialized from a pretrained depth head) with a two-channel mask head to predict left- and right-hand masks. These masks are converted into expanded bounding boxes. Inside each box, the model fuses two streams of information: geometric patch features cropped from the backbone, and appearance features from a RGB crop processed by a WiLoR encoder. The fused representation feeds a MANO head that predicts pose, shape, global orientation, and depth, plus a head that predicts 2D landmarks. A differentiable least-squares projection solves for lateral hand translation by aligning projected 3D joints with the predicted 2D landmarks.
Training proceeds in two stages. Stage I learns camera-space hand localization and MANO reconstruction on all valid camera-space annotations. Stage II adds the camera pose head and trains jointly on 36-frame layouts consisting of four anchor frames followed by two consecutive 16-frame windows, mixing ground-truth and predicted bounding boxes in a 1:1 ratio so the model sees realistic localization noise. Losses combine camera pose and field-of-view supervision, relative motion between frame pairs, world-space joint error, temporal consistency across camera parameters, MANO parameters, masks and landmarks, plus the Stage I mask, 2D and MANO terms.
To control drift over long sequences, a DROID bundle adjustment backend refines camera poses using a binary keyframe pool that retains selected past frames for sparse refinement, rather than optimizing over the entire history. Refined poses then map hand predictions into the world frame.
Why This Matters
The paper frames world-space hand motion estimation as a prerequisite for embodied learning — converting abundant first-person human video into structured 3D geometric supervision that can train or condition robot policies. Its main practical claim is that doing this in one streaming model, rather than a cascade of separate estimators, improves both accuracy and speed.
Real-world applications (as motivated by the paper):
- Human-to-robot motion transfer: Using reconstructed hand trajectories as demonstrations for training world action models.
- Robot imitation from retrieved human examples: Conditioning in-context robot imitation on retrieved human demonstrations.
- Scalable 3D annotation of egocentric video: Turning large volumes of human video into structured hand-motion supervision.
- Dexterous data augmentation and manipulation policy learning: The paper lists these as downstream uses enabled by faster, more reliable annotation.
Industry relevance: Robotics and embodied AI groups that need large-scale hand demonstration data, and egocentric/AR-VR and wearable-camera companies interested in reconstructing hand interaction in a shared world frame, are the most likely beneficiaries. Efficiency (11.19 FPS, 2.04× over HaWoR) matters for any pipeline that must process video at scale or on constrained hardware.
Future Directions
- Metric scale recovery. The underlying LingBot-Map framework lacks inherent metric scale, so the method depends on an auxiliary post-processing alignment model whose errors can propagate into hand positions and trajectories. Removing or improving that dependency is an open problem.
- Better data coverage for hard interactions. The authors note that high-quality egocentric datasets with complex, large-amplitude two-hand interactions and camera dynamics remain scarce, limiting generalization to rapid viewpoint changes, severe occlusion, and intermittent hand visibility.
- Temporally consistent supervision. Monocular reconstruction pipelines inherit time-varying scale drift from derived monocular annotations, creating temporally inconsistent training targets that hurt long-sequence trajectory accuracy and smoothness even after global scale alignment.
- Extending the pipeline beyond hands. The paper does not report extensions to full-body pose, object state, or contact estimation, which would be logical next steps for embodied learning applications.
Target Audience
Researchers and engineers working on 3D hand reconstruction, egocentric vision, streaming 3D reconstruction, and SLAM, as well as robotics and embodied-AI practitioners who need scalable 3D human motion data. Readers evaluating accuracy-versus-throughput tradeoffs in world-space motion estimation will find the benchmark tables and efficiency analysis most relevant; readers new to MANO models, bundle adjustment, or streaming memory architectures should treat the paper as advanced material.
Authors’ abstract
World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhead. To address these limitations, we present InfiniHand, an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video. InfiniHand integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture. We train InfiniHand in two progressive stages by first learning robust camera-space hand priors and then extending to streaming world-space reconstruction. To support this process, we aggregate a pretraining corpus of approximately 5,000 hours of egocentric data across multiple public datasets. Extensive evaluations demonstrate that InfiniHand outperforms state-of-the-art baselines on in-domain benchmarks, achieving a 21.4% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift. Furthermore, InfiniHand generalizes robustly to in-the-wild videos and operates at 11.19 FPS, delivering more than twice the throughput of HaWoR.