Skip to content
AI.info

Research

FlowHMR: Physically Plausible Motion Capture from Video

FlowHMR: Physically Plausible Motion Capture from Video Overview Research area: Computer vision / computer graphics — monocular (single-camera) human motion capture, generative modeling, and physics-b

FlowHMR: Physically Plausible Motion Capture from Video
arXiv
2610.03691
Published
2026-10-02
Authors
Zhanke Wang, Chengfeng Zhao, Qing Shuai, Jingzhong Lin, Heng Li, Zeyu Ling, Yuxin Wen, Jing Li, Di Kang, Chunchao Guo, Linchao Bao

AI summary

FlowHMR: Physically Plausible Motion Capture from Video

Overview

Research area: Computer vision / computer graphics — monocular (single-camera) human motion capture, generative modeling, and physics-based character control.

Technical level: Advanced. The paper assumes familiarity with diffusion and flow-matching generative models, reinforcement learning with policy gradients (GRPO), SMPL-family body representations, and physics simulation of humanoid controllers.

Scope: The paper proposes FlowHMR, a two-stage framework that pretrains a video-conditioned flow-matching generator of global 3D human motion and then post-trains it with reinforcement learning using fidelity and physics-tracking rewards, evaluated on a new 4,182-clip internet video benchmark called Wild-4K.

What This Paper Is About

Recovering the full 3D motion of a person — body pose, hands, and their global trajectory through the world — from an ordinary handheld video is deeply ambiguous, because depth information is lost, bodies are frequently occluded, and human movement is entangled with camera movement. Existing learning-based methods regress motion directly with geometric losses, which tends to produce over-smoothed, averaged results that are not guaranteed to be executable under real physics, so simulation controllers often fail to track them. FlowHMR instead treats video motion capture as a conditional generation problem and uses physics simulation as a training-time reward signal, so the recovered motion stays faithful to the video while becoming physically trackable.

Key Contributions

  1. FlowHMR, a video-conditioned flow matching framework that jointly models body pose, hand articulation, and global root trajectory in a unified world coordinate frame, using a 0.46B-parameter multimodal diffusion transformer (MM-DiT) conditioned on visual features and camera orientation.

  2. An RL post-training scheme with two rewards — a fidelity reward measuring agreement with ground-truth motion, and a physical tracking reward measuring how accurately a frozen, pretrained PHC+ humanoid controller tracks the generated motion in simulation — optimized with MixGRPO.

  3. Wild-4K, a large and diverse evaluation set of 4,182 internet video clips (curated from Koala-36M) covering everyday activities, highly dynamic movements, and severe occlusions, for assessing motion fidelity and physical tracking success-rate in the wild.

  4. An analysis of how the two rewards interact, showing that optimizing tracking alone causes reward hacking and severe reconstruction degradation, while the joint reward produces physical plausibility at moderate reconstruction cost.

Main Findings

  • Tracking success on Wild-4K: FlowHMR reaches a physical tracking success rate of 82.47%, versus 62.82% for the strongest baseline, GVHMR. The pretrained model without RL ("Ours w/o RL") reaches 78.79%, so post-training adds 3.68 percentage points.

  • Baseline comparison (Wild-4K success rate): PromptHMR 60.33%, GENMO 58.80%, DuoMo 56.71%.

  • Geometric artifacts (Wild-4K, medians over 4,182 clips): FlowHMR reduces Penetration from 0.60 mm to 0.06 mm, Floating from 0.53 mm to 0.20 mm, and Skating from 2.48 mm to 2.17 mm after post-training. GVHMR still has lower Skating at 1.59 mm.

  • Blinded human evaluation: FlowHMR wins 61.5–79.2% of valid annotations against the four external methods, with losses of 3.8–10.0%. Against its own pretrained version (129 valid annotations), the post-trained model gets 34 wins (26.4%), 76 ties (58.9%), and 19 losses (14.7%).

  • RICH benchmark (185 sequences): Post-training raises PHC+ success rate from 48.65% to 51.35% and reduces tracking MPJPE from 133.5 mm to 118.8 mm, the best among evaluated methods. WA-MPJPE and PA-MPJPE rise slightly relative to the pretrained model (55.70 → 56.68 mm and 30.78 → 31.40 mm), while Jitter and foot sliding decrease.

  • Cross-controller execution (SONIC on Unitree G1, 615 clips): FlowHMR reaches 77.72% success, above GVHMR at 54.15%, showing the gains transfer to a controller never used during post-training.

  • Data scale matters: Increasing pretraining data from 200 to 3,000 hours reduces WA-MPJPE from 149.44 to 67.82 mm and hand-rotation error from 25.17 to 16.08 degrees in the Studio ablation.

  • Generative formulation vs. regression: Flow matching reduces WA-MPJPE from 73.28 to 67.82 mm and improves RTE, Jitter, and foot sliding, though hand-rotation error is slightly higher (16.08 vs. 14.55 degrees) and MPJPE is nearly identical (43.49 vs. 43.44 mm).

  • Model size: Scaling from 0.2B to 1B parameters reduces WA-MPJPE from 67.82 to 62.84 mm and Jitter from 11.75 to 10.93, but gains are not monotonic — the 0.46B model has the lowest MPJPE, RTE, and foot sliding.

  • Reward hacking without fidelity: Tracking-only GRPO reaches 99.62% success rate but inflates WA-MPJPE to 255.67 mm and PA-MPJPE to 139.20 mm, marked as reward hacking in the paper.

  • SFT alternatives are weaker: Filtered-original SFT gives 78.38% success (no improvement over pretrained), and projected-target SFT reaches 80.25% success but raises WA-MPJPE to 70.80 mm and PA-MPJPE to 42.13 mm.

  • Adding a kinematic Phys-Err reward (penetration, floating, skating) to the joint rewards yields comparable success (82.57% versus 82.47%) but higher reconstruction errors and a slightly lower Tracking Score, so the two-reward configuration is preferred.

Methodology in Plain English

Stage 1 — Pretraining a generator. The authors formulate motion capture as generating a motion that fits a video, rather than regressing a single answer. They render roughly 3,000 hours of synthetic video from 650 hours of source motion, randomizing cameras, textures, lighting, and occluders so the model sees occlusions and partial visibility. All motion is fitted to a common SMPL-H skeleton. The generator is trained with flow matching, which learns to gradually transform random noise into a motion conditioned on per-frame visual features (from a frozen SAM 3D Body encoder) and camera orientation. The motion representation encodes root velocity and height, joint positions, 6D joint rotations, body shape, and foot-contact labels. Training ran for 500K steps with a batch size of 1,024, with a learning rate cosine-decayed from 10⁻⁴ to 10⁻⁵, using a weighted Smooth-L1 flow-matching loss plus forward-kinematics losses on global joint rotations and positions. At inference, the model integrates the learned velocity field with 20 uniform Euler steps and no classifier-free guidance, producing one motion per video without simulation-based selection or correction.

Stage 2 — Post-training with physics feedback. Because supervision only encodes geometry, not dynamics and contact, the authors fine-tune the generator with reinforcement learning. Each training video is used to sample a group of G = 8 motions using stochastic (SDE) sampling inside a window of denoising steps and deterministic (ODE) steps elsewhere, following MixGRPO. Two rewards are computed: a fidelity reward based on the exponential of the negative mean joint-position error against ground-truth motion in the world frame with no alignment, and a tracking reward based on how closely a frozen PHC+ controller reproduces the motion in Isaac Gym (with per-step error capped at d_max = 0.5 m; τ = 0.15 m for both rewards). Only groups containing both successful and failed tracking outcomes are kept, and within each group the two rewards are standardized separately, summed with equal weights, and clipped. Only the generator is updated; the visual encoder and controller remain frozen.

Evaluation. Since Wild-4K has no ground-truth motion, the authors rely primarily on blinded pairwise human comparisons of direct predictions, supplemented by simulation tracking success, geometric artifact measurements (Penetration, Floating, Skating), and comparisons on the annotated RICH benchmark.

Why This Matters

Impact on research. The paper makes a case for treating video motion capture as conditional generation rather than deterministic regression, and demonstrates that physics simulation can be used as a training-time reward instead of a post-hoc correction step. It also shows that reward design is delicate: optimizing only for trackability produces motions that are easy to simulate but wildly unfaithful to the video, a concrete reward-hacking case study for the generative-motion community. Wild-4K adds a large, unannotated, in-the-wild evaluation resource that complements laboratory benchmarks like Human3.6M and the 60-sequence 3DPW.

Real-world applications:

  • Character animation for film, games, and virtual production, where animators need world-space motion with a global trajectory from ordinary footage.
  • Humanoid robot learning from human demonstrations, where recovered human motion must be physically executable to serve as a control reference.
  • Sports and clinical movement analysis, using consumer video to analyze motion without a mocap studio.
  • Motion data mining from internet video, converting the huge volume of existing footage into physically plausible motion assets.

Industry relevance. The pipeline targets a practical bottleneck for studios and robotics teams: the gap between visually plausible and physically executable motion. The reported cross-controller result with SONIC on a Unitree G1 humanoid is directly relevant to robotics, and the authors release code and checkpoints at https://github.com/flowhmr/flowhmr with a project page at https://flowhmr.github.io/.

Future Directions

  • Removing or replacing the ground-truth fidelity reward. The current post-training depends on paired ground-truth motion for its fidelity term (~571k pairs passed the tracking filter), which limits how far the method can extend to unlabeled internet video.

  • Generalizing beyond flat ground-plane simulation. The authors explicitly exclude underwater, aerial, and suspended settings to match their simulation environment, leaving non-flat terrain scenario coverage as an open problem.

  • Higher-fidelity hand and fine-detail reconstruction. Finger poses are not assessed in human evaluation (a shared default is used), and the generative formulation slightly increases hand-rotation error relative to regression in the ablation (16.08 vs. 14.55 degrees).

  • Robustness to controller choice and reward design. The tracking reward depends on a specific frozen PHC+ controller and succeeds less often on RICH (51.35%) than on Wild-4K, suggesting that transfer across controllers, motions, and scenes remains an open question — as does the paper's finding that an added kinematic Phys-Err reward did not improve the joint-reward result.

Target Audience

Researchers and graduate students working on 3D human pose and motion capture, generative motion synthesis, and physics-based character animation; reinforcement learning researchers interested in reward design for generative models with non-differentiable objectives; and robotics or VFX engineers who need motion recovered from ordinary video to be directly executable by a simulated or real humanoid.

Authors’ abstract

We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and train the network with geometric supervision. However, recovering human motion from monocular video is inherently ambiguous in depth, and direct regression tends to collapse toward an averaged solution. Moreover, the recovered motions are not guaranteed to be physically plausible, so physics-based tracking of them often fails. To address these challenges, we formulate video motion capture as a video-conditioned motion generation problem and first pretrain a flow matching model for this task. Given an input video, the pretrained model generates diverse motion candidates, but not all of them are faithful to the video or physically trackable. We therefore post-train the model using Group Relative Policy Optimization (GRPO) with two rewards. A fidelity reward encourages consistency with the input video. A tracking reward favors motions that a physics-based controller can track successfully. Together, these rewards shift the model's output preference, so the post-trained model stays faithful to the input video while producing more physically plausible motion. We further introduce Wild-4K, a large and diverse dataset of about 4K internet videos, for evaluating human motion recovery in the wild. Qualitative and quantitative experiments on Wild-4K show that our method outperforms state-of-the-art methods in overall motion fidelity and achieves a physical tracking success rate of 82.47%, compared with 62.82% for the strongest baseline, GVHMR.

Read the original paper