Research
PAMI: Part Anchored Motion for Text to Human-Object Interaction Generation
Overview Research area: Computer vision and generative motion synthesis, specifically text-conditioned full-body human–object interaction (HOI) generation. Technical level: Advanced. The paper assumes

- arXiv
- 2609.38466
- Published
- 2026-09-29
- Authors
- Chuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger, Gerard Pons-Moll
AI summary
Overview
Research area: Computer vision and generative motion synthesis, specifically text-conditioned full-body human–object interaction (HOI) generation.
Technical level: Advanced. The paper assumes familiarity with variational autoencoders, latent diffusion/flow matching, SMPL-family body models, and contact-based geometric losses.
One-sentence scope: The paper introduces PAMI, a coarse-to-fine framework that represents object motion as body-part "votes" in a learned interaction latent space, generates a coarse interaction from text, and then refines contact geometry with hybrid long- and short-range surface sensing.
What This Paper Is About
Generating a human motion and an object trajectory that both match a text prompt and stay physically coordinated over time is hard, because small relative errors produce floating grasps, object drift, missed contact, or penetration. Most prior full-body methods represent the human and the object as two separate trajectories and try to learn their coupling implicitly through feature exchange, contact prediction, or hand-tuned distance weights. PAMI instead asks what representation is actually suitable for generating interaction motion, and answers with a part-anchored voting scheme plus a hierarchical generator that first gets the interaction roughly right and then fixes the fine contact geometry.
Key Contributions
-
The first part-based latent generative framework for text-conditioned full-body HOI motion generation. Generation is hierarchical, coarse-to-fine, and the authors state it achieves the best performance on the benchmark. Code and model are stated to be publicly released.
-
A body-part based object voting representation. Object motion is anchored to body parts, and the anchors vote for object movement through learned weights, allowing efficient generation in latent space while preserving subtle human–object relations and producing semantically well-aligned coarse interaction motion.
-
A hybrid local surface sensing module used to train a RefineNet (PamiRefiner). It reasons about fine human–object geometry and contacts to improve the coarse latent-generation output. Trained on mixed data from generation and ground truth, it improves not only the authors' own latent generation results but also baseline generations zero-shot.
Main Findings
-
Contact recall improvement: PAMI achieves 14.5% higher contact recall than the previous state-of-the-art method on the InterAct benchmark (0.836 ± 0.006 for hands versus LIGHT's 0.730 ± 0.003).
-
Best performance on almost all metrics: Against HOI-Diff, CHOIS, InterDiff, Text2HOI, InterAct, and LIGHT, PAMI reaches R-Precision (Top 1) 0.840 ± 0.003, FID 0.092 ± 0.003, Multimodal Distance 2.195 ± 0.016, penetration 0.062 ± 0.001, Contact 0.193 ± 0.001, and hand interaction precision/recall/F1 of 0.843/0.836/0.821. Ground truth is 0.859, 0.000, 1.475, 0.051, 0.060, 0.219, 1.000, 1.000, 1.000 respectively.
-
Ablation of the voting representation: Removing part factorization drops R-1 to 0.822 ± 0.002 and FID to 0.171 ± 0.006; removing anchors drops hand F1 to 0.686 ± 0.001; removing learned voting weights drops R-1 to 0.831 ± 0.001; and removing the absolute root representation degrades most severely, with FID rising to 0.401 ± 0.011 and R-1 falling to 0.780 ± 0.005. Part factorization lowers autoencoder reconstruction error by 60.8%.
-
Refinement nearly halves penetration: Applying four PamiRefiner steps to the same coarse generations reduces penetration from 0.115 ± 0.003 (no refinement) to 0.062 ± 0.001, and improves full-body F1 from 0.389 ± 0.004 to 0.457 ± 0.004 while R-1 stays stable and FID improves.
-
Both training streams matter: Removing the generation stream or the corruption/recovery stream degrades performance, indicating the two data streams are complementary.
-
Both sensor ranges matter: Removing long-range probes hurts overall interaction quality; removing short-range probes mainly degrades hand contacts.
-
Zero-shot generalization of the refiner: Applying a frozen PamiRefiner to LIGHT output reduces penetration from 0.134 to 0.069 and raises hand F1 from 0.704 to 0.755; applied to InterAct output it reduces penetration from 0.126 to 0.077 and raises hand F1 from 0.658 to 0.738.
-
Anchor weights are meaningful: The paper reports that the learned body-part anchor weights strongly correlate with the contact body parts.
Methodology in Plain English
The starting idea comes from the classic Hough Transform: instead of detecting a hard-to-find global structure directly, let many small local elements each vote for it and take their consensus, which is robust to noise and irrelevant votes. Here the "global structure" is the object's pose, and the "local voters" are body-part anchors.
Concretely, the authors define K = 6 anchors — selected body joints of the SMPL model plus one extra "free" anchor that is identical to world space so that a detached object can vote independently of the human. Object translation is written relative to each anchor's own local frame. A variational autoencoder, PamiVAE, encodes human body parts and the object separately (with a temporal encoder per body part and one for the object, each downsampling by a factor of four so each latent token covers four motion frames), couples them with spatiotemporal attention that stacks all tokens and applies causal temporal attention, and then decodes object rotation, per-anchor relative translations, and per-frame routing logits. The final object translation is a softmax-weighted sum of the anchor-specific translations — a soft, differentiable voting scheme that can down-weight body parts not involved in the interaction.
The VAE is trained with L2 reconstruction losses on human and object motion, a KL term, a composition loss comparing the voted translation to the observed object translation, and a weight loss against pseudo ground-truth contact weights (a part's weight is one if any of its vertices is within 1 cm of the object; if no part contacts the object, the free anchor is assigned one).
Generation happens in two stages. First, PamiGen, a rectified-flow (flow matching) transformer conditioned on text and the canonical object mesh, samples a coarse interaction latent. Beyond the standard rectified-flow loss, the authors decode the estimated clean latent and add an explicit Euclidean-space self-consistency loss that pulls the free-anchor trajectory toward the body-part-anchored trajectories on contact frames (using a Huber distance), plus a trajectory loss. At inference, sampling goes from noise at τ = 1 backward to τ = 0 with classifier-free guidance.
Second, PamiRefiner refines the coarse motion in explicit space using hybrid surface sensors. Long-range probes cover 22 body joints, recording the direction and distance from each joint to its closest object surface point along with the surface normal there; these capture overall body-part influence. Short-range probes sit on hand joints and fingertips and additionally collect the object surface points within a given radius r, transformed into the local coordinate of their query point in the style of GEARS and encoded with a PointNet-like encoder. An MLP plus transformer predicts delta updates to human pose, human rotation/translation, and object rotation/translation. The refiner is trained on mixed data: a "recovery stream" of corrupted ground-truth motion (encoded by the VAE, decoded, and perturbed with temporally smooth noise, supervised against ground truth with L2) and a "generation stream" refined with geometric objectives for contact, penetration, support, smoothness, and preservation of non-interacting regions. At inference the same feed-forward refiner is applied recursively for a small fixed number of steps; the paper reports four steps.
Evaluation follows LIGHT on the standard InterAct benchmark across semantic alignment (R-precision, multimodal distance), motion quality (FID, diversity, foot-skating ratio), and interaction geometry (penetration depth, contact ratio, interaction precision/recall/F1). R-precision is evaluated with a batch size of 64. Unlike LIGHT, which evaluates the geometry metrics only on hand joints, PAMI reports them for both hand and all body joints.
Why This Matters
Impact on research: The paper argues that the dominant practice of representing a human and an object as two separate trajectories is a structural bottleneck, and offers a different representational prior — part-anchored voting in latent space — that is compatible with, and improves, other generators. The demonstration that a frozen PamiRefiner improves InterAct and LIGHT outputs zero-shot suggests a reusable, model-agnostic refinement stage for the HOI community rather than a single monolithic model.
Real-world applications:
- Robotics and embodied AI, where language-specified interaction plans must translate into physically coherent manipulation and whole-body motion.
- Character animation for film, games, and virtual production, replacing motion capture or manual authoring of object interactions.
- AR/VR content creation, where users describe an interaction in text and get a coordinated human and object animation.
- Simulation and synthetic data generation, producing contact-rich interaction data that satisfies geometric constraints.
Industry relevance: The paper positions text-conditioned HOI generation as a cornerstone of physical intelligence with direct implications for robotics, animation, and AR/VR, and text as a scalable interface that removes the need for motion capture or manual authoring. Public release of code and models is stated, which lowers the barrier for studios and robotics teams to adopt the part-anchored representation or the refiner as a post-processing step on top of existing generators.
Future Directions
- Closing the gap to ground truth. The paper's own numbers show a persistent gap on several metrics (for example hand interaction F1 of 0.821 versus the ground-truth 1.000, and penetration of 0.062 versus 0.060), leaving room for further refinement.
- Improving foot-skating. PAMI's foot-skating ratio (0.078 ± 0.001) is higher than several baselines and above the ground-truth value of 0.051, an unresolved trade-off against its gains in penetration and contact.
- Broader generalization of the refiner. So far it is verified zero-shot on LIGHT and InterAct only; whether it transfers to hand–object methods or multi-object settings such as HIMO is not reported.
- Stated limitations. The paper's final section is titled "Conclusion and limitations," but the provided content does not enumerate specific limitations, and the dataset size for InterAct is not reported in this content.
Target Audience
Researchers and graduate students in computer vision, graphics, and embodied AI working on motion generation, human–object interaction, or diffusion/flow-based generative models; practitioners in animation, AR/VR, and robotics who need text-driven, contact-accurate interaction synthesis; and readers interested in how classic computer-vision machinery (the Hough Transform) can be repurposed as a representation prior inside a modern latent generative model.
Authors’ abstract
Text-conditioned full-body human-object interaction (HOI) generation requires synthesizing human motion and object trajectories that match the input text while remaining precisely coordinated over time. Most methods represent the human and object as separate trajectories and predict the global human-object couplings. Learning this complex, dynamically changing relationship implicitly, however, often yields object drift, missed contact, and penetration. We introduce PAMI, a Part-Anchored Motion framework for Interaction generation. Inspired by the classic Hough Transform, our key idea is to localize object motion by letting body-part anchors vote for it: we express object motion relative to multiple body-part anchors and use PamiVAE to learn an interaction latent space, decoding frame-wise weights that aggregate these part-specific votes. Building on this representation, PAMI generates interactions in a coarse-to-fine hierarchy. PamiGen first generates a coarse human-object interaction from text in this structured latent space, and PamiRefiner then recursively resolves fine-grained contact geometry using a hybrid surface-sensing representation, combining long-range probes that capture overall body-part influence with short-range sensors that resolve detailed contacts near the object surface. Experiments on InterAct show that PAMI generates more faithful interactions and more accurate human-relative object motion than previous methods, achieving 14.5% higher contact recall than the previous state of the art. Extensive ablations validate the contributions of both the part-anchored voting representation and hybrid surface-sensing refinement.