Skip to content
AI.info

Research

VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation

VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation Overview Research area: Computer vision and 3D hand–object interaction, specifically human grasp generation and contact representatio

VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation
arXiv
2610.10197
Published
2026-10-07
Authors
Zhuo Chen, Yihua Cheng, Aleš Leonardis, Hyung Jin Chang

AI summary

VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation

Overview

  • Research area: Computer vision and 3D hand–object interaction, specifically human grasp generation and contact representation for MANO-parameterized hands.
  • Technical level: Advanced. The paper assumes familiarity with signed distance fields, variational autoencoders, latent diffusion models, MANO, continuous surface embeddings, and physics-based simulation metrics.
  • Scope: This summary covers the paper's proposed volumetric contact representation, its diffusion-based generation framework, and its reported reconstruction, grasp-generation, ablation, and runtime results on the GRAB and HO3D datasets.

What This Paper Is About

Human grasp generation methods typically predict where a hand should touch an object using contact maps defined only on the object's surface. Because those maps discard information about how far the hand sits relative to the surface, existing methods must rely on hand-crafted, nearest-neighbor penetration penalties at test time, which are unstable and can leave grasps either floating loosely or penetrating deeply. The paper introduces VolCo, a contact representation that surrounds each object surface point with a small 3D volumetric grid encoding both contact likelihood and hand correspondence, and VolCoDiff, a latent diffusion framework that generates grasps from this representation.

Key Contributions

  1. VolCo representation: A hierarchical, volume-based contact representation built by sampling N points on the object surface and expanding each into a k×k×k local grid, where every grid point stores a contact likelihood and a 4-D Continuous Surface Embedding (CSE). This extends the Mosaic-SDF idea from general object geometry to hand contact with fixed topology and correspondence.
  2. VolCoDiff framework: A volume-wise latent diffusion model that follows VolCo's hierarchy, using a 3D volumetric autoencoder (VolumeVAE) to model local hand configurations conditioned on the object SDF, and a prior-guided latent diffusion model based on Scene Diffuser to model global combinations of local configurations.
  3. Force-aware and prior-guided supervision: An auxiliary hand branch supervises generated contact latents so they can be fitted by a valid MANO hand, and an explicit spring-damper contact-force formulation enables a stability loss with slack variables.
  4. Empirical validation: State-of-the-art results on hand reconstruction from ground-truth contact and on grasp generation for both an in-domain dataset (GRAB) and an out-of-domain dataset (HO3D), with code released at https://github.com/chzh9311/volco.

Main Findings

  • Reconstruction precision: On the GRAB test set (sampled every 64 frames), the normal setting (128 × 8³) achieves the best End Point Error of 6.33 mm compared with ContactOpt (89.42), ContactGen (59.38), ManiDext (11.19), and the sparse setting (64 × 4³, 6.70). The sparse setting also beats ManiDext by 40.1% in EPE at the same number of sampled points (4096), and the normal setting by 43.4%.
  • Other reconstruction metrics: Normal setting reaches F@5mm of 0.752, F@15mm of 0.971, and AUC of 0.875, versus ManiDext's 0.616, 0.923, and 0.785. Reported improvements for the sparse and normal settings are 20.0%/22.1% (F@5mm), 5.1%/5.2% (F@15mm), and 10.6%/11.4% (AUC).
  • In-domain grasp generation (GRAB): The method attains the lowest Simulation Displacement (0.52 ± 0.64 cm) and Penetration Depth (0.27 ± 0.28 cm) and the largest Contact Area (31.1 ± 12.9 cm), exceeding the previous state of the art by 14.8%, 32.5%, and 1.6% respectively. Contact Ratio is 1.00, Entropy 2.77, and Cluster Size 3.61.
  • Out-of-domain grasp generation (HO3D): The method achieves SD of 1.06 ± 1.68 cm, PD of 0.70 ± 0.64 cm, and CA of 45.4 ± 29.2 cm, surpassing the previous state of the art by 12.3% (SD), 28.6% (PD, as listed in the table; the abstract and text state 28.5%), and 18.2% (CA). Intersection Volume is 10.9 ± 27.3 cm, CR 1.00, Entropy 2.85, Cluster Size 3.85.
  • Penetration versus tightness trade-off: The authors state that their Intersection Volume is deliberately not the smallest among compared methods, because tighter grasps with larger contact areas naturally produce more penetration volume; they report that over 28% improvement in PD on both datasets indicates a larger portion of penetrations are plausible.
  • Ablation results: Replacing point-based contact with VolCo (setting 1 to 2) improves SD from 1.72 to 1.28 cm. Adding prior guidance lowers PD from 0.42 to 0.23 cm and raises CA from 19.5 to 26.5 cm. Adding initialization improves SD from 1.63 to 0.78 cm in the no-stability-loss path. The full configuration (setting 7) reaches SD 0.52 cm, PD 0.27 cm, IV 3.78 cm, CR 1.00, and CA 31.1 cm. The stability loss reduces displacement clearly only when prior guidance is present (settings 4 to 6 and 5 to 7), which the authors interpret as geometric plausibility being a prerequisite for physical stability.
  • Runtime characteristics: VolCoDiff uses 74.04M parameters and 358.1M peak memory versus 24.41M and 191.4M for the point-contact baseline, yet requires 2.01 TFLOPs versus 13.2 TFLOPs. Inference time barely grows from 5.28 s to 5.47 s when moving from 1 to 10 batched samples, while the baseline rises from 2.90 s to 19.57 s. Total per-sample time is 10.7 s for VolCoDiff and 10.4 s for VolCoDiff-S (64 × 4³), whose SD and PD are 0.67 cm and 0.32 cm.
  • Comparison of diffusers: Among diffusion-based methods, the paper reports the second-fastest inference time and the second-fastest optimization time overall, behind only the point-contact baseline.

Methodology in Plain English

The method starts from motion-capture data. For each grasp, every point near the object is described by two things: how likely it is to be in contact (a value that saturates at 1 as the distance to the hand shrinks, computed as min{d₀²/d², 1}) and where on the hand surface that nearest point falls, encoded as a 4-D Continuous Surface Embedding. The authors then sample N points on the object surface and expand each into a local volume of half size s, discretized at resolution k. A default configuration uses k = 8, s = 1 cm, N = 128, with d₀ = s/(k−1).

Because each small volume contains only limited shape variation, a 3D variational autoencoder (VolumeVAE) compresses the 5-channel volume (1 contact + 4 CSE channels) into a 128-dimensional latent conditioned on the object SDF at each grid point. Its loss combines contact MSE (weight 5.0), contact-weighted CSE L1 (weight 1.0), a CSE-weight (barycentric) loss (weight 0.5), and a KL term with β = 10⁻⁶.

A latent diffusion model operating over the concatenation of a 16-dimensional hand latent and the N × 128 volume latent then learns which combinations of local configurations form a valid grasp, using object SDFs as conditioning and a PointNet-style encoder for positional features. Two forms of supervision shape the generation: an auxiliary branch that predicts a MANO hand so that the generated contact can actually be fitted (consistency and reconstruction losses, weight 0.1 each), and a stability loss (weight 0.1) derived from a linear spring-damper contact model F = kδ, with a 1 mm margin and slack variables of ±1 mm that tolerate small penetration errors. Finally, the predicted hand pose initializes an optimization that fits the mesh to the generated VolCo through a weighted reconstruction loss, a repulsive loss for non-contact volumes, and a regularization term that limits deviation from the initial pose.

Training uses AdamW with learning rates of 2 × 10⁻⁴ for 20 epochs (VolumeVAE) and 10⁻⁴ for 100 epochs (VolCoDiff), with 1,000 optimization iterations for pose fitting, on a workstation with an Intel Xeon w7-3445 CPU, 64 GB RAM, and two 4090 GPUs. MANO vertices are handled in a 778-dimensional weight space.

Why This Matters

This work reframes contact as a volumetric rather than surface-only signal, which removes the reliance on non-differentiable nearest-neighbor penetration penalties that the paper identifies as a source of unstable, local-minimum-prone grasp fitting. It shows that preserving fine contact detail from motion-capture data translates into measurably tighter yet less severe penetrations, and it does so with a latent diffusion model whose cost scales with token count rather than with volumetric resolution.

Real-world applications include:

  • AR/VR hand-interaction systems that need physically believable grasps for virtual objects.
  • Robotic manipulation, where a generated grasp should be both stable under simulation and realistically tight around an object.
  • Human-computer interaction and animation pipelines that require plausible hand poses for arbitrary object templates.
  • Dexterous robotic deployment trained from human motion-capture demonstrations.

Industry relevance stems from the released code, the use of standard benchmarks (GRAB, HO3D, YCB objects), and the practical runtime profile: batched inference that stays nearly flat as the sample count increases is directly useful in generative pipelines that must produce many grasps.

Future Directions

  • Controllability: The authors state their method currently prefers only tighter grasps, and propose exploring control over VolumeVAE latents as future work.
  • Scale adaptivity: The fixed number (N = 128) and fixed scale (s = 1 cm) of volumes limit coverage of object scales; adaptive volume numbers and scales are named as open directions.
  • Generalization beyond current benchmarks: HO3D objects are noted to be generally larger than GRAB objects, so extending robust behavior across wider object-scale and shape ranges remains open.
  • Balancing tightness and penetration: Since larger contact areas tend to accompany higher intersection volumes, the interplay between plausibility, depth of penetration, and grasp stability is left as an ongoing design question.

The paper also notes a broader-impact risk: while the method could improve AR/VR experiences and dexterous robotics, it risks fostering misuse of robots in harmful contexts.

Target Audience

This paper is most useful to computer vision and graphics researchers working on hand–object interaction, grasp synthesis, and contact modeling; to robotics researchers interested in generative grasp priors and simulation-based stability metrics; and to graduate students or engineers already comfortable with MANO, diffusion models, and 3D implicit or volumetric representations. Readers seeking a beginner-level introduction to grasp generation will find the prerequisites steep, though the core idea of replacing surface contact maps with local volumetric grids is conceptually accessible.

Authors’ abstract

Accurate contact modeling is fundamental to understanding hand-object interaction, yet existing contact representations are typically restricted to object surfaces and rely on hand-crafted rules to recover contact details, leading to severe penetrations and implausible results. To better exploit the rich detail in motion-capture data, we introduce Volumetric Contact (VolCo), a representation that expands surface points to a set of 3D volumetric grids. VolCo encodes 3D contact that allows precise hand part recovery, and is organized in an inherent hierarchy: local contact details within each volume and global hand geometry across all volumes. Our framework, VolCoDiff, employs two modules to capture local and global features following this hierarchy. For local contact details, we use a 3D variational autoencoder to model the possible hand configurations conditioned on the local object signed distance field (SDF). For global hand geometry, we design a prior-guided diffusion model that learns the distribution of compressed latent features aggregated from the volumetric grids. We evaluate our method on two benchmark datasets and demonstrate state-of-the-art performance in penetration and stability, indicating the capability to generate tight grasps with much less severe penetrations. Our code is available at https://github.com/chzh9311/volco.

Read the original paper