Skip to content
AI.info

Research

DexAvatar: 3D Sign Language Reconstruction with Hand and Body Pose Priors

DexAvatar: 3D Sign Language Reconstruction with Hand and Body Pose Priors Overview Research area: Computer vision — 3D human mesh recovery (whole-body and hand pose estimation) applied to sign languag

DexAvatar: 3D Sign Language Reconstruction with Hand and Body Pose Priors
arXiv
2512.21054
Published
2025-12-24
Authors
Kaustubh Kundu, Hrishav Bakul Barua, Lucy Robertson-Bell, Zhixi Cai, Kalin Stefanov

AI summary

DexAvatar: 3D Sign Language Reconstruction with Hand and Body Pose Priors

Overview

  • Research area: Computer vision — 3D human mesh recovery (whole-body and hand pose estimation) applied to sign language video.
  • Technical level: Advanced. The paper assumes familiarity with SMPL-X/MANO parametric body models, variational autoencoders, differentiable optimization, and standard pose-estimation benchmarks.
  • Scope: DexAvatar is an optimization framework that reconstructs bio-mechanically accurate 3D body and hand poses from monocular sign language video by replacing generic pose priors with two sign language-trained priors, SignBPoser (body) and SignHPoser (hands).

What This Paper Is About

Most sign language datasets provide video and automatically reconstructed 2D keypoints only, and 2D representations cannot capture depth or hand-to-body contact — different 3D hand configurations can project to identical 2D keypoints. Existing 3D mesh recovery methods are trained on general-purpose human motion, so they fail on the rapid, intricate hand articulations, frequent self-occlusions, and motion blur typical of signing. The goal of this work is to reconstruct accurate 3D signing avatars from in-the-wild monocular videos by learning pose priors specifically from sign language data.

Key Contributions

  1. Two sign language-aware pose priors. SignHPoser for hands and SignBPoser for the body, both trained to learn compact latent spaces that preserve phonologically meaningful pose variation and discourage anatomically implausible configurations.
  2. A new mocap dataset for hands. The authors collected their own sign language motion capture data using a Vicon setup of 9 high-resolution cameras and Manus gloves, from 8 signers (six proficient in Auslan, two fluent in ASL), each fingerspelling a curated list of 93 words letter-by-letter, retargeted to an SMPL-X rig.
  3. The DexAvatar optimization pipeline. An optimization framework that uses the two priors as differentiable regularizers together with temporal consistency and contact-aware terms, stabilizing estimation under self-occlusion and noisy 2D keypoints, including upper-body-only and one-handed signing.
  4. Strong benchmark results. Experiments on the SGNify motion capture dataset, described as the only benchmark available for this task, show lower reconstruction error than whole-body and hand-only mesh recovery baselines.

Main Findings

  • Benchmark improvement: On the SGNify dataset, DexAvatar reports an improvement of 35.11% in the estimation of body and hand poses relative to the state of the art. In the results section this is stated specifically as a substantial 35.11% improvement on the upper body.
  • Per-region gains over the closest baseline: DexAvatar surpasses Neural Sign Actors on the left and right hands by 16.32% and 14.11% respectively, and on the upper body by 35.11%.
  • TR-V2V error table (mm), Upper Body excluding face / Left Hand / Right Hand: FrankMoCap 78.07 / 20.47 / 19.62; PIXIE 60.11 / 25.02 / 22.42; PyMAF-X 68.61 / 21.46 / 19.19; SMPLify-SL 56.07 / 22.23 / 18.83; SGNify 55.63 / 19.22 / 17.50; OSX 47.32 / 18.34 / 18.12; Neural Sign Actors 46.42 / 16.17 / 15.23; EVA* 40.38 / 13.73 / 13.68; DexAvatar 30.13 / 13.53 / 13.08. EVA* is the authors' modification of EVA to accommodate one-handed signs.
  • Bio-mechanical filtering of body data helps: SignBPoser trained on bio-mechanically filtered data (BP_f) outperformed the unfiltered variant (BP_u) across all four vertex subsets, with relative error reductions of 2.0% (FBody), 10.6% (UBody), 7.5% (UBody without head), and 11.1% (UBody excluding face).
  • Adding the body bio-mechanical loss during training slightly hurts: The BP_f+bio variant degraded relative to BP_f by 0.14% (FBody), 0.56% (UBody), 1.28% (UBody without head), and 0.53% (UBody excluding face), which the authors attribute to mild over-regularization.
  • Adding the body bio-mechanical loss during optimization helps: Retaining BP_f and adding the bio-mechanical loss at optimization time (not shown in the ablation table) produced the best results on all subsets, with relative error reductions of 0.17% (FBody), 0.37% (UBody), 0.05% (UBody without head), and 0.33% (UBody excluding face) compared to BP_f.
  • Correcting hand data helps: The hand prior trained on bio-mechanically corrected data (HP_f) outperformed the uncorrected variant (HP_u) on all metrics, with relative error reductions of 3.7% on Upper Body, 4.5% on Left Hand, and 6.2% on Right Hand.
  • Mixed effect of the hand bio-mechanical loss: HP_f+bio improved Upper Body (0.13%) and Left Hand (0.15%) relative to HP_f, but Right Hand performance degraded slightly (0.2%).
  • Qualitative behavior: On the signs Sonne, BesuchenEinmischen, and Muell, DexAvatar preserves fine-grained hand articulations and finger orientations and maintains structural consistency, whereas baselines such as PIXIE, OSX, PyMAF-X, SGNify, and EVA* produce misaligned wrists, unnatural limb orientations, or noticeable deviation from ground truth.
  • Ground truth is imperfect: The authors inspected the evaluation dataset and found some implausible hand poses in the ground truth itself.
  • Hyperparameter and evaluation details: The evaluation used 57 German signs and computed quantitative results only on the central portions of each sign, totaling 2,872 frames. Metrics are TR-V2V (mean vertex-to-vertex error) restricted to vertices above the pelvis, plus MPJPE and MPVPE for the priors separately.

Methodology in Plain English

DexAvatar does not train an end-to-end network. Instead it starts from off-the-shelf tools — SMPLerX and HaMeR for initial body and hand poses, Sapiens for body 2D keypoints, HaMeR for hand 2D keypoints, and existing camera estimates — and then refines those estimates by optimization.

The refinement minimizes a weighted sum of seven terms: a 2D joint reprojection loss, a body prior loss, a hand prior loss, an interpenetration penalty, a temporal consistency loss, a body bio-mechanical loss, and a hand bio-mechanical loss. The interpenetration penalty detects colliding face pairs with a bounding volume hierarchy and penalizes how deeply vertices intrude. Lower-body joints are excluded by setting their confidence weight to zero, because signing involves minimal lower-body motion. A Hand Decision Maker, using a sign classifier from prior work, detects one-handed signs and disables the non-dominant arm and hand so the optimizer does not introduce spurious updates.

The novelty is in the two priors. Both are encoder-decoder VAEs with 3 linear layers and an embedding size of 512, trained with a combination of KL divergence, reconstruction, mesh vertex, orthogonality, parameter regularization, and bio-mechanical losses, with loss weights 0.001, 0.999, 0.999, 0.01, 0.0001, and 1.5 for SignBPoser and 0.0001, 0.999, 0.999, 0.01, 0.0001, and 1.5 for SignHPoser, using Adam at a learning rate of 1e-3. The bio-mechanical loss is applied to 6 body joints for SignBPoser and 15 hand joints for SignHPoser, and the latent code is 33-dimensional.

Training data was cleaned before use, because both sources are noisy. Body data comes from the 3D data published by SignAvatars, reconstructed from How2Sign; frames are removed when shoulder, elbow/forearm, or wrist angles fall outside physiological ranges of motion or outside a torso-anchored "signer space" volume. Hand data comes from the authors' own mocap capture and is rectified by enforcing per-joint limits on bending, splaying, and twisting across the 15 hand joints, with axes aligned to match MANO.

DexAvatar is implemented in PyTorch and optimized with LBFGS, running on an NVIDIA RTX 4090 with 24 GB GPU memory and 64 GB CPU memory. The paper's supplementary material covers additional hyperparameter search, an ablation of SignHPoser against VPoser, and evaluations on motion blur, noisy, and self-occluded frames.

Why This Matters

Impact on research. The paper argues that domain-specific priors matter more than generic ones for sign language: VPoser and other general-purpose priors, trained on everyday human motion, can produce poses outside the space in which signs are actually produced. The authors state that the two pretrained priors can be plugged into any regression- or optimization-based approach for whole-body mesh recovery from sign language videos, which makes them reusable infrastructure rather than a single closed pipeline. The paper also surfaces a data problem — implausible hand poses in the ground truth of the benchmark it evaluates on.

Real-world applications (the paper's framing implies these; it does not enumerate them explicitly):

  • Generating 3D signing avatars for accessibility tools and automated sign language interfaces.
  • Producing 3D-annotated sign language datasets to replace the current video-and-2D-keypoints standard.
  • Animation and content production that needs anatomically plausible signing motion rather than generic hand gestures.
  • Linguistic and phonological analysis of sign, since the priors are designed to preserve phonologically meaningful hand shape and orientation variation.

Industry relevance. The work targets a communication need for approximately 466 million Deaf or hard-of-hearing individuals worldwide. Companies building sign language avatars, accessibility interfaces, or motion-capture-driven character pipelines could use the priors and the reconstruction pipeline; the authors also flag the cost of the problem, since the mocap data behind SignHPoser required a 9-camera Vicon rig and instrumented gloves.

Future Directions

  • Scale the training data. The authors state explicitly that future work will scale the training data behind the priors.
  • Strengthen the priors for broader coverage. The stated goal is to cover a broader range of signers and signing styles, addressing a limitation of data collected from a small, specific signer pool.
  • Improve ground truth quality. Since the authors found implausible hand poses in the ground truth of the evaluation dataset, reconstruction quality is partly capped by annotation and pseudo-ground-truth quality.
  • Extend the prior design. Because the ablation showed over-regularization effects (BP_f+bio degraded, and the hand bio-mechanical loss degraded the right hand), the balance between filtering, bio-mechanical constraints, and optimization-time regularization remains an open question.

Target Audience

Researchers and graduate students in computer vision working on 3D human mesh recovery and pose estimation; computer scientists building sign language generation, translation, or avatar systems; linguists and Deaf-community-facing technologists who need accurate 3D representations of signing; and practitioners in animation or motion capture who need domain-specific pose priors rather than generic human motion priors. Readers without background in SMPL-X, VAEs, or differentiable optimization will find the method sections demanding.

Authors’ abstract

The trend in sign language generation is centered around data-driven generative methods that require vast amounts of precise 2D and 3D human pose data to achieve an acceptable generation quality. However, currently, most sign language datasets are video-based and limited to automatically reconstructed 2D human poses (i.e., keypoints) and lack accurate 3D information. Furthermore, existing state-of-the-art for automatic 3D human pose estimation from sign language videos is prone to self-occlusion, noise, and motion blur effects, resulting in poor reconstruction quality. In response to this, we introduce DexAvatar, a novel framework to reconstruct bio-mechanically accurate fine-grained hand articulations and body movements from in-the-wild monocular sign language videos, guided by learned 3D hand and body priors. DexAvatar achieves strong performance in the SGNify motion capture dataset, the only benchmark available for this task, reaching an improvement of 35.11% in the estimation of body and hand poses compared to the state-of-the-art. The official website of this work is: https://github.com/kaustesseract/DexAvatar.

Read the original paper