Skip to content
AI.info

Research

SPRig: Self-Supervised Pose-Invariant Rigging from Mesh Sequences

Overview Research area: Computer Vision / 3D character animation — automatic character rigging (skeleton and skinning estimation) from dynamic, animated mesh sequences. Technical level: Advanced. The

SPRig: Self-Supervised Pose-Invariant Rigging from Mesh Sequences
arXiv
2602.12740
Published
2026-02-13
Authors
Ruipeng Wang, Langkun Zhong, Miaowei Wang

AI summary

Overview

Research area: Computer Vision / 3D character animation — automatic character rigging (skeleton and skinning estimation) from dynamic, animated mesh sequences.

Technical level: Advanced. The paper assumes familiarity with autoregressive Transformer decoders, token-space and geometry-space losses, rigid Procrustes alignment, knowledge distillation, and Linear Blend Skinning.

Scope: The paper proposes SPRig, a teacher–student fine-tuning framework that turns existing single-frame (static) rigging models into pose-invariant, temporally consistent riggers by using unlabeled animated mesh sequences as a self-supervised signal.

What This Paper Is About

State-of-the-art automatic rigging models such as UniRig and Puppeteer are built for static assets and assume a predefined canonical rest pose (a standard T-pose). Dynamic mesh sequences — such as those in DyMesh or DT4D — have no such canonical T-pose, so running these models independently frame by frame produces temporally inconsistent rigs, visible as topological flickering, joint drift, and erratic changes in surface connectivity.

SPRig addresses this by fine-tuning a pretrained static rigging model so that its output on every frame of a sequence matches the rig it predicts on one automatically chosen canonical anchor frame, enforcing pose invariance without any ground-truth skeleton or skinning labels.

Key Contributions

  1. SPRig framework: A self-supervised fine-tuning framework that improves temporal stability and pose invariance in state-of-the-art rigging models using only unlabeled mesh sequences. It is built on a frozen-teacher / trainable-student paradigm and covers both skeleton and skinning generation.

  2. Dual consistency losses for skeleton generation: Token-space consistency regularization (with an anchor-frame self-loss and a cross-frame loss under teacher forcing, weighting parent tokens at w_i = 5 versus w_i = 1 for other tokens) combined with a geometry-space loss that uses rigid Procrustes alignment to remove global motion, with three terms: bone direction, bone-length distribution, and joint-endpoint Chamfer mismatch.

  3. Articulation-invariant skinning loss: A consistency-distillation objective (masked symmetric KL, masked L1, and an anchor-frame loss) modulated by a soft-support mask that suppresses teacher noise, plus structural regularization through an entropy penalty and a geometric proximity prior averaged over a sliding window of 3 frames.

  4. Automatic anchor-frame selection: Instead of using the conventional first frame, SPRig selects the anchor as the frame with maximum total surface area, c = argmax_k A(M^k), to avoid self-occluded or curled initial poses.

Main Findings

  • Large reduction in skeleton jitter: On DT4D, SPRig reduces pairwise joint distance deviation (PJDD) from 17.46 (Puppeteer) and 15.76 (UniRig) to 0.68, while GSD falls from 0.062 and 0.060 to 0.056. On DyMesh, PJDD goes from 16.53 (Puppeteer) / 19.56 (UniRig) to 0.72, and GSD from 0.068 / 0.071 to 0.054.

  • Improved skinning temporal stability and fidelity: On DT4D, temporal L1 Error drops from 1328.80 (Puppeteer) and 1310.31 (UniRig) to 982.35, with LBS RMSE improving from 0.007560 / 0.007576 to 0.007552. On DyMesh, L1 Error drops from 1379.32 / 1430.43 to 936.58 and LBS RMSE from 0.007835 / 0.007763 to 0.007447.

  • Static per-frame quality is preserved or improved: On out-of-domain static test sets, SPRig gives lower CD-J2J, CD-J2B, and CD-B2B than both baselines. On Articulation-XL 2.0: 0.0270 / 0.0213 / 0.0188 versus 0.0311 / 0.0237 / 0.0198 (Puppeteer) and 0.0331 / 0.0261 / 0.0218 (UniRig). On ModelsResource: 0.0325 / 0.0258 / 0.0235 versus 0.0377 / 0.0280 / 0.0241 and 0.0396 / 0.0302 / 0.0257. On Diverse-Pose: 0.0191 / 0.0178 / 0.0125 versus 0.0251 / 0.0199 / 0.0160 and 0.0325 / 0.0257 / 0.0208.

  • Skinning gains on the challenging Diverse-pose set: SPRig reports Precision 0.842, Recall 0.734, and Avg. L1 Dist. 0.378, compared with Puppeteer's 0.836 / 0.722 / 0.405 and RigNet's 0.747 / 0.654 / 0.746. On Articulation-XL 2.0 the paper reports 0.863 / 0.745 / 0.355 and on ModelsResource 0.732 / 0.883 / 0.462, which the authors describe as competitive rather than uniformly better.

  • Both loss spaces are necessary: Removing the token loss raises PJDD from 0.68 to 1.20; removing the geometry loss raises it to 1.38 and GSD from 0.056 to 0.067; removing the parent-token weighting (w_i = 5) raises PJDD to 1.36 and GSD to 0.073. Qualitative ablations show missing limbs, broken hierarchical connectivity, and poor geometric alignment respectively.

  • Skinning loss components are critical: Removing the symmetry loss raises temporal L1 Error to 3897.24 and LBS RMSE to 0.016587; removing the L1 loss gives 4007.13 / 0.023254; removing the anchor-frame loss gives 4006.53 / 0.023305; removing the entropy loss gives 4007.74 / 0.023311, all versus the default 982.35 / 0.007552. These variants were trained for 24 epochs. The ablation for removing the geometric prior loss is truncated in the provided content.

  • Automatic anchor selection helps: Table 4 reports that using the automatic anchor rather than the first frame improves skeleton PJDD from 0.68 to 0.59, GSD from 0.056 to 0.046, and CD-B2B from 0.0188 to 0.0175, and skinning L1 Error from 1001.56 to 982.35 and LBS RMSE from 0.007558 to 0.007552. The skeleton ablation table separately lists the default configuration at PJDD 0.68 and GSD 0.056, which differs from the anchor-frame table's default row.

  • Efficient convergence: Fine-tuning converges within 10 to 12 epochs, completing in about 20 hours on a single NVIDIA A100 GPU.

Methodology in Plain English

The core idea is that an animated sequence depicts one object, so it should have one single rig that stays the same no matter what pose the object is in. This gives a free supervisory signal: pick one frame as the reference, ask a frozen pretrained model to rig it well, then teach a copy of that model to reproduce the same rig on all the other frames.

Stage 1 — Skeleton. The mesh from each frame is sampled as a point cloud and fed to an autoregressive decoder that emits a token stream grouped into per-joint quadruples of x, y, z coordinates and a parent index, for a total of L = 4J tokens. The frozen teacher runs only on the anchor frame and produces a canonical token sequence. During fine-tuning, the student is always given the anchor's prefix tokens as context (teacher forcing) and is scored on both the anchor frame itself and every other frame, with parent tokens weighted more heavily so that kinematic topology stays stable. Because token-level supervision alone can drift, a second loss evaluates the decoded skeletons after rigidly aligning each frame to the anchor with Procrustes analysis, comparing bone directions, sorted bone-length distributions, and joint endpoints via one-sided Chamfer distance.

Stage 2 — Skinning. N points are sampled on the canonical mesh by triangle-area importance sampling and tracked through every frame using the same face index and barycentric coordinates, yielding point position and normal features. The frozen teacher predicts skinning weights on the canonical frame only; a trainable tri-stream Transformer student conditioned on the anchor skeleton, a valid-joint mask, and joint-tree shortest-path distances predicts weights on every frame. A soft support mask keeps only the top-K_s joints per point fully active while leaving a small residual on the rest, filtering out negligible teacher noise. The student is pushed toward the masked teacher weights by a masked symmetric KL, a masked L1, and an anchor-frame loss, and is regularized by an entropy penalty that sharpens assignments and a geometric proximity prior that prefers bones near each surface point.

Anchor selection. For each frame the total triangle surface area is computed and the most extended pose is chosen as the anchor, since it most closely resembles a T-pose and exposes the most surface area. All meshes are then normalized using the anchor's axis-aligned bounding box.

Evaluation setup. The method is instantiated by fine-tuning pretrained Puppeteer models — freezing the geometry point cloud encoders and updating only the Transformer decoders — in PyTorch with AdamW. The skeleton network uses a learning rate of 3×10⁻⁶ and batch size 20; the skinning network uses 1×10⁻⁵ and batch size 12 with automatic mixed precision. Dynamic evaluation uses a 125-sequence validation set built by sampling one sequence per object category from DeformingThings4D (DT4D) and a curated subset of 1007 DyMesh motion sequences. Dynamic metrics are PJDD and GSD for skeletons, and temporal L1 Error and LBS RMSE for skinning; static metrics CD-J2J, CD-J2B, CD-B2B, Precision, Recall, and Avg. L1 Dist. are evaluated on Articulation-XL 2.0, ModelsResource, and Diverse-pose.

Why This Matters

Research impact. The paper reframes unlabeled dynamic mesh sequences — which are abundant thanks to 4D reconstruction and generative mesh models — as a self-supervision signal, rather than treating them as a reason to train a new rigging model from scratch. It shows that temporal consistency can be injected into an existing static foundation model by fine-tuning, and that doing so does not cost static accuracy. It also introduces dynamic evaluation metrics (PJDD, GSD, temporal L1 Error on skinning weights) for a setting where static rigging benchmarks do not apply.

Real-world applications:

  • Game development pipelines, where flickering rigs break animation tooling and require manual cleanup.
  • Film and visual effects, where rigging is described as a fundamental but time-consuming bottleneck.
  • Virtual reality content creation.
  • 4D video reconstruction and real-world 4D scanning captures, which produce sequences that lack a canonical T-pose or a matching natural pose.

Industry relevance. Automatic rigging tools exist to accelerate content creation, but their T-pose assumption does not match how dynamic capture and generative data actually arrive. A fine-tuning step that can be layered on top of an existing pretrained rigger — with roughly 20 hours on one A100 and an average of 10 to 12 epochs — is a practical retrofit rather than a replacement. The authors state the code is available in the supplemental material and will be made publicly available upon publication.

Future Directions

  • Extending beyond explicit mesh sequences. Related work cited by the paper (MoRig for point cloud sequences, LASR for monocular video, Reacto for casual videos) suggests applying the same anchor-based consistency idea to inputs that are not clean, vertex-corresponding meshes.
  • Reducing dependence on a pretrained teacher. The framework relies on a frozen model producing a high-quality rig on the anchor frame; what happens when the teacher itself is weak on that frame is not resolved beyond the soft support mask, and the soft support mask, top-K_s, γ, and β choices are not fully ablated in the provided content.
  • Rethinking anchor selection. The current criterion is maximum total surface area; whether other geometric criteria produce better anchors across topologies is an open question raised by the anchor ablation.
  • Closing gaps in the reported ablations. The skinning ablation table is truncated in the provided content for the geometric prior loss, and the skeleton default configuration appears with different PJDD/GSD values in the skeleton ablation table versus the anchor-frame ablation table — a point that would benefit from clarification.

Target Audience

Researchers and engineers working on 3D character rigging, skeletal animation, and skinning weight prediction; practitioners in game, film, VR, and 4D reconstruction pipelines who need temporally stable rigs from captured or generated mesh sequences; and machine learning researchers interested in self-supervised fine-tuning, teacher–student distillation, and enforcing temporal consistency in models initially designed for static inputs.

Authors’ abstract

State-of-the-art rigging methods typically assume a predefined canonical rest pose. However, this assumption does not hold for dynamic mesh sequences such as DyMesh or DT4D, where no canonical T-pose is available. When applied independently frame-by-frame, existing methods lack pose invariance and often yield temporally inconsistent topologies. To address this limitation, we propose SPRig, a general fine-tuning framework that enforces cross-frame consistency across a sequence to learn pose-invariant rigs on top of existing models, covering both skeleton and skinning generation. For skeleton generation, we introduce novel consistency regularization in both token space and geometry space. For skinning, we improve temporal stability through an articulation-invariant consistency loss combined with consistency distillation and structural regularization. Extensive experiments show that SPRig achieves superior temporal coherence and significantly reduces artifacts in prior methods, without sacrificing and often even enhancing per-frame static generation quality. The code is available in the supplemental material and will be made publicly available upon publication.

Read the original paper