Research
RigAnyFace: Scaling Neural Facial Mesh Auto-Rigging with Unlabeled Data
Overview Research area: Computer Vision and Computer Graphics, specifically 3D facial mesh auto-rigging (automatic generation of blendshape rigs from a neutral mesh) using neural surface learning and

- arXiv
- 2511.18601
- Published
- 2025-11-23
- Authors
- Wenchao Ma, Dario Kneubuehler, Maurice Chu, Ian Sachs, Haomiao Jiang, Sharon Xiaolei Huang
AI summary
Overview
Research area: Computer Vision and Computer Graphics, specifically 3D facial mesh auto-rigging (automatic generation of blendshape rigs from a neutral mesh) using neural surface learning and 2D supervision.
Technical level: Advanced. The paper assumes familiarity with mesh learning networks, the Laplace-Beltrami operator, FACS blendshape systems, differentiable rendering, and generative 2D animation models.
One-sentence scope: The paper introduces RigAnyFace (RAF), a neural framework that deforms neutral facial meshes of arbitrary topology, including meshes with multiple disconnected components such as eyeballs and teeth, into FACS-driven blendshape rigs, and scales its training using unlabeled meshes through a 2D supervision strategy.
What This Paper Is About
Facial rigging, the process of making a static face mesh animatable, normally requires skilled artists tens of hours per asset. Existing auto-rigging methods mostly transfer blendshapes from a predefined template mesh, which limits accuracy when the target shape differs from the template and generally cannot handle meshes with multiple disconnected parts. The goal of RAF is to rig any neutral facial mesh directly from explicitly controllable FACS parameters, without a template and without requiring the mesh to be a single connected humanoid head, while training on far more data than the small number of expensively hand-rigged assets would allow.
Key Contributions
- A template-free neural auto-rigging framework (RAF) built on DiffusionNet that deforms a neutral mesh into 96 FACS poses (48 FACS poses plus 48 corrective poses) conditioned on a one-hot-like FACS vector, producing a linear blendshape rig.
- Two architecture modifications to DiffusionNet: a conditional diffusion block that injects the FACS vector (concatenated with a global feature) into each block, and a global encoder (a smaller 2-layer DiffusionNet with global average pooling) that captures holistic mesh information across disconnected components.
- A 2D supervision strategy that scales training to unlabeled neutral meshes, combining appearance supervision (rendered image and binary mask losses) with motion supervision (a differentiable 2D displacement loss analogous to optical flow), generated using a MegActor-based 2D face animation model, RAFT flow estimation, and image segmentation.
- A curated dataset of artist-crafted facial meshes with multiple disconnected components, including a subset rigged by professional artists, plus a UV-layout-based head interpolation augmentation strategy to increase dataset size.
Main Findings
- Full model accuracy on artist-crafted data: With both 2D and 3D supervision and the global encoder, MAE is 1.92 mm and MAE Q95 is 5.63 mm on the full set of 96 FACS poses.
- Ablation results: Removing the global encoder raises MAE to 2.14 mm and MAE Q95 to 6.64 mm; removing the 2D loss gives 2.08 mm / 5.84 mm; removing unrigged data gives 2.01 mm / 5.81 mm; removing the 2D displacement loss gives 1.95 mm / 5.89 mm.
- Comparison against baselines: On 12 artist-annotated humanoid heads, RAF achieves 1.01 mm MAE and 2.94 mm MAE Q95, versus NFR at 2.77 mm / 7.21 mm and Deformation Transfer at 2.93 mm / 8.41 mm. Deformation Transfer requires additional inputs (an exemplar expression mesh and user-annotated correspondences).
- Penetration reduction: With the global encoder, penetration between inner components and the outer face surface drops to 0.166 (without "all other components") and 0.173 (with them), compared to 0.377 and 0.405 without the global encoder.
- Global features encode component layout: t-SNE visualization of global features from samples with randomly perturbed or removed components forms separate clusters from the original samples.
- In-the-wild generalization: Qualitative results on meshes from ICT FaceKit, Objaverse, and CGTrader show better accuracy and generalizability than NFR; although NFR was trained on ICTFaceKit data and RAF was not, RAF's results there are comparable to NFR's, and RAF generalizes to a non-humanoid head that NFR leaves largely undeformed.
- Efficiency: The model has 5.4M parameters, trains in about 2 days on 8 NVIDIA A100 GPUs, and generates a rig in 8.72s on an Apple M2 Max CPU and 3.1s on an Nvidia T4 GPU on the test set (1,750 vertices and 3,362 faces on average).
- NFR preprocessing dependence: NFR requires keeping only the largest connected component and trimming inner-lip and eyelid surfaces; without this, retaining multiple components causes self-penetration and untrimmed inner-lip surfaces cause artifacts. RAF needs no such preprocessing.
Methodology in Plain English
The team starts with a neutral face mesh and a FACS vector that lists which action units are active. A neural network predicts a per-vertex displacement that moves the neutral mesh into the pose described by that vector; doing this for all 96 poses yields a full blendshape rig.
The backbone is DiffusionNet, chosen because its heat-diffusion operation depends only on surface geometry, so the same weights work on meshes with different resolutions and triangulations. Because diffusion cannot jump between disconnected pieces like eyeballs and teeth, the authors add a global encoder that compresses the whole mesh into a single feature vector, which is then combined with the FACS vector and injected into every diffusion block.
Training happens in two stages. Stage one trains on both rigged and unrigged heads using only 2D supervision: a rendered image loss and mask loss for visible appearance changes, plus a 2D displacement loss computed from vertex deformations through differentiable rendering, which catches subtle poses like "Jaw Left" that barely change pixel colors. For unrigged heads, the 2D ground truth is manufactured: a fine-tuned MegActor-based animation model transfers an expression from a rendered rigged head onto the neutral unrigged head, a segmentation model produces masks, and RAFT estimates pixel offsets. Stage two fine-tunes only on rigged heads with both 2D losses (image, mask, landmarks, eye closure) and a 3D MSE loss on vertex positions.
Dataset expansion uses a standardized UV layout so vertices correspond across heads, allowing linear interpolation between head geometries: 137 rigged heads are interpolated by a factor of 25 to give 2,929 samples, and 137 unrigged heads by a factor of 50 to give 5,457 samples (8,386 total in stage one).
Why This Matters
For research, RAF shows that 2D supervision from generative animation and flow models can substitute for scarce 3D rigging ground truth, and that conditioning a triangulation-agnostic surface network on FACS parameters plus a global feature can handle disconnected geometry that previously broke both diffusion-based and template-transfer methods.
Real-world applications the authors demonstrate or describe:
- User-controlled animation, where a user edits FACS parameters to pose a mesh.
- Video-to-mesh retargeting, transferring a subject's tracked FACS expressions onto an unrigged mesh.
- Animating a facial mesh produced by a text-to-3D model, turning a static neutral mesh into an animatable avatar.
- Lowering the barrier for small studios, educators, and assistive-tech developers to create expressive avatars, as noted in the broader impact discussion.
Industry relevance is direct: the work is a collaboration with Roblox and targets avatar pipelines where asset variety, non-humanoid and stylized heads, and eyeball or teeth components are common, and where manual rigging cost is a bottleneck. The paper also notes the dual-use risk that the same ease of use can lower the barrier for deepfake production, arguing for dataset curation, usage licenses, and watermarking.
Future Directions
- Extending the dataset to cover mesh structures that deviate significantly from the training data, such as shell-like meshes lacking the fine geometric detail needed for high-quality animation.
- Improving robustness to poor discretization, where the main facial mesh fragments into multiple disconnected components that the model fails to keep spatially coherent, for instance by using a diffusion operator defined on a high-quality background triangulation.
- Establishing quantitative evaluation for unrigged in-the-wild heads: the paper reports only qualitative results there because 3D ground truth is unavailable.
- Reducing reliance on the quality of the 2D supervision pipeline, which depends on a fine-tuned MegActor animation model, RAFT flow estimation, and segmentation, and which was fine-tuned on a small set of rigged heads to handle stylized faces.
Target Audience
Computer vision and graphics researchers working on 3D generative models, mesh deformation, and character animation; technical artists and rigging pipeline engineers in the game and film industries; and graduate students with a background in geometry processing or surface neural networks who want to understand how 2D generative supervision can be used to scale a 3D deformation task.
Authors’ abstract
In this paper, we present RigAnyFace (RAF), a scalable neural auto-rigging framework for facial meshes of diverse topologies, including those with multiple disconnected components. RAF deforms a static neutral facial mesh into industry-standard FACS poses to form an expressive blendshape rig. Deformations are predicted by a triangulation-agnostic surface learning network augmented with our tailored architecture design to condition on FACS parameters and efficiently process disconnected components. For training, we curated a dataset of facial meshes, with a subset meticulously rigged by professional artists to serve as accurate 3D ground truth for deformation supervision. Due to the high cost of manual rigging, this subset is limited in size, constraining the generalization ability of models trained exclusively on it. To address this, we design a 2D supervision strategy for unlabeled neutral meshes without rigs. This strategy increases data diversity and allows for scaled training, thereby enhancing the generalization ability of models trained on this augmented data. Extensive experiments demonstrate that RAF is able to rig meshes of diverse topologies on not only our artist-crafted assets but also in-the-wild samples, outperforming previous works in accuracy and generalizability. Moreover, our method advances beyond prior work by supporting multiple disconnected components, such as eyeballs, for more detailed expression animation. Project page: https://wenchao-m.github.io/RigAnyFace.github.io