Skip to content
AI.info

Research

MotionPhysics: Learnable Motion Distillation for Text-Guided Simulation

Overview Research area: Computer Vision and Computer Graphics — specifically physics-based 3D dynamic simulation, differentiable rendering, video diffusion models, and multimodal large language models

MotionPhysics: Learnable Motion Distillation for Text-Guided Simulation
arXiv
2601.00504
Published
2026-01-01
Authors
Miaowei Wang, Jakub Zadrożny, Oisin Mac Aodha, Amir Vaxman

AI summary

Overview

Research area: Computer Vision and Computer Graphics — specifically physics-based 3D dynamic simulation, differentiable rendering, video diffusion models, and multimodal large language models.

Technical level: Advanced. The paper assumes familiarity with Gaussian Splatting, the Material Point Method (MPM), score distillation sampling (SDS), and video diffusion models.

Scope: A text-guided, end-to-end differentiable framework that automatically infers physical material parameters for 3D scenes so that a simulator reproduces the dynamics described in a natural language prompt, with no ground-truth trajectories or annotated videos required.

What This Paper Is About

Simulating how a 3D object deforms — bouncing, splashing, stretching, flowing — normally requires an expert to hand-tune physical parameters such as density, Young's modulus, and yield stress through slow trial and error. MotionPhysics replaces that manual process: given a 3D scene and a sentence like "A rubber ball-like object hits the ground and bounces back," it estimates a plausible full set of material parameters automatically and produces a physically grounded animation. The central difficulty the authors target is that pretrained video diffusion models — the usual source of "how things should move" supervision — entangle motion with appearance and geometry, so they give wrong answers for AI-generated, human-designed, or newly viewed real objects.

Key Contributions

  1. A learnable motion distillation (LMD) loss. A lightweight, trainable two-layer convolutional motion extractor isolates pure motion signals from a pretrained video diffusion model while suppressing the model's appearance and geometry inductive biases.

  2. Constraint-aware LLM initialization. Rather than naively querying GPT-4 for material values, the authors prompt it with domain-specific parameter ranges drawn from material-property handbooks (for example, Young's modulus of 10⁷–4×10¹¹ Pa for elastic materials and 10⁶–5×10⁶ Pa for plasticine), which anchors estimates in physically plausible values and suppresses hallucination.

  3. A fully automatic, text-guided simulation system. MotionPhysics combines Gaussian Splatting reconstruction, a differentiable GS-adapted MLS-MPM simulator, and the LMD objective into one pipeline that requires no ground-truth dynamics, motion-capture markers, or videos.

  4. Broad evaluation across more than thirty scenarios. The method is tested on elastic materials, plasticine, metals, foams, sands, and both Newtonian and non-Newtonian fluids, spanning real-world scans, human-designed assets, and AI-generated meshes, and it is shown to extend to heterogeneous multi-object scenes.

Main Findings

  • User preference favors MotionPhysics on both metrics. In a two-alternative forced-choice study with 79 participants comparing 15 randomly selected video pairs, preferences exceeded 50% across all datasets and exceeded 80% on human-designed and AI-generated scenes. Against PhysDreamer, realism preference was 96.77% (human-designed), 77.63% (real world), and 93.80% (AI-generated); prompt-adherence preference was 95.88%, 86.08%, and 96.45%. Against DreamPhysics: 80.27%, 69.77%, 94.65% realism and 82.35%, 75.05%, 93.80% adherence. Against OmniPhysGS: 97.62%, 92.73%, 84.60% realism and 96.48%, 91.43%, 80.75% adherence. Against PhysFlow: 90.30%, 66.13%, 84.20% realism and 91.30%, 71.53%, 85.60% adherence.

  • Best objective metric scores with competitive runtime. On the Bird scene, MotionPhysics attained Overall Consistency (OC ×10⁻²) of 18.18 versus 17.00 (PhysDreamer), 18.02 (DreamPhysics), 17.10 (OmniPhysGS), and 17.96 (PhysFlow); CLIPSIM (×10⁻²) of 21.69 versus 21.62, 21.64, 21.27, and 21.32; and ECMS (lower is better) of 11.37 versus 27.48, 15.76, 13.70, and 13.07. Optimization times were 18.39 min for MotionPhysics, 16.48 min for PhysDreamer, 18.20 min for DreamPhysics, approximately 9 h for OmniPhysGS, and 18.14 min for PhysFlow.

  • LMD outperforms alternative supervision losses. Against optical flow loss and SDS loss, models optimized with the LMD objective aligned most closely with PAC-NeRF's manually tuned reference behavior on the Playdoh (plasticine) and Bird (elastic) scenes.

  • Initialization matters substantially. In the Hat scene, PhysFlow-style initialization produced exaggerated early deformations and unrealistic artifacts even after optimization, while median-value initialization with a large Young's modulus of 2×10¹¹ produced overly rigid behavior with minimal deformation. The constrained-range initialization gave a stable starting point.

  • Robustness to changed conditions. With a different force direction and a shifted viewpoint, MotionPhysics still produced the prompted vibration of a telephone cord while PhysFlow remained nearly static. When a plane's propeller rotation speed was increased fivefold, MotionPhysics preserved structural integrity and yield stress, whereas PhysFlow detached the propeller.

  • Ablation scores. The reported OC (×10⁻²) values in the ablation table are 18.04, 18.05, 18.13, and 18.18, with the complete method achieving the highest value.

  • Heterogeneous materials work. By combining GS segmentation with SAM2 and a multimodal LLM, per-object material parameters can be inferred from a single text prompt; in the reported example the axe was treated as metal and the toy as elastic rubber in the MPM simulation, which assigns parameters per particle.

  • Stated limitation. The method does not model shadow effects, and the estimated parameters enable plausible simulations but are not intended for accurate real-world material measurement.

Methodology in Plain English

The pipeline has three stages.

1. Reconstruction. The input scene — multi-view images, video, a single image, or a mesh — is converted into a collection of 3D Gaussians using methods such as PGSR, GIC, Splat3R, or GaMes. Each Gaussian stores a center, opacity, covariance, and color coefficients.

2. LLM-based initialization. A multimodal LLM (GPT-4) is prompted with the user's text and, secondarily, a reference rendered image, to predict the material class, density, and class-specific coefficients. Crucially, the prompt template contains handbook-derived value ranges for each material type, so the model selects from realistic intervals instead of inventing numbers. This initial guess is important because poor starting values waste computation and hinder convergence.

3. Optimization with motion distillation. A differentiable Gaussian-Splatting-adapted MLS-MPM simulator (built on NVIDIA WARP) evolves the splats over time under external forces, and rendered frames are passed to a pretrained video diffusion model (CogVideoX, with classifier-free guidance of 100). The key insight is that the same simulated motion produces globally consistent latent codes in the video encoder even when texture or shape changes — only local details differ. So the authors add Gaussian perturbations to the Gaussians' centers and colors, compute one-step denoised latents, and train a small two-layer convolutional "motion extractor" (initialized as the identity map, learning rate 2×10⁻⁵, kept synchronized by exponential moving averaging) to predict the same motion target from both perturbed and denoised versions. The Charbonnier-form loss with β = 10⁻³ then backpropagates gradients to the material coefficients.

Implementation details: each simulation spans 5 seconds and yields 150 rendered frames with 256 internal substeps per frame; a frame-boosting scheme splits frames into M = 8 interleaved subsequences with 256 × M intermediate updates between adjacent frames; training converges in roughly 40 iterations at about 28 seconds per forward–backward pass on an NVIDIA A100 80 GB GPU.

Evaluation: three dataset groups — eight human-designed PAC-NeRF models (Torus, Bird, Playdoh, Cat, Trophy, Droplet, Letter Cream, Toothpaste) using GIC reconstructions; real-world scenes (PhysDreamer's Alocasia, Carnation, Hat, and Telephone, plus Fox, Plane, Kitchen, Jam, and Sandcastle); and four AI-generated meshes (Urchin, Alien, Gentleman, Axe) from Hunyuan3D processed with GaMes. Baselines are PhysDreamer, DreamPhysics, PhysFlow, and OmniPhysGS, run with identical simulation settings.

Why This Matters

Impact on research. The paper shows that motion supervision can be decoupled from appearance and geometry supervision in diffusion-guided simulation, and that LLM knowledge becomes far more useful when bounded by physical constraints. It also provides evidence that current automatic metrics (OC, CLIPSIM, ECMS) do not fully track human perception of physical plausibility, pointing to a need for better evaluation in physics-grounded generation.

Real-world applications:

  • Animation and visual effects pipelines, where artists could describe a material's behavior in text instead of hand-tuning simulator parameters.
  • Video game and interactive content development, where props and destructible objects need plausible responses to forces without per-asset parameter authoring.
  • Robotic and embodied simulation, where text-described material properties could seed contact and manipulation scenarios.
  • Digital twins and product prototyping, where different material choices (foam, metal, plasticine) could be compared through language-driven simulation rather than physical measurement.

Industry relevance. The work is positioned as an accessible alternative to expert-driven parameter tuning, the authors acknowledge funding from Huawei Technologies Co. Ltd., and the code and project page are released publicly, lowering the barrier for companies that already hold 3D assets but lack simulation specialists.

Future Directions

  • Supporting fully automatic configuration of fine-grained simulations from text, as stated in the conclusion as the authors' next aim.
  • Extending the motion distillation loss to other animation tasks, specifically character rigging and deformation.
  • Incorporating shadow effects, which the authors list as a current omission affecting visual realism.
  • Developing evaluation protocols or ground-truth distributions of physically plausible outcomes that better correlate with human judgment, since the reported OC, CLIPSIM, and ECMS scores sometimes disagreed with perceived correctness (for example, the Droplet water-splashing scene scored slightly lower than some baselines).
  • Scaling heterogeneous, multi-object simulation further; the authors note that adapting the pipeline to cinematics or video games requires engineering beyond the scope of this work, and multi-object code for OmniPhysGS was unavailable for comparison.

Target Audience

Researchers and practitioners in 3D vision, computer graphics, and physically based animation who work on differentiable simulation, Gaussian Splatting, video diffusion priors, or physics-grounded generation. It is also relevant to technical artists and simulation engineers seeking automatic material parameter estimation, and to readers interested in how LLM priors can be constrained by domain knowledge to avoid hallucinated values.

Authors’ abstract

Accurately simulating existing 3D objects and a wide variety of materials often demands expert knowledge and time-consuming physical parameter tuning to achieve the desired dynamic behavior. We introduce MotionPhysics, an end-to-end differentiable framework that infers plausible physical parameters from a user-provided natural language prompt for a chosen 3D scene of interest, removing the need for guidance from ground-truth trajectories or annotated videos. Our approach first utilizes a multimodal large language model to estimate material parameter values, which are constrained to lie within plausible ranges. We further propose a learnable motion distillation loss that extracts robust motion priors from pretrained video diffusion models while minimizing appearance and geometry inductive biases to guide the simulation. We evaluate MotionPhysics across more than thirty scenarios, including real-world, human-designed, and AI-generated 3D objects, spanning a wide range of materials such as elastic solids, metals, foams, sand, and both Newtonian and non-Newtonian fluids. We demonstrate that MotionPhysics produces visually realistic dynamic simulations guided by natural language, surpassing the state of the art while automatically determining physically plausible parameters. The code and project page are available at: https://wangmiaowei.github.io/MotionPhysics.github.io/.

Read the original paper