Skip to content
AI.info

Research

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

Overview Research area: Physical fidelity evaluation for embodied AI — specifically, benchmarking numerical physics engines and generative video world models against real-world measurements. Technical

arXiv
2608.05948
Published
2026-08-06
Authors
Shuai Wang, Yaxin Feng, Xuekun Jiang, Shihan Tian, Ningyu Yan, Xing Shen, Chaoyang Lyu, Hui Wang, Yunsong Zhou, Hanqing Wang, Jiangmiao Pang, Yang Xiang, Xing Gao, Chunhua Shen, Weinan Zhang

AI summary

Overview

Research area: Physical fidelity evaluation for embodied AI — specifically, benchmarking numerical physics engines and generative video world models against real-world measurements.

Technical level: Intermediate. The paper assumes familiarity with physics simulation concepts (contact solvers, cloth/deformable solvers) and with video generation models, but its core argument — that looking realistic is not the same as being physically correct — is accessible without deep expertise.

Scope in one sentence: GAUGE is a real-world, measurement-grounded benchmark of 22 controlled task families that diagnoses how physics engines and video world models reproduce or deviate from measured physics, rather than how visually plausible their outputs appear.

What This Paper Is About

Physics engines are used to train and evaluate robots at scale, and generative video models are increasingly proposed as implicit simulators of future states. Existing evaluations of "physical realism" are typically done in isolation for each system type and lean on perceptual similarity or human judgment, which cannot say which physical principle or parameter was violated. GAUGE addresses this by building a single real-world experimental suite with motion-capture ground truth and calibrated physical parameters, then running two separate evaluation protocols on top of the same data — one for numerical simulators and one for video world models.

Key Contributions

  1. A unified real-world task suite. 22 controlled task families spanning rigid bodies, flexible cables, textiles, and volumetric deformable objects, containing approximately 1,560 motion-capture trials together with uncertainty estimates and task-specific physical observables.

  2. Cross-regime physical parameter annotations. To the authors' knowledge, GAUGE is the first real-world benchmark dataset to provide task-associated, experimentally characterized parameters spanning rigid-body contact, cloth constitutive response, and volumetric soft-body mechanics within a single standardized collection.

  3. Two complementary evaluation protocols. Using the same real-world experimental foundation, the paper separately benchmarks numerical physics engines and generative video world models. The engine protocol measures task-specific sim-to-real discrepancy; the world-model protocol separates agreement with the form of a physical law from parameter accuracy and temporal stability.

  4. A diagnostic comparison against prior benchmarks. Table 1 positions GAUGE against existing resources (Physics-IQ/Verified, IRIS, WorldBench, RIGIDBENCH, Cloth Sim-to-Real, PokeFlex, RGBench, FysicsEval), showing that prior work typically covers one mechanical regime, lacks task-paired real dynamic references, or lacks experimental characterization of parameters. Fluids are explicitly omitted from GAUGE's scope.

Main Findings

  • No engine is uniformly faithful. Isaac Sim (v6.0.0), Genesis (v1.12.0), and Newton (v1.3.0) each excel in different regimes, and none is consistently accurate across rigid, textile, and volumetric tasks. The largest discrepancies arise in impulsive contact, rapid textile motion, and volumetric deformation.

  • Simple contact and sliding are handled reasonably well. Isaac Sim gives the lowest errors for slope contact, nonsmooth contact, and turntable motion, with turntable normalized RMSE and DTW of 0.17 and 0.61. Genesis performs best on slope slider (0.58 RMSE, 0.69 DTW), while Newton is less accurate in that friction-dominated setting. Rankings depend on contact geometry and reference frame.

  • Demanding rigid-body events expose a large gap. Even the best bouncing-ball result is 15.63 times the baseline RMSE and 5.58 times the baseline DTW. On Newton's cradle, Isaac Sim and Newton produce zero longest stationary duration and normalized momentum-transfer efficiencies of only 0.20 and 0.26, and Genesis does not produce a valid rollout.

  • Matching low-frequency motion does not imply accurate impact or energy behavior. Isaac Sim and Newton obtain pendulum normalized periods of 1.10 and 1.09, but all engines accumulate energy error over the longer rollout, with raw energy-loss values ranging from −0.041 to 0.034.

  • Textile fidelity is regime-dependent. For textile stretching, Isaac Sim and Newton achieve normalized errors close to or below one; for textile bending, the best RMSE and DTW rise to 7.94 and 11.78. In textile flinging, Genesis reaches the lowest RMSE of 8.54 while Isaac Sim reaches 128.26 — quasi-static agreement does not predict behavior under rapid, spatially varying deformation.

  • Volumetric deformable bodies remain roughly an order of magnitude off. Genesis gives the lowest errors for stretching, shearing, and twisting, and Newton performs best for bending, but even the best deformable-body simulations show errors approximately one order of magnitude above the real-world baselines.

  • World models can get the equation form right and the physics wrong. Satisfying a law's structural form and recovering its parameters are distinct capabilities. For slope sliding, the lowest QFI comes from a different model for each material (Cosmos3-Super-I2V for wood, Seedance 2 for plastic, Cosmos3-Nano for metal), indicating limited consistency across materials.

  • Accelerations are systematically underestimated. For the wood slope slider, the best estimate of 2.06 m/s² is reasonably close to the measured 2.58 m/s², but for plastic and metal the closest reported accelerations are only 0.75 and 0.43 m/s² against real values of 2.57 and 2.67 m/s². On the bouncing ball, Cosmos3-Super-I2V with the negative prompt achieves the lowest QFI of 12.50 while inferring an acceleration of only 0.088 m/s²; Seedance 2's 1.84 m/s² is the closest among evaluated models yet still far below gravitational acceleration (baseline 9.81 m/s²).

  • Momentum transfer is hard. Six of the ten reported model configurations fail to produce a valid Newton's-cradle sequence, and the best momentum-transfer efficiency, 0.76 from Wan-2.2 with the negative prompt, stays below the real-world value of approximately one.

  • Oscillation can look right while timing is far off. Wan-2.2 and Genie 3 both reach an R² of 0.99 on the pendulum, but their periods are 1.93 s and 1.90 s rather than the measured 1.06 s. The closest period is 1.83 s from Wan-2.7, still about 73% longer than the baseline. Fitted damping and amplitude vary substantially across models, including different damping signs.

  • Prompts change physics results inconsistently. The negative prompt reduces Cosmos3-Super-I2V's bouncing-ball QFI from 270.69 to 12.50 and enables Wan-2.2 to produce a Newton's-cradle rollout, but it increases that model's wood slope-slider QFI from 13.61 to 569.36. Direction and magnitude of the change depend on both model and task, so the paper argues physics evaluation should use fixed prompt templates and report paired results.

Methodology in Plain English

The authors first build a physical laboratory rather than a dataset of internet videos. A 2 m × 2 m × 2 m motion-capture volume is instrumented with 16 NOKOV Mars9H infrared cameras arranged across three height levels, running at 180 Hz with sub-millimeter 3D localization accuracy, on a 0.9 m × 0.9 m black optical breadboard. Rigid objects carry 6 mm retroreflective spheres (about 0.5 g each) so a body-fixed frame can be recovered as 6-DoF pose; textiles and deformable objects carry 6 mm reflective adhesive markers whose connections form a tracked surface mesh. Recordings are truncated to valid intervals, downsampled to 30 fps, and 20 independent trials are collected per evaluation task, stored as JSON alongside calibrated properties.

Calibration is separate from evaluation. Friction and restitution for rigid bodies come from inclined-plane and collision tests; textile tensile, shear, and bending stiffness come from fabric tests conducted by Style3D; volumetric objects get Young's modulus and Poisson's ratio from tensile tests and Digital Image Correlation.

To evaluate engines, the team reconstructed matched scenes in Isaac Sim, Genesis, and Newton using the calibrated parameters, leaving all other settings at defaults to measure out-of-the-box behavior. Rigid bodies ran on PhysX (Isaac Sim), Genesis's native solver, and MuJoCo (Newton); textiles on Surface FEM, PBD, and VBD; volumetric bodies on FEM, explicit MPM, and implicit MPM. Simulated vertices do not correspond to physical markers, so the Hungarian algorithm assigns each marker to a nearby vertex or particle with matching error below 1 cm.

Because rigid bodies, cloth, and soft bodies have incompatible states, each rollout is converted into a "generalized trajectory": rigid-body position P(t) ∈ R³, textile marker Gaussian curvature K(t), or deformable mesh face areas A(t). Errors are computed as RMSE against the real mean trajectory and as dynamic time warping (DTW) distance, with the real-world within-trial RMSE and its per-trial standard deviation serving as the baseline that simulator errors are normalized against. For pendulum-like tasks where trajectory error is not enough, the paper adds Longest Stationary Duration and Momentum Transfer Efficiency for Newton's cradle, and Period Duration and Energy Loss for the simple pendulum.

For world models, the team kept to rigid-body tasks because current models struggle with 3D consistency and deformable dynamics. Six models received the same frontal initial frame and a standardized text prompt: five image-to-video models (Cosmos3-Nano, Cosmos3-Super-I2V, Wan-2.2, Wan-2.7, Seedance 2.0) plus the interactive world model Genie 3. SAM3 segments and tracks the target object, the binary mask centroid gives 2D pixel coordinates, and known object dimensions provide an image-to-world scale that converts the centroid sequence into a 2D experimental-coordinate trajectory. Evaluation uses Dynamic Error (deviation from Newton's second law), R² (how well the expected physical model explains the trajectory), and Quadratic Form Improvement (whether adding a quadratic term to the expected linear relation meaningfully reduces residual).

Inference setup: input images at 832 × 480 except Genie 3 at 1300 × 750; Cosmos3-Nano on a single NVIDIA GeForce RTX 4090, Cosmos3-Super-I2V and Wan-2.2 on eight RTX 4090 GPUs each; Wan-2.7 and Seedance 2.0 through remote APIs and Genie 3 through its online interface. Engine simulations ran on a Windows 11 workstation with a single RTX 4090, at 180 Hz for rigid and textile tasks and 900 Hz for 3D deformable tasks, with states recorded at 30 Hz.

Why This Matters

The paper's central claim is that physical fidelity cannot be characterized by visual quality, a single trajectory distance, or parameter estimation alone. High visual fidelity does not imply accurate dynamics, and a simulator that looks convincing while misrepresenting motion, contact, or deformation can encourage policies to exploit simulation artifacts or produce incorrect rankings of policy performance — directly undermining real-to-sim-to-real pipelines such as those validated by SIMPLER.

Real-world applications:

  • Robot policy training and evaluation. Benchmarks like this tell practitioners which engine is trustworthy for a given contact or deformation regime before they commit compute to large-scale policy training.
  • Simulator selection and calibration in industry. Teams choosing between Isaac Sim, Genesis, and Newton get per-task evidence about where each engine's error is concentrated, rather than a single aggregate number.
  • Evaluating and improving generative video world models. The separation of law-form agreement from parameter accuracy gives model developers a diagnostic signal about what to fix, not just a plausibility score.
  • Manufacturing and material simulation. The experimentally characterized cloth and soft-body parameters (from Style3D fabric tests and DIC-based modulus measurements) are relevant to textile and soft-goods simulation, where material properties must come from instrumented tests rather than fitted from the same trajectories used for evaluation.

Industry relevance: the benchmark sits at the intersection of embodied AI, digital twins, computer graphics, and generative video, and its authors span academic labs and industry research organizations. It provides a standardized comparison point for a market currently split between numerical simulation vendors and video-generation providers, both of which claim physical realism.

Future Directions

  • Broaden materials and parameter ranges. The current benchmark's materials and calibrated parameter ranges are limited; materials in the same nominal category can differ through surface treatment, internal structure, manufacturing process, temperature, and wear. Future versions should cover more materials and wider ranges of friction, stiffness, density, restitution, and damping.

  • Extend to fluids and coupled processes. The task set should grow to fluid-rigid and fluid-soft-body interactions to better represent the physical diversity embodied agents encounter. Fluids are currently outside GAUGE's scope.

  • Move the world-model track beyond 2D rigid-body trajectories. Two-dimensional image trajectories are insufficient for textiles and volumetric deformable bodies, whose states involve distributed deformation, depth variation, self-occlusion, and self-contact. Future metrics should operate on reconstructed point clouds, meshes, or tracked surface elements, measuring local strain, area and volume change, Gaussian and mean curvature, bending energy, geodesic distortion, self-contact, oscillation modes, damping, and energy dissipation.

  • Model perception uncertainty explicitly. Correspondence and reconstruction uncertainty must be modeled so that errors in 3D perception are not incorrectly attributed to the world model.

  • Standardize prompt protocols. Because prompt changes shift results in both directions depending on model and task, the paper argues for fixed prompt templates and paired reporting rather than conclusions from a single favorable generation.

Target Audience

Researchers and engineers working on embodied AI, robot learning, and real-to-sim-to-real pipelines; developers and evaluators of physics engines and soft-body/cloth solvers; and teams building or assessing generative video world models. It is also useful for practitioners in graphics, digital twins, and material simulation who need to know where simulated dynamics diverge from instrumented measurement, and for benchmark designers interested in separating physical-law form from parameter accuracy in evaluation.

Authors’ abstract

Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated. We introduce GAUGE, a real-world-grounded diagnostic benchmark for jointly evaluating how numerical simulators and generative video world models reproduce or deviate from real-world physics. It comprises 22 controlled task families covering rigid bodies, flexible cables, textiles, and volumetric deformable objects. Grounded in real-world trajectories and paired with calibrated physical metadata, uncertainty annotations, and task-specific observables, these tasks cover fundamental physical processes including collision, friction, momentum transfer, oscillation, self-contact, and deformation across diverse materials and conditions. We benchmark Isaac Sim, Genesis, and Newton on 14 task families using generalized trajectory errors, and evaluate 6 image-to-video models on 5 rigid-body tasks by testing physical-law consistency and the temporal stability of inferred parameters. Our results reveal no uniformly faithful physics engine, with the largest discrepancies arising in impulsive contact, rapid textile motion, and volumetric deformation. We further find that video world models can produce trajectories with the expected equation form while recovering incorrect accelerations, momentum transfer, and oscillation timing. GAUGE lays the groundwork for developing more physically faithful simulators and world models for embodied intelligence.

Read the original paper