Research
World Models' Last Exam in Physics
Overview Research area: Computer vision / video generation evaluation, at the intersection of generative video models and physics-grounded benchmarking for embodied AI. Technical level: Advanced. The

- arXiv
- 2610.08791
- Published
- 2026-10-06
- Authors
- Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin, Ziming Qin, Zheng Jiang, Wenyi Li, Calvin Xiao, Youjie Zheng, Kaisen Yang, Qinhuai Na
AI summary
Overview
Research area: Computer vision / video generation evaluation, at the intersection of generative video models and physics-grounded benchmarking for embodied AI.
Technical level: Advanced. The paper assumes familiarity with image-to-video generation, vision-language models (VLMs), tracking and segmentation tools, and quantitative physical measurement.
Scope: This paper introduces World Models' Last Exam in Physics, a measurement-based benchmark of 40 controlled tasks across nine physical categories that scores video world models by directly measuring physical quantities in generated videos rather than by VLM judgment or reference-video comparison.
What This Paper Is About
Video world models can produce visually convincing but physically inconsistent sequences, which undermines their use for prediction and planning in embodied AI. Existing evaluations rely either on VLM judgments (limited by the evaluator's own physical reasoning) or on comparisons to reference videos (sensitive to reference quality and alignment), while direct physical tests largely focus on mechanics. The paper's goal is to test observable physical relationships across many domains using task-specific measurements, without requiring reference videos or simulator ground truth.
Key Contributions
-
A measurement-based benchmark of 40 controlled video generation tasks covering mechanics, optics, fluid behavior, thermal and phase-change phenomena, electromagnetism, and surface-tension effects, organized into nine task categories.
-
An interpretable evaluation protocol that combines temporal-consistency and task-observability screening with task-specific physical measurements, explicitly separating physical test outcomes from cases with insufficient measurement evidence.
-
A systematic evaluation of eight video generation models over the benchmark's 40 tasks and 1,280 videos, revealing a gap between temporal coherence and performance on explicit physical tests, plus substantial variation across tasks and physical domains.
-
Validation of the evaluator on synthetic videos with known physical relationships, and a human-agreement study against a direct VLM scoring baseline.
Main Findings
-
Best model scores 57.76 out of 100. Under the strict per-video consistency gate (C ≥ 80), the strongest model, Seedance, achieves an average composite score of approximately 57.76. The full ranking of average scores is Seedance 57.76, MiniMax 54.89, Cosmos 42.46, VBVR 37.79, Wan 35.66, LingBot 32.25, Hunyuan 31.97, and CogVideoX 18.81.
-
Large variation across tasks and models. Melting ice containing a stone (P26) is challenging for every model: model-averaged composite score of 3.45, maximum 15.00, and all eight models receive zero physical scores, so nonzero composite scores come entirely from the consistency term. Light reflection (P18) splits the field sharply: Seedance scores 96.15 and VBVR 95.75, while all other models score at most 49.82. Charged-sphere equilibrium (P31) shows a similar gap: Seedance 98.34 and MiniMax 97.24, with all other models at most 15.00.
-
Some tasks are solved broadly. Communicating-vessel equilibrium (P21) is an example of consistently strong performance: all eight models score between 93.09 and 99.01 and pass the automatic consistency gate on every available video.
-
Temporal consistency does not imply physical correctness. All available free-fall (P2) videos pass the automatic consistency gate (C ≥ 80), yet the average composite score across models is only 26.61.
-
Difficulty groups differ in gate pass rates. Only 56.9% of available Hard videos pass the automatic consistency gate, compared with 89.4% of Easy videos. Easy tasks include communicating-vessel equilibrium (P21), floating-ice immersion (P22), and hanging-chain equilibrium (P7); Medium tasks include free fall (P2) and pure rolling (P8); Hard tasks include equal-mass collisions (P5), rough-incline round trips (P11), and melting ice containing a stone (P26).
-
Model strengths differ by difficulty. MiniMax leads on Easy tasks, while Seedance achieves the highest task-averaged Hard score of 37.76, substantially above MiniMax's 19.87, although its absolute score remains low.
-
The measurement module agrees with known physics under controlled conditions. On 480 synthetic reference videos across 40 tasks (12 per task), with the VLM consistency gate bypassed, mean physical scores are 97.93 on Easy, 96.64 on Medium, and 98.25 on Hard tasks, with an overall mean of 97.53 out of 100.
-
The evaluator agrees with human judgments more than a direct VLM baseline. Using 320 videos across 40 tasks and ten annotators, within-task ranking agreement is 473 of 917 pairs (51.58%) for the evaluator, 398 (43.40%) for Direct VLM (Qwen3.6-27B), and 519 (56.60%) for Human–Human. Confirmed pairwise agreement over 160 pairs is 85 (53.12%) for the evaluator, 73 (45.62%) for Direct VLM, and 58.54% for Human–Human. This is an improvement of 8.18 percentage points over Direct VLM in ranking agreement and 7.50 percentage points in pairwise agreement, while remaining 5.02 and 5.42 percentage points below the Human–Human reference, respectively.
Methodology in Plain English
The benchmark treats each task as a controlled experiment. A task is defined by: physical assumptions (materials, initial conditions, observation geometry), an initial image paired with a generation prompt, and predefined measurable criteria. Initial frames are generated with GPT-Image-2.5 to establish a clear initial configuration, then paired with task-specific prompts refined through pilot generation and human inspection, forming a feedback loop between input construction and pilot review.
The design principle is to test relationships that cancel out unknown factors rather than requiring absolute calibration. Three test types are used: spatial relationships among geometric quantities (for example, a projectile launched at 45° and returning to launch height should satisfy H/R = 1/4), temporal relationships among periods and event timings (for example, two equal-length small-angle pendulums with different bob masses should have nearly equal periods), and coupled spatiotemporal relationships between motion and geometry (for example, pure rolling requires v = ωR).
Each generated video is first screened for temporal consistency and semantic coherence by Qwen3.6-27B, producing a score C(V) in [0, 100]. Physical measurements are attempted independently of this screen, using CoTracker3 for point tracking, SAM 2 for segmentation, and dedicated routines for geometric fitting, region-of-interest photometry, and temporal analysis. Extracted quantities (trajectories, periods, ray angles, contours, liquid levels) are compared against predefined physical criteria to produce residuals, which are mapped to agreement scores, typically via s(e; a) = 1/(1 + |e|/a) with a metric-specific error scale a (for example, a = 0.20 for the pendulum-length task). The physical score P is the aggregate of a task's indicators; the composite score is S = 0.15 C + 0.85 P̃, where P̃ = P only when C ≥ 80 and a physics score is recorded, otherwise P̃ = 0. Intermediate measurements and visual evidence are retained to help distinguish extraction errors from physical violations.
Eight image-to-video models are evaluated: CogVideoX1.5-5B (832×480, 81 frames), Cosmos3-Super (832×480, 81 frames), HunyuanVideo-1.5 (1264×720, 129 frames), LingBot-Video (832×480, 81 frames), MiniMax-H3 (1344×768, 124 frames), Seedance-2.5 (1270×726, 121 frames), VBVR-Wan2.2 (832×480, 81 frames), and Wan2.2-I2V-A14B (832×464, 81 frames). Each of the 40 tasks uses the same initial frame and task description for all models, with four samples per model–task pair, giving a planned total of 1,280 videos. Random seeds are fixed to 42–45 where supported.
For the human study, the Direct VLM baseline (Qwen3.6-27B) receives 24 sampled frames, the original generation prompt, and the reference image, resized to a maximum side length of 640 pixels, and predicts a physical correctness score in [0, 1] at temperature zero. The task-level difficulty groups (E/M/H, 15/15/10 tasks) are empirical, defined by cross-model task-mean scores.
Why This Matters
Impact on research. The benchmark shifts physical-consistency evaluation away from model-based judgments and reference-video similarity toward direct, task-specific measurement of observables. It supplies an interpretable basis for diagnosing where models fail, a validity check via synthetic videos with known physical relationships, and an explicit statement of measurement limitations. It also provides evidence that measurement-based scoring aligns with human judgment more closely than a direct VLM baseline (51.58% vs. 43.40% ranking agreement; 53.12% vs. 45.62% pairwise agreement), while still leaving room before the Human–Human reference.
Real-world applications:
- Embodied AI and robotics: determining whether a video world model can be relied on for prediction and planning, since physically implausible motions and interactions may undermine the reliability of simulated outcomes and the plans derived from them.
- World action models (WAMs): models that couple future-state prediction with action generation depend on simulated futures that respect physical constraints.
- Video generation model development and selection: per-task, per-domain scores let developers identify which phenomena (for example, light reflection or charged-sphere equilibrium) a model handles and which it does not.
- Simulation content and scientific visualization: lower-stakes but directly relevant uses where generated sequences should reproduce effects such as capillary rise, communicating vessels, or eddy-current braking plausibly.
Industry relevance. The evaluation covers eight widely used or commercially developed image-to-video systems, including MiniMax-H3, Seedance-2.5, HunyuanVideo-1.5, Cosmos3-Super, Wan2.2-I2V-A14B, VBVR-Wan2.2, LingBot-Video, and CogVideoX1.5-5B. The finding that the best average score is 57.76, with all eight models at zero physical score on at least one task, gives a concrete measure of the gap between visual coherence and physical consistency for teams building video world models.
Future Directions
- Closing the consistency–physics gap. Free-fall videos all passed the consistency gate yet averaged only 26.61, so future work must determine which failures are genuinely physical versus extraction or observability problems.
- Extending beyond observable relationships. The benchmark assesses only observable relationships under stated assumptions; unmeasured physical properties remain unverified, leaving calibration-dependent and quantitative-magnitude tests as an open area.
- Improving hard, multi-event tasks. Hard tasks such as equal-mass collisions (P5), rough-incline round trips (P11), and melting ice containing a stone (P26) require maintaining physical relationships through interactions and successive events, with a 56.9% consistency-gate pass rate versus 89.4% on Easy tasks.
- Raising automatic–human agreement. The evaluator remains 5.02 percentage points below Human–Human ranking agreement and 5.42 points below Human–Human pairwise agreement, motivating better measurement and aggregation.
Target Audience
Researchers and engineers working on video generation models, video world models, and world action models; benchmark and evaluation researchers interested in measurement-based rather than judgment-based scoring; embodied AI and robotics practitioners who depend on simulated futures for planning; and physicists or domain experts interested in how classical relationships such as projectile geometry, pendulum isochronism, Archimedes' principle, Faraday's law, and Jurin's law can be operationalized as testable video criteria.
Authors’ abstract
Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and planning in embodied AI systems. Existing evaluations often rely on model-based judgments or reference videos, while direct physical tests largely focus on mechanics. We introduce World Models' Last Exam in Physics, a measurement-based benchmark for evaluating physical consistency in video world models. The benchmark comprises 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism, and surface tension. Each task pairs an initial image and a generation prompt with predefined physical criteria, enabling interpretable tests of observable physical relationships without requiring reference videos. Its evaluator combines task-observability screening with task-specific quantitative physical measurements. Experiments on eight video generation models across 1,280 videos reveal persistent physical inconsistencies and substantial variation across tasks, with the best model achieving an overall score of 57.76 out of 100. Evaluation on synthetic videos with known physical relationships provides evidence for the validity of the measurement module under controlled conditions. The evaluator also achieves higher agreement with human judgments than a direct vision-language model baseline in both within-task rankings and pairwise comparisons. By combining coverage across physical domains with scores grounded in measurable evidence and explicit measurement limitations, the benchmark provides an interpretable basis for diagnosing physical inconsistencies and tracking progress toward physically consistent video world models.