Skip to content
AI.info

Research

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Overview Research area: Computer vision, specifically 4D inverse graphics — reconstructing dynamic scenes from video as executable graphics programs — evaluated through a benchmark for multimodal codi

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
arXiv
2610.03715
Published
2026-10-02
Authors
Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu

AI summary

Overview

Research area: Computer vision, specifically 4D inverse graphics — reconstructing dynamic scenes from video as executable graphics programs — evaluated through a benchmark for multimodal coding agents, with connections to physical scene understanding, world modeling, and simulation.

Technical level: Intermediate. The task definition, metric families, and results are accessible without deep graphics expertise, though familiarity with 3D rendering, Chamfer distance, optical flow, and Elo ratings helps.

Scope: The paper introduces 4DCodeBench, a 200-scene benchmark (100 real videos, 100 synthetic scenes with 4D ground truth) on which 18 multimodal coding models are asked to reconstruct dynamic scenes by writing executable code, and reports how those models perform on appearance, geometry, and dynamics.

What This Paper Is About

Most inverse graphics work fits a prescribed model to visual observations. This paper instead asks an agent to write the model itself: given only a reference video, the agent must produce an executable graphics program that builds the scene's 3D geometry and its evolution over time, then renders it back to video. The central question is whether today's multimodal coding agents can infer dynamics — deformation, flow, fracture, and interacting materials — from video, and whether strong static reconstruction ability carries over to moving scenes.

Key Contributions

  1. A benchmark and dataset for 4D inverse graphics through code generation. 4DCodeBench comprises 200 dynamic scenes — 100 real videos and 100 synthetic scenes with 4D ground truth — spanning diverse materials, interactions, and physical phenomena (collisions, deformation, flow, fracture). Agents must reconstruct each video as an executable graphics program that exposes explicit geometry and dynamics over time.

  2. A comprehensive evaluation of frontier and open-weight models. 18 multimodal coding models are benchmarked under a common execution environment and evaluation protocol, measuring visual fidelity, 2.5D and 3D geometry, and 2D and 3D dynamics. The analysis shows large differences across models and a consistent weakness in reconstructing dynamics even when appearance and static geometry are recovered well.

  3. Human-aligned evaluation and failure analysis. A large-scale human preference study (3,587 judgments from 76 participants) supports automated VLM- and reconstruction-based metrics that closely track human judgments, and is used to analyze which scene properties and physical regimes expose different reconstruction failures.

  4. An analysis of how agents implement motion. Executable submissions are classified into analytic motion, custom simulation, Blender physics, or keyframing, revealing that leading models differ substantially in strategy even on the same reconstruction task.

Main Findings

  • Static reconstruction beats dynamic reconstruction. Across all 18 models, Perceptual and 3D Geometry scores average higher than 2D and 3D Dynamics scores. Even the strongest model, Astra [Max], scores 0.91 on the static families versus 0.67 on the dynamic families.

  • Leaderboard ordering. Astra [Max] leads the Overall ranking, followed by Opus 5.5 [High], Astra [High], Fable [High], and Astra [Low]. Open-weight models generally trail proprietary ones, but the leaderboard distinguishes substantial differences among open-weight models as well.

  • Automated metrics track human preference closely. Across 3,587 pairwise judgments from 76 participants, human and VLM Elo correlate with Spearman rho = 0.980. On 916 shared individual comparisons, human and VLM judgments agree in 89.3% of cases (kappa = 0.761), near the 92.1% agreement between human annotators on repeated comparisons (kappa = 0.816). Human Elo also correlates with the Overall score (rho = 0.96). VLM judgments are less reliable for closely matched models: agreement is near chance within 50 Elo points, reaching 75% at a 132-point separation and 90% at 343 points.

  • Most submissions execute successfully. 90.1% of all submissions are fully executable, with at least 96.5% for every proprietary model and 29.0–95.5% for open-weight models.

  • VQA reveals large per-scene variation. Astra [Max] reaches 87.6% mean per-scene accuracy, versus 78.8% for Fable [High] and 29.5% for GLM [Max]. Astra [Max] answers every question on 59% of scenes, whereas Mistral answers none on 94%. First-frame questions are generally answered more reliably than questions about the final frame and key events.

  • Scene category changes what fails. Relative to synthetic scenes, real scenes have lower VQA and Perceptual scores (standardized effects of −0.73 and −0.88). Scenes with multiple matter types score lower than single-matter scenes (−0.57 and −0.39). Co-dimensional scenes have lower 2D Dynamics and 2.5D Geometry (−0.35 and −0.76), while driven scenes have lower VQA but higher 3D Geometry (−0.30 and +0.46). Flowing scenes also show lower 3D Dynamics (−0.21), though this effect does not meet the figure's consistency criterion.

  • Agents prefer analytic motion. Overall, 67% of solutions use analytic motion, 19% custom simulation, 10% Blender physics, and 3% keyframing. The distribution varies sharply by model: Opus 5.5 [High] uses custom simulation in 61% of its solutions, while Astra [Low] uses analytic motion in 85% and Terra [High] in 100%.

  • More reasoning helps within a model, but more tokens across models does not. Raising Astra's reasoning effort from Low to High to Max improves Overall from 0.73 to 0.77 to 0.79, 2D Dynamics from 0.48 to 0.55 to 0.60, 3D Dynamics from 0.63 to 0.69 to 0.73, and VQA from 78.2% to 85.0% to 87.6%. Across different models, however, token volume and agent steps have little correlation with Overall score.

  • Mesh-quality checks disagree with the leaderboard. Five additional reconstruction metrics broadly agree with the leaderboard rankings, whereas watertightness and related mesh-quality checks do not — simple but inaccurate geometry can score well. Those checks are therefore reported separately as structural diagnostics rather than included in the Overall score.

  • Kimi K3 was excluded from reported results. The authors note that its unusually long agent trajectories cost about 50% more than Fable without a commensurate gain in reconstruction quality.

Methodology in Plain English

The benchmark poses 4D inverse graphics as a code generation problem. For each scene, an agent receives the reference RGB video and a fixed task prompt describing the reconstruction task and required output format — no scene names, semantic descriptions, or object lists. The agent writes an executable program that constructs the scene's 3D geometry and its evolution over time. Executing that program must expose a geometric state at each evaluated timestep and render it in Blender; the render must match the reference's resolution, frame count, and frame rate, and the program must run without access to the reference video. How dynamics are implemented is left open: animation, Blender physics, simulation libraries such as Taichi or Warp, or custom code are all permitted.

The dataset combines two sources with complementary roles. The 100 real-world clips are curated from physics video benchmarks, robot-manipulation datasets, and web footage (51, 21, and 28 clips respectively), selected for stationary cameras, limited occlusion, and continuous footage, and excluding humans, animals, and visible hands. The 100 synthetic scenes are built with physical simulators and provide ground-truth geometry and motion at every frame, enabling direct 3D and 4D evaluation rather than only comparing rendered pixels. More than two-thirds of the synthetic examples use a simulator richer than Blender, and the authors extended the Genesis simulator with three new material-point-method constitutive models: a viscoplastic model, the snow model of Stomakhin et al. (2013), and a Drucker–Prager model for sand. All scenes are rendered in Blender. Videos are capped at 60 fps, and real videos at 300 frames.

Evaluation spans five metric families. Perceptual (DINOv3 similarity), 2D Dynamics (Dynamic IoU), and 2.5D Geometry (depth error) are measured against the reference video; 3D Geometry (Chamfer distance at frame 0) and 3D Dynamics (Trajectory DTW and EMD step) are measured against the reference world on synthetic scenes. The Overall score averages the five families. On the reconstruction side, quantities are computed analytically from the 4D world rather than estimated from rendered pixels. Failed runs receive the worst applicable score. Mesh-validity checks (open boundaries, non-manifold edges, degenerate faces, self-intersections, interpenetrations) are reported as diagnostics outside Overall. Six real scenes with transparent main objects are excluded from metrics that compare rasterized geometry to the reference.

The 18 models run one time on each of the 200 scenes. Proprietary models run in their provider's native agent CLI; open-weight models run in the open Stirrup harness. Each run executes in an isolated container with one GPU, an empty workspace, and read-only access to the reference video and task files, with Blender, common numerical and simulation libraries, and an offline copy of the Blender API documentation available. Each submission must include a rendered video, camera parameters, per-frame geometry, and temporal correspondence for dynamic matter. Human preference was collected by showing the reference video followed by two models' renders side by side in randomized order, with model pairs sampled adaptively to shrink the widest confidence intervals; judgments were aggregated into Elo ratings with a Bradley–Terry model, as for the VLM judge. The human study covers 17 of the 18 models (all but Opus 5.5 [High], which was not part of the study due to its time of release).

Why This Matters

Impact on research. The paper reframes scene understanding as world-program generation, requiring agents to commit their physical understanding to an executable 4D scene that can be simulated, inspected, and verified. It also shows that strong static reconstruction ability does not yet translate into reliable reconstruction of complex dynamics, giving the field a measurable target. The close agreement between automated metrics and human preference supports using the automated suite to compare new models.

Real-world applications.

  • Robotics and embodied AI, where agents must infer materials, interactions, and driving forces from video of manipulation and everyday activity.
  • Content creation pipelines for games, visual effects, and animation, where a video could be converted into editable, simulatable scenes rather than pre-baked footage.
  • Physical-plausibility evaluation for video generation, by checking whether generated or reconstructed scenes obey dynamics consistent with the observation.
  • Simulation and digital-twin authoring, where the recovered program exposes explicit geometry and dynamics that can be re-run under new conditions.

Industry relevance. Proprietary and open-weight model developers get a leaderboard that separates appearance, geometry, and motion rather than collapsing them into one number, plus a cost–quality view: the paper reports that greater token use across different models does not consistently yield higher Elo, while greater reasoning effort within Astra does. That distinction matters for anyone deciding how much inference budget to spend on a reconstruction task. The finding that simple but inaccurate geometry can score well on watertightness-style checks is also a caution for teams building mesh-quality validators.

Future Directions

  • Test transferable physical understanding. Current evaluation focuses on reconstruction fidelity, leaving physical mechanisms underconstrained; prescribed trajectories can achieve high scores while the generalization benefits of simulation remain largely untested. Interventions on initial conditions and external forces, plus longer-horizon prediction, are proposed as ways to probe this.
  • Track intermediate reconstructions. Evaluation currently focuses on final submissions. Following render-and-compare rounds could show how quality improves with iteration and token expenditure, when gains saturate, and how effectively agents use visual feedback.
  • Determine when simulation helps. Leading models adopt markedly different strategies, from predominantly analytic motion to custom physical simulation, raising the question of whether requiring physics-based simulation would preserve their relative performance.
  • Improve reasoning efficiency. Since token expenditure across models does not consistently yield better reconstructions, the paper suggests stronger physical intuition could help agents form plausible hypotheses earlier, anticipate the consequences of modeling choices, and refine solutions in fewer render-and-compare iterations.

Target Audience

Researchers and engineers working on inverse graphics, world models, physical scene understanding, and multimodal coding agents, as well as benchmark designers who need evaluation protocols that separate appearance, geometry, and dynamics and that align with human judgment. Model developers comparing frontier and open-weight systems under a shared execution environment will find the leaderboard, metric definitions, and diagnostic analyses directly useful. Readers seeking an introduction to 4D scene reconstruction will benefit from the task formulation, but should expect moderate technical density in the metric and simulator sections.

Authors’ abstract

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench

Read the original paper