Skip to content
AI.info

Research

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

Overview Research area: Computer Vision / multimodal agent evaluation, at the intersection of video understanding and programmatic 3D scene generation. Technical level: Advanced. The benchmark design

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
arXiv
2609.15478
Published
2026-09-14
Authors
Yolo Y. Tang, Daiki Shimada, Jiayue Meng, Jing Bi, Pinxin Liu, Yicheng Wang, Yunzhong Xiao, Zhangyun Tan, Zeliang Zhang, Chao Huang, Susan Liang, Qianxiang Shen, Luchuan Song, Ali Vosoughi, Mingqian Feng, Melika Filvantorkaman, Chenliang Xu

AI summary

Overview

  • Research area: Computer Vision / multimodal agent evaluation, at the intersection of video understanding and programmatic 3D scene generation.
  • Technical level: Advanced. The benchmark design and agentic setup assume familiarity with multimodal models, video benchmarks, and 3D authoring tools, though the central idea is explainable in plain terms.
  • Scope: One sentence — the paper proposes a benchmark that judges whether an AI agent understands a video by requiring it to rebuild that video as a coded Blender animation, and reports that current models reproduce appearance far better than they preserve facts.

What This Paper Is About

Video understanding is normally graded by asking a model questions about a clip. This paper argues that a stronger test is reconstruction: if an agent truly understands a video, it should be able to write code that recreates it as an animated 3D Blender scene. The goal is to build and validate a benchmark that measures that ability fairly across many models, and to see how well today's agents actually do.

Key Contributions

  1. A new benchmark task, BVB (Blender-VideoBench). Instead of question answering, agents must reconstruct real-world videos as animated Blender scenes written through code, with diffusion models explicitly not used.
  2. A shared evaluation harness, Mini-BVB. Each agent programs its reconstruction inside an identical sandbox under a common cost limit, so comparisons are made under matched conditions rather than at each agent's convenience.
  3. A two-axis scoring scheme. Reconstructions are rendered from their animated camera and judged by (1) Dual VQA, which counts how many spatiotemporal facts survive, and (2) Latent Similarity, which measures perceptual closeness to the source video. The overall score is a square-root mean, deliberately rewarding balanced performance rather than strength on only one axis.
  4. A large empirical and human-validated study. 51 configurations from 10 model families are evaluated and analyzed for semantic retention, perceptual similarity, reasoning effort, and cost. A blind study with 15 raters across five configurations checks that the automated perceptual measure tracks human preference.

Main Findings

  • Perception is strong, semantics are weak. The best model reaches 88.6 Latent Similarity while retaining only 53.7% of the source-correct spatiotemporal answers — it looks like the video much more than it knows what happened in it.
  • More reasoning helps appearance, not facts. Allowing additional reasoning improves visual similarity but does not close the gap in factual accuracy, suggesting the failure is not simply a matter of insufficient thinking time.
  • The automated similarity score is human-aligned. In a blind study with 15 raters and five configurations, Latent Similarity correlated strongly with human preference.
  • Programmatic reconstruction works as a test. The authors conclude it is a viable way to probe agentic video understanding.
  • Semantic retention is the open problem. The abstract identifies factual preservation, not visual fidelity, as the main remaining challenge. Details of cost breakdowns and per-family rankings are not given in the abstract.

Methodology in Plain English

The researchers took real-world videos and asked AI agents to write code that builds an equivalent animated scene in Blender. To keep the contest fair, every agent worked inside the same lightweight harness, in the same sandbox, with the same spending ceiling — no agent got more attempts, better tools, or a bigger budget than another. Each finished scene was then rendered from its own animated camera, so the output could be judged as a video rather than as static geometry. Scoring happened on two fronts: a fact-based check counting how many of the source video's spatiotemporal details the reconstruction preserved, and a perceptual check measuring how close the result looked to the original. These two were combined with a square-root mean so that a model cannot win by maxing out one measure while neglecting the other. Finally, human raters were shown a subset of configurations in a blind setup to confirm that the perceptual score matched what people actually preferred.

Why This Matters

Evaluation shapes what models learn. Grading video understanding by multiple-choice questions rewards models that can extract a few salient facts; grading it by reconstruction demands a full, executable model of what happened, where, and in what order. The paper's results suggest current agents can imitate the look of a video without capturing its content — a distinction that question-answering benchmarks tend to hide.

Real-world applications:

  • Robotics and embodied agents: an agent that can reconstruct a demonstration video as a scene plan is closer to one that can act on what it observed.
  • Simulation and synthetic data: reconstructing real footage into editable 3D scenes supports generating training environments without diffusion-based generation.
  • Animation, VFX, and game pipelines: automating the conversion of reference footage into animated scene setups could reduce manual authoring work.
  • Video search, editing, and forensics: fact-level reconstruction offers a finer-grained way to verify what a clip contains than caption-based retrieval.

Industry relevance: the benchmark targets labs building multimodal agents, as well as 3D content, animation, and robotics companies that need models to convert video into structured, editable representations. The shared cost limit also makes it relevant where compute budgets — not just accuracy — determine which approach is practical.

Future Directions

  • Closing the semantic gap. Since extra reasoning improves only appearance, the paper's framing invites work on architectures, training objectives, or tools that specifically raise factual retention.
  • Diagnosing where facts are lost. The abstract reports the retention rate but not which spatiotemporal facts agents systematically drop; identifying those failure modes is a natural next step.
  • Extending beyond Blender and beyond coding agents. The benchmark relies on programmatic reconstruction in one specific toolchain under one cost limit — how the findings transfer to other 3D engines, limits, or non-agentic models is an open question.
  • Refining the scoring scheme. The authors could examine whether the square-root mean and the Dual VQA design remain the right balance as models improve, and broaden the human validation beyond five configurations.

Target Audience

Researchers and engineers working on multimodal agents, video understanding benchmarks, and 3D scene generation will get the most from this paper, as will practitioners in animation, VFX, robotics, and simulation who care about turning video into editable 3D representations. Readers looking for a concise example of how evaluation design can expose the difference between imitation and understanding will also find it useful.

Authors’ abstract

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introduce BVB, Blender-VideoBench, a benchmark that tests this ability by asking agents to reconstruct real-world videos as animated Blender scenes. To ensure fair comparison, each agent programs the reconstruction through a lightweight harness, Mini-BVB, in an identical sandbox under a shared cost limit. The benchmark renders each reconstruction from its animated camera and evaluates it on two axes: (1) Dual VQA measures how many spatiotemporal facts the reconstruction preserves. (2) Latent Similarity measures how closely the reconstruction matches the source video perceptually. Our overall score, a square-root mean, favors balanced performance. We evaluate 51 configurations from 10 model families and analyze semantic retention, perceptual similarity, reasoning effort, and cost. The best model reaches 88.6 Latent Similarity but retains only 53.7% of the source-correct spatiotemporal answers. Additional reasoning improves visual similarity but does not close this gap in factual accuracy. In a blind study with 15 raters and five configurations, Latent Similarity correlates strongly with human preference. These results show that programmatic reconstruction is a viable test of agentic video understanding, and that semantic retention remains the main challenge.

Read the original paper