Research
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
Overview Research area: Multimodal AI agents (vision-language models and vision-language-action models) evaluated on interactive 3D environment inspection, positioned between embodied AI benchmarking

- arXiv
- 2609.40325
- Published
- 2026-09-30
- Authors
- Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang, Shiyu Chang
AI summary
Overview
Research area: Multimodal AI agents (vision-language models and vision-language-action models) evaluated on interactive 3D environment inspection, positioned between embodied AI benchmarking and automated game/simulation quality assurance.
Technical level: Advanced. The paper assumes familiarity with VLM/VLA architectures, agent tool use, model-context-protocol style tool interfaces, VLM-as-a-judge evaluation, and interactive simulation environments (Unreal Engine 5, Three.js).
Scope (one sentence): The paper introduces WorldAuditBench, a 213-task, 13-environment benchmark for "3D world auditing," and uses it to measure whether five frontier multimodal models can couple navigation with visual reasoning to find anomalies that humans detect at 83.4% success.
What This Paper Is About
Interactive 3D worlds can look convincing while containing defects: objects floating above the floor, walls an agent can walk straight through, or objects that vanish once the viewer looks away. Detecting these defects requires more than looking at a scene, because some anomalies only become apparent through interaction, viewpoint changes, or revisiting a location over time. The paper's goal is to build a benchmark that forces AI agents to actively gather evidence about a 3D world and then decide whether a real anomaly exists, and to measure how far current models are from human auditors at that task.
Key Contributions
-
A benchmark for interactive world auditing. WorldAuditBench contains 213 tasks across 13 interactive 3D environments, covering five anomaly families — static physics, interactive physics, spatial consistency, temporal consistency, and semantic consistency — with the families subdivided into fifteen categories. Each task contains exactly one pre-defined labeled anomaly and pairs it with expected normal behavior, scene context, and an evaluation rubric.
-
Two auditing paradigms evaluated on the same tasks. The paper compares a single VLM agent auditor that reasons and acts in a closed loop against a two-stage VLA–VLM auditor that first explores with a VLA model and then analyzes the fixed recorded trajectory with a VLM, isolating how much the coupling of action and visual reasoning matters.
-
A systematic analysis of failure modes. The authors separate the failure to expose an anomaly during exploration from the failure to identify it once exposed, and run ablations on starting distance, exploration budget, and task guidance, plus a multi-anomaly experiment. They report a specific finding that VLA models struggle to even reach anomalies in large 3D worlds, and that memory tools substantially improve VLM auditor anomaly discovery.
-
A human baseline. Roughly ten computer science PhD students, two per task, with up to ten minutes per task, reach 83.4% success, establishing a concrete gap against the best model result of 42.3%.
Main Findings
-
Models fall far short of humans. Across five backbones and two paradigms, success rates range from 6.6% to 42.3%, compared with 83.4% for human auditors.
-
Interactive VLM auditing beats two-stage VLA–VLM auditing. VLM-only agents score 28.2% to 42.3% overall, while two-stage VLA–VLM auditors score 6.6% to 17.4%. The authors attribute this to the VLM agent's ability to choose new viewpoints and actions to verify a suspected defect, whereas the two-stage auditor analyzes a trajectory that is already fixed.
-
Best individual results. Among VLM auditors, GPT-6 Astra reaches 42.3% overall success, followed by Gemini 3.8 Flash at 32.4%, Claude Opus 5 at 28.2%, Muse Spark 1.3 at 15.0%, and Qwen 3.8 Flash at 8.5%. Among VLA–VLM auditors, Gemini 3.8 Flash leads at 17.4%, followed by GPT-6 Astra at 16.4%, Claude Opus 5 at 15.0%, Qwen 3.8 Flash at 12.7%, and Muse Spark 1.3 at 6.6%.
-
Anomaly families differ sharply in difficulty. Static physics and semantic consistency are easiest for VLM auditors (GPT-6 Astra: 59.3% and 68.2%). Spatial and temporal consistency are harder because they require comparing observations across viewpoints or time. Temporal consistency is the hardest family, with success at or below 12.5% for every model and every paradigm.
-
VLM agents expose more anomalies. Using GPT-6 Astra on the 213 tasks, geometric coverage of the anomaly is 91.1% for the VLM agent versus 47.4% for the VLA explorer, and human-rated exposure is 68.1% versus 34.7%.
-
The VLM advantage persists after exposure. On tasks geometrically covered by both paradigms, the VLM agent succeeds 38.5% of the time versus 26.0% for the VLA–VLM auditor; on tasks with human-rated exposure, the gap is 62.1% versus 47.3%. The authors attribute this to inspecting additional viewpoints, testing anomalies through interaction, and using memory tools to retrieve earlier frames and notes.
-
Efficiency is a trade-off, not a free win. Given prerecorded trajectories, VLA–VLM analysis is 2.7 to 10.3 times faster than interactive VLM auditing and has lower inference costs, but achieves lower success rate.
-
Distance matters. Moving the start position closer to the anomaly, reducing the navigable route to roughly one third of its original length, raises VLM success from 33.3% to 38.1% and VLA–VLM success from 5.6% to 7.1% on the 126 Unreal tasks with Gemini 3.8 Flash.
-
Budget matters, with diminishing returns. Halving the budget (20 steps for VLM, 30 seconds for VLA–VLM) drops VLM success from 33.3% to 23.0% and VLA–VLM from 5.6% to 4.0%. Increasing it to 1.5x (60 steps for VLM, 90 seconds for VLA–VLM) raises VLM to 34.9% and VLA–VLM to 7.1%.
-
The anomaly-type hint matters more than the in-context example. Removing the in-context example while keeping the type hint changes VLM success from 33.3% to 31.0% and leaves VLA–VLM at 5.6%. Removing both drops success to 15.9% for VLM and 2.4% for VLA–VLM.
-
Coexisting anomalies overwhelm agents. On 21 anchor tasks (three from each of seven Unreal Engine 5 environments) with two additional anomaly conditions, anchor SR stays roughly stable at 42.9% for one or two anomalies and 47.6% for three, but complete discovery is rare: the agent identifies every target in only 3 of 42 two-anomaly runs (7.1%) and 2 of 42 three-anomaly runs (4.8%), with target recall of 29.8% and 29.4%.
Methodology in Plain English
The authors first collected 13 publicly available interactive 3D environments built in Unreal Engine 5 or Three.js, spanning indoor, urban, historical, industrial, and natural settings. Larger environments were split into bounded regions, yielding 27 scenes in total. All 213 tasks were hand-crafted by the authors. For each task they started from a clean scene, fixed the agent's starting position, orientation, exploration boundaries, and available interactions, then deliberately introduced exactly one anomaly matching their taxonomy — for example, by floating an object, breaking a collision boundary, changing an object's appearance across viewpoints, making an object change over time, or placing an object that conflicts with the scene's historical setting or intended function.
Task quality was checked independently. Reviewers who did not build the tasks explored each scene through a browser interface and rated it Pass, Fail, or Uncertain; failed or uncertain tasks were revised and re-reviewed until the task was rated Pass by at least two judges on the same task version.
For evaluation, an agent receives auditing instructions, a scene description, an anomaly-type hint, a matching in-context example, and a first-person RGB view. It then explores and submits a report with supporting images. In the VLM-only paradigm the agent gets 40 environment actions and can reason and act in a loop, using environment actions, memory tools, and anomaly-reporting tools. In the VLA–VLM paradigm, Open-P2P 1.2B collects 60 seconds of simulated exploration with observations sampled every 0.5 seconds, and a VLM then analyzes the recorded trajectory. Success is judged by GPT-6 Astra acting as a VLM judge against each task's rubric. The authors note that Table 7 reports agreement between the VLM judge and human judgments, but the numeric agreement value is not included in the provided content.
Why This Matters
Research impact. The paper reframes anomaly detection as an evidence-gathering problem rather than a passive perception problem. By separating anomaly exposure from anomaly identification, it gives the field a diagnostic instrument: a model can fail because it never reached the anomaly, or because it reached it and still could not recognize it. This distinction is difficult to make with existing benchmarks, most of which supply fixed video or image observations in advance. It also connects multimodal agent research to the long-standing game-bug literature that classifies defects by whether they appear in a single state or only over time.
Real-world applications (potential, based on the settings the paper studies):
- Automated quality assurance for game and simulation content, where defects like missing collision, floating props, and broken object persistence currently require human testers.
- Validating the fidelity of simulation environments used to train embodied agents, since an agent that succeeds by walking through a broken wall is not a reliable navigation result.
- VR/AR and digital-twin content inspection, where spatial and temporal consistency across viewpoints matter.
- Regression testing of generative or procedurally constructed 3D scenes, where new assets are added continuously and defects are hard to enumerate by hand.
Industry relevance. The task is a plausible automation target for game studios, simulation vendors, and companies building embodied AI training pipelines. The paper's efficiency measurements matter commercially: two-stage trajectory analysis is 2.7 to 10.3 times faster and cheaper than interactive auditing, which is the trade-off a production audit pipeline would have to weigh against the large drop in success rate.
Future Directions
- Improve VLA exploration. The paper attributes much of the two-stage paradigm's failure to VLA models not reaching anomalies in large 3D worlds, so better exploration policies for embodied models are an obvious next step.
- Close the temporal consistency gap. Temporal consistency success sits at or below 12.5% for every model and paradigm. Building agents that reliably collect and compare before-and-after evidence remains unresolved.
- Scale auditing to multiple coexisting anomalies. Anchor detection held steady as anomaly count grew, but complete discovery collapsed to 7.1% and 4.8% in the two- and three-anomaly conditions. Methods for coverage and prioritization under a fixed budget are needed.
- Strengthen the coupling between reasoning and memory. The memory tool that stores and retrieves observations substantially improved VLM anomaly discovery, which points toward studying how agents should decide which observations to seek and whether they support a conclusion.
- Broaden the anomaly taxonomy and judge validation. The provided content does not report the numeric agreement between the VLM judge and human judgments, so judge reliability across the fifteen categories remains an open question for follow-up work.
Target Audience
Researchers and engineers working on multimodal agents, embodied AI, and VLM/VLA systems who need a stress test for coupled perception and action; benchmark designers interested in interactive evaluation methodology; game and simulation QA engineers exploring automated defect detection; and graduate students entering the area of agentic evaluation in 3D environments. The paper is written for readers already comfortable with the vocabulary of frontier model evaluation, so it is less suited to readers seeking an introductory treatment.
Authors’ abstract
As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.