Skip to content
AI.info

Research

HappyWorld-Bench

Overview Research area: Evaluation of world models across computer vision, video generation, 3D/spatial scene generation, and embodied AI. Technical level: Advanced. Scope: The paper introduces HappyW

HappyWorld-Bench
arXiv
2609.24308
Published
2026-09-21
Authors
Zhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu

AI summary

Overview

  • Research area: Evaluation of world models across computer vision, video generation, 3D/spatial scene generation, and embodied AI.
  • Technical level: Advanced.
  • Scope: The paper introduces HappyWorld-Bench, a three-track benchmark (video, spatial, embodied) built on a six-level world-capability hierarchy (W1–W6), paired with human A/B arena Elo ratings and automated behavioral metrics, used to evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates.

What This Paper Is About

Evaluating a world model requires more than judging how realistic a generated frame looks; it requires testing whether the world stays consistent when an agent moves through it, acts on it, revisits it, or intervenes to change it. Existing benchmarks tend to study controllability, physics, memory, or interaction in isolation, and each community (video generation, 3D reconstruction, embodied control) uses its own protocols, so results cannot be compared under a shared capability framework. HappyWorld-Bench addresses this fragmentation by defining one hierarchical capability taxonomy and instantiating it across three independent evaluation tracks, so that generated worlds are judged by state consistency and correctness of responses to actions and interventions rather than by visual quality alone.

Key Contributions

  1. A hierarchical capability framework (W1–W6) that formalizes what it means for a generated world to be reliable, giving a common vocabulary across video generation, 3D reconstruction, and embodied control. The six levels are: W1 Perceptual World (coherent world representation from visual or multimodal conditions with correct semantics, spatial structure, and short-term temporal continuity), W2 Interactive World (action-conditioned state transitions preserving local geometric, physical, and causal consistency), W3 Persistent World (global spatial structure, object identity, and accumulated state over long-horizon interaction, viewpoint change, occlusion, and revisitation), W4 Programmable World (explicit interventions on objects, events, behaviors, or world rules with causal propagation while unaffected content stays consistent), W5 Scalable World (infinitely extensible shared world states for multiple agents with communication, synchronization, cooperation, and conflict handling under partial observability), and W6 Universal World (generation, simulation, persistent state modeling, interaction, and planning unified into a system that generalizes across environments, tasks, modalities, and embodiments).

  2. HappyWorld-Bench itself, comprising 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases — 1,692 track-specific instances in total — with a dedicated suite of automated metrics for each of the three domains to support systematic and reproducible evaluation.

  3. HappyWorld-Arena, an arena platform that organizes human A/B comparisons of model outputs and derives model-level Elo ratings within each track, intended as a shared evaluation resource for community participation and exchange of results.

  4. A comprehensive empirical study of 14 video world models, 9 spatial systems, and 8 embodied candidates, revealing significant gaps in interactive simulation, state persistence, and programmable dynamics.

Main Findings

  • Arena rankings: HappyOyster achieves the highest Arena Elo at 1263, followed by Genie 3 at 1206. The full video-track ranking is HappyOyster 1263, Genie 3 1206, JoyAI-Echo 1147, Alaya-EVOKE 1132, Lingbot-World-v2 1115, Cosmos3 1092, Yume-1.5 1063, Lyra 2.0 1059, DreamX-World 1041, MatrixGame-3.0 943, SANA-WM 917, MatrixGame-2.0 856, ABot-World 843, and Open-Oasis 324.

  • Arena preference and level-wise capability diverge: Genie 3 obtains the highest W1 score at 82.8, slightly above HappyOyster at 82.4, while HappyOyster achieves the highest reported W2 and W3 scores at 78.0 and 76.8. The paper explicitly motivates level-wise analysis rather than relying solely on overall Arena preference.

  • W4 is sparsely covered: Programmable-world results are currently available only for HappyOyster and Lingbot-World-v2, with scores of 66.7 and 59.3.

  • Perception does not imply consistency (W1): Many models achieve strong visual quality and instruction following, yet these strengths do not translate into stable geometry, appearance, or world state over time. Models with competitive perceptual scores can still show substantially weaker temporal or state consistency.

  • Control following does not imply coherent consequences (W2): High trajectory accuracy (TA) or action execution (AE) scores indicate a model can follow requested control, but do not guarantee the resulting state stays coherent; models with similar control-following performance differ substantially in state consistency and causal consistency.

  • Persistence is hard (W3): Preserving established layout, appearance, temporal continuity, world state, and subject identity over extended interaction and revisitation remains difficult; visually plausible frames may still show geometric drift, appearance changes, or altered states. The paper notes W2 and W3 use different case pools, so their aggregate differences should not be read as a pure duration effect.

  • Spatial track gaps: Spatial world models perform poorly in physical plausibility, editability, and scene expansion, reaching at best 70.14% placement accuracy and 73.33% edit execution.

  • Embodied track gaps: Embodied models struggle to preserve state across multi-step actions and to respond precisely to altered action conditions and physical rules.

  • Overall conclusion: Video models exhibit reduced consistency during extended rollouts and revisits, and current world models face substantial challenges in maintaining reliable state under repeated change despite impressive generative quality.

  • Dataset characteristics: The video pool spans W1–W5 with 218, 294, 342, 180, and 104 cases at each level, organized into 74 fine-grained evaluation facets across five capability dimensions — perception and representation (342 cases), consistency and state retention (228 cases), causality and causal rollout (299 cases), controllable interaction and counterfactuals (212 cases), and multi-subject coordination (57 cases). The spatial track uses 300 scenes: 266 for evaluating visual quality, physical usability, consistency, and editing, and 34 for spatial expansion. The embodied track uses 254 formal test cases spanning atomic actions, multi-stage sequences, and action pairs. Data span nine application domains (robotics, urban and indoor environments, nature, transport, game worlds, daily life, industry, fantasy, and materials), rollouts range from short interactions to 60-second sequences, and first-frame resolutions span from below 1 MP to above 8 MP.

Methodology in Plain English

The authors first define what a "reliable world" means by breaking world modeling into six escalating capability levels, from merely building a coherent representation (W1) up to a fully unified universal simulator (W6). They then operationalize the same taxonomy in three separate tracks that match how different model families are actually used.

The video track builds 1,138 structured test cases. Each case pairs a first-frame image with a textual description of the scene and any selected subject, plus a time-indexed control sequence where applicable. Controls for W2 and W3 specify directions and time intervals rather than target positions, because the same input can produce different amounts of movement across models; W4 cases add natural-language instructions for subject actions, interactions, or rule edits; W5 and W6 allow controls assigned to specific subjects. Cases are annotated by target capability, W level, application domain, image source, and viewpoint, and each undergoes human review to confirm the task is clear, executable, and assessable, filtering out ambiguous targets, infeasible controls, or insufficient observable evidence.

The spatial track evaluates exported scenes for construction quality, navigation and stable placement, scene- and object-level consistency, editing that preserves non-target content, and expansion that retains existing regions and connecting paths.

The embodied track keeps the interface purely generative: each candidate is conditioned only on a single egocentric reference frame and an action prompt, so no simulator, action decoder, or robot is needed. Evidence is organized along a scoring axis of perception, consistency, causality, and controllability, and a capability axis rising from atomic action response (W2) to state persistence under ordered multi-step interaction (W3) to responses to an edited action or physical condition (W4), measured through matched-pair intervention from a shared initial state.

Measurement combines two sources. Human A/B comparisons in HappyWorld-Arena yield model-level Elo ratings that summarize overall preference. Automated metrics capture behavioral correctness: video quality (MUSIQ), perceptual preference (HPSv3), imaging stability, dynamic degree (RAFT optical flow), instruction following, background consistency (masked CLIP), geometric and texture consistency (DA3-based reprojection and photometric alignment), state and subject consistency (DINOv2/CLIP features on SAM2.1 masks), physical and content causality (Gemini-based checklists), interpenetration, trajectory accuracy (DA3-estimated camera motion versus time-aligned translation commands), action execution, and for W4 interaction validity and state fidelity. The video-track comparison covers 14 interactive world models, with five additional video-generation baselines added for W1 and W4 restricted to the two systems supporting the required intervention setting.

Why This Matters

Impact on research. The paper argues that without a unified benchmark probing how a world behaves under action, memory, and intervention across model forms, it is hard to tell whether systems support reliable world modeling beyond surface-level generation. The W1–W6 vocabulary gives otherwise separate communities — video world models, spatial world models, and embodied world models — a shared yardstick, and the results show that overall human preference and capability-level performance can rank models differently, which argues against single-number leaderboards.

Real-world applications (drawn from the nine domains the benchmark data spans and the three model tracks it evaluates):

  • Robotics and embodied manipulation: training or evaluating policies in action-conditioned imagined rollouts where state must persist across multi-step actions and physical rules must hold.
  • Urban and indoor environments, transport, and driving: generating navigable, geometrically consistent scenes that can be explored, revisited, and modified without corrupting existing content.
  • Game worlds and interactive media: sustaining controllable, persistent, programmable environments over long interaction horizons and revisitation.
  • Industrial, nature, fantasy, and materials domains: producing simulation-grade visual content that supports physical contact, placing objects validly, and editing scenes while preserving everything outside the edit's scope.

Industry relevance. All three tracks target deployment-relevant failure modes — drift over long rollouts, loss of object identity on revisit, invalid contact, and edits that leak into unaffected content. The Arena provides an operational mechanism (human A/B comparisons translated into Elo ratings) that industrial labs can reuse for comparative model assessment, and the benchmark's generative-only embodied interface means candidates can be scored without simulators or robots.

Future Directions

  • Closing the persistence gap in video world models: reducing the consistency loss observed during extended rollouts and revisits, and separating duration effects from other causes by aligning the W2 and W3 case pools, which the paper notes currently differ.
  • Advancing spatial physical plausibility, editability, and expansion: the three areas where the paper reports the weakest spatial performance, including the roughly 70% placement and 73% edit-execution ceilings.
  • Strengthening embodied state keeping and intervention response: improving preservation of state across multi-step actions and precision under altered action conditions and physical rules.
  • Extending evaluation upward through the hierarchy: W4 results are currently available only for two models, and the framework defines W5 (Scalable World, with multi-agent communication, synchronization, cooperation, and conflict handling under partial observability) and W6 (Universal World), which the reported evaluation does not yet reach.

Target Audience

Researchers and engineers working on video world models, interactive video generation, 3D/spatial scene generation, and embodied AI who need a shared, reproducible way to measure reliability under interaction rather than visual quality alone. It is also useful for benchmark designers and evaluation teams who want a capability taxonomy spanning multiple model families, and for practitioners in robotics, simulation, and interactive media deciding which world model is dependable enough for long-horizon, action-driven use.

Authors’ abstract

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

Read the original paper