Research
To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
Overview Research area: Computer Vision, specifically text-driven 3D scene and world generation, spanning urban-scale exterior synthesis and building-scale interior synthesis. Technical level: Advance

- arXiv
- 2608.05879
- Published
- 2026-08-06
- Authors
- Xiaobin Huang, Zilong Huang, Yang Luo, Hongchao Fan, Yiping Chen, Ting Han
AI summary
Overview
- Research area: Computer Vision, specifically text-driven 3D scene and world generation, spanning urban-scale exterior synthesis and building-scale interior synthesis.
- Technical level: Advanced. The paper describes a multi-module pipeline that couples large language and image models with image-to-3D conversion, and defines a hierarchical "world context" representation.
- Scope: The paper proposes HoloWorld, a framework that generates a city's exteriors and the interiors of its buildings as corresponding parts of one coherent 3D urban world, and evaluates it against urban generation baselines and an independent indoor-generation baseline.
What This Paper Is About
Text-driven 3D generation can already produce large outdoor cities and detailed indoor rooms, but these are typically synthesized separately, so an interior rarely matches the building it is supposed to sit inside. The authors formulate unified indoor-outdoor urban scene generation as a single world-consistent task: given one text prompt, generate an urban exterior grid and, for selected buildings, interiors that inherit those buildings' semantics, appearance, and footprint. HoloWorld is their framework for doing this by carrying a continuously updated, cross-scale context from city planning down to individual buildings.
Key Contributions
- Task formulation and framework. The paper formulates unified indoor-outdoor urban scene generation and introduces HoloWorld, which generates exteriors and corresponding interiors as one coherent 3D urban world. The authors state this is the first framework to unify indoor and outdoor generation within a coherent 3D urban world.
- A cross-scale "living context." A hierarchical world context with world, block, and building levels that is initialized from the user description and progressively updated, supporting coherent autoregressive generation across blocks and explicit correspondence between urban exteriors and building interiors.
- Building-grounded indoor generation. A strategy that localizes context to individual building instances so interior synthesis is conditioned on the building's identity, inherited appearance, asset library, and recovered footprint, with the generated interior constrained to lie within that footprint.
- An evaluation of the unified setting. Experiments comparing against urban generation methods (CityCraft, SynCity, MajutsuCity) and an independent-generation baseline (TRELLIS), plus ablations on the dynamic world context and on autoregressive neighborhood conditioning.
Main Findings
- Urban exterior quality: Under GPT-5.5-based evaluation, HoloWorld improves AQS over the strongest baseline, MajutsuCity, by relative margins of 9.38%, 4.65%, 9.00%, and 8.11% on SVC, SRC, MTF, and LA respectively, and obtains the highest RDR scores in all four dimensions. Human evaluation with 20 experts shows the same trend. The abstract reports an average AQS improvement over the SOTA of 7.68% and the highest average RDR score.
- Indoor-outdoor correspondence vs. independent generation: Against TRELLIS given matched functional and stylistic descriptions, the paper reports Visual AQS increasing from 5.88 to 8.21 and RDR from 1.83 to 24.29 under GPT-5.5, with the same trend in human judgments. HoloWorld additionally reports Spatial AQS of 8.77 (GPT-5.5) and 8.65 (human) and a Shape IoU of 0.997, while TRELLIS's shape and spatial entries are marked as unavailable because its independently generated interior is not grounded in the exterior footprint.
- Dynamic context ablation: Freezing the world context reduces functional, visual, and spatial AQS from 7.67, 8.17, and 7.75 to 5.75, 4.58, and 3.25 under GPT-5.5, and Shape IoU falls from 0.994 to 0.670. Human average AQS drops from 8.21 to 5.54. The authors interpret this as showing that the initial city description alone does not supply sufficient building-specific semantic, visual, and geometric constraints.
- Autoregressive neighborhood conditioning ablation: Removing access to previously generated neighboring blocks lowers continuity AQS from 8.25 to 7.25 and RDR from 24.04 to 16.66 under GPT-5.5, and from 7.75 to 5.38 (AQS) and 21.97 to 19.07 (RDR) under human evaluation, with visible boundary discontinuities shown in the qualitative figure.
- Qualitative observations: The authors report that CityCraft and MajutsuCity produce well-structured layouts but less diversity in public spaces and fine-grained landscape elements, while their method jointly synthesizes buildings, roads, parks, plazas, and layered vegetation; compared with SynCity's tile-based strategy, they report more coherent organization and visual consistency across neighboring blocks.
- Not reported: The paper does not report dataset sizes, training procedures, compute budgets, or a count of generated cities. It states that the system uses GPT-5.4 for language-based reasoning, GPT-Image-2 for block image generation and editing, and the Meshy API for image-to-3D conversion, and that all experiments use 3 x 3 block grids for comparability while the framework supports arbitrary R x C grids.
Methodology in Plain English
HoloWorld works like a growing record of a world that is refined as generation proceeds.
- Start from text. A user description sets the city theme, functional organization, and inter-block relationships over a grid of blocks, and also fixes shared exterior and interior style references. This becomes the world-level context.
- Plan each block. A block design module turns the world context into block-specific spatial layouts stored at the block level.
- Generate exteriors autoregressively. Each block image is generated conditioned on the world context, its own block design, and the blocks already generated around it. The block images are converted to 3D models and assembled into the urban exterior.
- Ground buildings. A building region in a block image is associated with its 3D instance, giving the building an identity and a footprint. This creates the building-level context, which combines inherited semantics and appearance with instance identity and geometry.
- Generate interiors from that building-level context. An indoor module first plans a structured interior program (functional spaces, spatial relationships, required assets) and builds an asset library guided by exterior appearance cues and inherited style. Then it synthesizes the interior hierarchically, from room layout to furniture, architectural elements, and smaller objects, with the generated interior's footprint constrained to lie inside the exterior building's footprint.
- Validate before propagating. After each step, a candidate result and newly derived evidence are produced; only validated results are merged into the context, so unreliable information does not propagate to later stages.
Why This Matters
The paper reframes indoor and outdoor generation as one consistency problem rather than two separate synthesis problems. It supplies a concrete mechanism for propagating semantics, appearance, and geometry across spatial scales, together with evaluation protocols for building-level correspondence (functional, visual, spatial consistency and Shape IoU) and cross-block continuity, which other work in this space has not coupled in this way.
Real-world applications:
- Games, film, and virtual production: generating a navigable city whose building interiors match their exteriors, rather than assembling mismatched assets.
- Embodied-agent and robotics simulation: training environments where agents move continuously between streets and building interiors with consistent identity.
- Digital twins and urban planning: rapid, layout-constrained previews of districts and building interiors keyed to real footprints.
- AR/VR and interactive 3D: persistent virtual districts where the same building always contains the same interior, supporting spatial memory and navigation.
Industry relevance: the pipeline is built on existing commercial and open components (GPT-5.4, GPT-Image-2, the Meshy API, TRELLIS), which lowers the barrier for studios and simulation teams to adopt the context-based orchestration idea. The main practical payoff claimed is correspondence and continuity, the two properties that most often break when urban and indoor assets are produced by separate tools.
Future Directions
- Multi-floor and vertical structure: the current framework focuses on single-floor interiors and does not explicitly model vertical building structures.
- Richer structural constraints: the authors list this as future work, which would tighten how interiors respect more than a 2D footprint.
- Interactive physical validation: planned as a way to check generated worlds beyond appearance and layout consistency.
- Open questions the paper leaves: how the approach behaves beyond the 3 x 3 grids used in all experiments; how much the reported gains depend on the GPT-5.5-based evaluator, given that human scores are reported alongside but on a coarser trend; and why Shape IoU is reported as 0.997 in the correspondence comparison and 0.994 in the ablation for the full model.
Target Audience
Researchers and practitioners working on text-to-3D generation, urban and city-scale scene synthesis, indoor scene generation, and procedural world building. It is also relevant to developers building simulation, game, XR, or digital-twin pipelines who need interiors that correspond to specific buildings, and to researchers interested in cross-scale consistency and context-propagation mechanisms for generative systems.
Authors’ abstract
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, but these domains are usually synthesized independently, lacking the correspondence required for a coherent urban world. We present HoloWorld, a unified indoor-outdoor urban world generation framework built on a continuously updated cross-scale world context. Initializing from a user description, HoloWorld progressively represents and updates the diverse world information, from city-scale planning to individual buildings, allowing generated interiors to maintain explicit correspondence with their associated exterior buildings. Conditioned on the evolving context and previously generated neighboring blocks, HoloWorld autoregressively generates urban exteriors with consistent spatial organization and visual identity across blocks. The generated exterior representations are further grounded in 3D building instances and footprints, enabling building-specific indoor generation with geometry-constrained layouts and inherited appearance characteristics. To our knowledge, HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world. Extensive experiments demonstrate that HoloWorld achieves superior urban exterior generation performance, improving the average AQS score over the SOTA by 7.68\% and obtaining the highest average RDR score, while maintaining strong building-level indoor-outdoor correspondence and cross-block continuity within a unified 3D urban world. Our project page: https://huangxb326.github.io/HoloWorld/.