arXiv:2608.05879nvidia/nemotron-3.5-lightning-30b-a3bAugust 29, 2026

To See a World in a Living Context: Unified Indoor-Outdoor Urban World GenerationExplained for Beginners

Xiaobin Huang, Zilong Huang, Yang Luo +3 more

Computer Vision and Pattern Recognition

Abstract

Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, but these domains are usually synthesized independently, lacking the correspondence required for a coherent urban world. We present HoloWorld, a unified indoor-outdoor urban world generation framework built on a continuously updated cross-scale world context. Initializing from a user description, HoloWorld progressively represents and updates the diverse world information, from city-scale planning to individual buildings, allowing generated interiors to maintain explicit correspondence with their associated exterior buildings. Conditioned on the evolving context and previously generated neighboring blocks, HoloWorld autoregressively generates urban exteriors with consistent spatial organization and visual identity across blocks. The generated exterior representations are further grounded in 3D building instances and footprints, enabling building-specific indoor generation with geometry-constrained layouts and inherited appearance characteristics. To our knowledge, HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world. Extensive experiments demonstrate that HoloWorld achieves superior urban exterior generation performance, improving the average AQS score over the SOTA by 7.68\% and obtaining the highest average RDR score, while maintaining strong building-level indoor-outdoor correspondence and cross-block continuity within a unified 3D urban world. Our project page: https://huangxb326.github.io/HoloWorld/.

The Problem: A City Without Its Insides

Imagine a virtual city rendered with breathtaking photorealism. The streets stretch toward a glowing horizon, the buildings are architectural marvels, and the lighting is perfect. But if you walk up to one of those buildings and push open the door, you find… a generic, empty shell. Or worse, a completely different style of interior than what the outside promised.

This is the current reality of text-driven 3D generation. We have advanced rapidly at creating large-scale outdoor environments—think sprawling cityscapes generated from a single text prompt. We have also become very good at generating detailed indoor scenes—furnished living rooms, cluttered kitchens, and themed offices. However, these two domains exist in isolation. An urban exterior is generated without any awareness of what lies inside its walls, and an indoor scene is generated without any knowledge of which building it belongs to.

For a developer building a video game, a filmmaker previsualizing a sci-fi metropolis, or a researcher simulating autonomous agents, this lack of correspondence is a dealbreaker. It breaks immersion. If the virtual world doesn't behave like a real city—where the inside matches the outside—it feels fragmented and artificial. The paper “To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation” identifies this exact gap: the inability to generate a coherent urban world where every room is explicitly linked to its originating building, maintaining semantic, visual, and spatial consistency across the entire scale, from the city plan down to the individual footprint.

How It Works: The "Living Context" and the Autoregressive City

HoloWorld tackles this by treating the city not as a collection of separate scenes, but as a single, evolving organism. The core innovation is a cross-scale world context—think of it as a living, shared memory that the generation process consults and updates at every step.

Here is the mechanics broken down into a software-engineer-friendly analogy:

  • The Hierarchical Memory (The Context): Imagine a text-based adventure game where the "state" object tracks everything that has happened. HoloWorld maintains a hierarchical state with three layers:

    1. World-Level: The "vibe" of the city. Is it a Victorian steampunk metropolis or a futuristic chrome jungle? This layer defines global semantics, functional organization (where the industrial zones are vs. residential), and the shared visual style (the color palette, the architectural "feel").
    2. Block-Level: The city is divided into a grid of blocks (like city blocks). As the generation proceeds, this layer records the specific layout of each block—where the roads go, where the parks are, and how the buildings align. It ensures that Block B doesn't suddenly look like a random alien landscape next to Block A.
    3. Building-Level: This is the crucial link. When a specific building is generated, this layer locks onto that instance. It records the building's identity, its exact footprint (the shape of its base), its visual materials, and its functional role (is it a warehouse? a residence?).
  • The "Living" Update Mechanism: The context isn't just a static file; it’s dynamic. As each block is generated, the results are "validated" and written back into the context. If a block generates a certain type of streetlamp, that detail is noted so the next block can reference it. This is the "continuously updated" part of the framework.

  • Autoregressive Neighborhood Conditioning: This is the glue that holds the city together. To generate Block 3, the model doesn't just look at the text prompt; it looks at Block 1 and Block 2, which were generated moments before. It asks: "What did the road look like two blocks back? What was the visual style of the previous building?" This prevents the "seam seam"—those visible lines where one artist's style ends and another begins. It uses a concrete conditioning mechanism where the generator looks at previously generated neighbors before producing the current block.

  • Building Grounding: Once the exterior blocks are up, the system identifies specific 3D building instances within those blocks. It extracts the footprint (the 2D shape of the building's base). This footprint becomes a constraint for the interior generation. The interior is not just "a room"; it is a room shaped like the building's base. Furthermore, the interior inherits the appearance characteristics of the exterior—if the outside is red brick with white trim, the interior materials will reflect that same color scheme, though adapted for indoor lighting.

Key Results & Benchmarks: Beating the State-of-the-Art

The paper evaluates HoloWorld using two main metrics: AQS (Absolute Quantitative Score), which rates quality on a scale, and RDR (Relative Dimension Ranking), which compares the model against others in pairwise battles.

The results are compelling. When generating urban exteriors, HoloWorld outperforms the previous best method, MajutsuCity, by a relative margin of 7.68% on the average AQS score. To put that in plain language: if the previous best method was scoring an 8.0 out of 10 on structural consistency and visual richness, HoloWorld is effectively raising that bar to roughly 8.62. It also secures the highest average RDR score across the board, meaning human evaluators (and the AI judge, GPT-5.5) consistently preferred HoloWorld's cities over competitors.

For indoor-outdoor correspondence, the gap is even more significant. The baseline method, TRELLIS, generates interiors independently. HoloWorld, by contrast, generates them grounded in the building. On visual consistency—a measure of whether the inside matches the outside—HoloWorld scores 8.21 using GPT-5.5, while TRELLIS lagged behind at 5.88. That’s a massive jump, representing a significant leap in the model's ability to "remember" the building's style when designing the inside.

Why It Matters: The Takeaways

This framework matters because it bridges the gap between "generating a picture" and "building a world." Here are the key takeaways:

  • Explicit Correspondence: For the first time, you can generate a city and be confident that if you model the interior of Building A, it will match the architectural style and function of Building A's exterior. This is huge for game devs who need interior/exterior consistency without manual texturing.
  • Spatial Coherence: The autoregressive neighborhood conditioning means cities feel "whole." Roads flow, architectural styles transition naturally, and the city doesn't feel like a patchwork of unrelated generator outputs.
  • The "Living Context" is Key: The ablation studies (removing the dynamic updates) show that without this continuous feedback loop, the interiors degrade significantly. The initial text prompt alone isn't enough to maintain coherence across scales; the system must update its internal state as it generates.
  • Future Roadmap: The authors acknowledge a current limitation: the framework currently focuses on single-floor interiors and doesn't explicitly model verticality (apartments, mezzanines). Future work will target multi-floor architectural reasoning.

Summary

HoloWorld represents a shift from independent scene synthesis to unified world modeling. By introducing a cross-scale "living context" and autoregressive neighborhood conditioning, it ensures that a generated city is more than the sum of its parts—the streets, buildings, and interiors all belong to the same coherent reality. It improves upon state-of-the-art metrics by single-digit percentages, but more importantly, it solves the practical problem of disconnected indoor-outdoor generation, paving the way for more immersive virtual environments, better simulation datasets for robotics, and a new standard for text-to-3D worldbuilding.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →