SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout EvolutionExplained for Beginners
Xingjian Ran, Xiaoye Mo, Sihao Liu +3 more
Abstract
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-language models (VLMs) can generate scenes of high fidelity but require costly iterative object placement and refinement. Another mainstream paradigm, parametric image-to-3D scene models, produces scenes efficiently from strong priors learned from 2D images but often leads to imprecise and physically invalid scenes. More importantly, both paradigms struggle to output diverse scenes for a single input, making it hard for them to reflect the dynamically changing nature of real scenes. In this paper we propose SceneMosaic, a framework that combines the merits of both paradigms. It obtains the initial candidate from the learned image-based prior, and subsequently evolves the result through VLM agents, ensuring both efficiency and physical validity. Within the evolution process, SceneMosaic exploits the locality of natural scenes and decomposes a scene into independent local units, allowing separate evolution within each unit before composing the global scene via Cartesian product. On SceneEval-100, SceneMosaic matches the strongest agentic baseline in semantic layout quality with a 24x speedup, substantially reduces physical violations, and receives the highest human ratings. Our code is publicly available at https://github.com/rxjfighting/SceneMosaic.
SceneMosaic: Growing Diverse Worlds in a Flash
The Problem
Let’s start with a frustration anyone who has tinkered with 3D engines or AI research knows well. You want to populate a virtual apartment with furniture, or you need a bunch of varied office layouts for a robot to practice navigating. Generating these scenes used to be a choice between two unsatisfactory options.
On one side, you have the "artisanal" approach using agentic text-to-3D pipelines. Imagine a team of digital architects, each powered by a vision-language model, meticulously placing every lamp, chair, and plant. The results can look incredible—high fidelity, physically plausible, and semantically sensible. But it’s slow. Very slow. Each object requires iterative refinement, feedback loops, and verification. If you need a hundred different scene variations for a stress test, you’re looking at days of compute time.
On the other side, you have the "industrial" approach using parametric image-to-3D models. These are efficient—they can spat out a scene in a heartbeat by learning priors from 2D images. But they often produce scenes that look a bit... "off." Maybe a table is floating three inches off the floor, or a TV is levitating above a sofa. They are efficient, but physically invalid. Worse, if you ask them to generate diversity—to give you five different layouts for the same prompt—they tend to give you the same living room over and over, just with slightly different wallpaper. Real-world scenes change dynamically; AI needs to capture that variety.
The gap this paper addresses is the need for speed, diversity, and physical validity all at once. The authors ask: Can we have the efficiency of the parametric models with the control and quality of the agentic models? That is the core tension driving SceneMosaic.
How It Works (The Technical Mechanics)
So how does SceneMosaic bridge this gap? The secret sauce is a "hybrid" philosophy. Instead of starting from scratch or relying on a single brittle model, SceneMosaic acts like a creative director overseeing two teams.
Here is the workflow broken down:
-
The Seed (The Learned Prior): The process begins with an image-based prior. Think of this as the "rough sketch." SceneMosaic takes a semantic layout or a rough description and uses a parametric model to generate an initial, efficient candidate scene. This gives us a starting point that already has the basic furniture and big shapes in the right general zones.
-
The Evolution (The Agentic Refinement): This is where the "agentic" part comes in. Rather than leaving the AI to guess the whole room at once, SceneMosaic unleashes Vision Language Model (VLM) agents—but with a crucial constraint. These agents don't try to redesign the whole house; they focus on local units.
-
The Locality Trick (Decomposition): This is the cleverest part. The authors observe that in real indoor scenes, things are local. The sofa is near the TV; the lamp is on the side table. SceneMosaic decomposes the big scene into these independent local units. Imagine chopping the floor plan into distinct "zones" or "rooms."
-
The Cartesian Product (Composition): Here is the magic. The agents evolve each local unit independently. Because these units are treated as separate entities, the system can explore massive diversity within each zone. Once all the local units have been "evolved" and refined, SceneMosaic composes the final global scene using a Cartesian product.
What does this mean in plain English? If Zone A has 3 possible sofa styles and Zone B has 2 possible coffee table styles, the system can instantly generate unique combinations. It doesn't have to re-evaluate the whole room for every single variation; it just mixes and matches the best parts of each local evolution. This is what enables the diversity—you get many valid configurations without paying the computational cost of regenerating the whole scene from scratch every time.
- Physical Validity: Throughout this evolution, the agents check for physical violations (like that floating table mentioned earlier). Because the refinements happen locally and then composed, the system maintains a high rate of physically plausible scenes.
Analogy for the Software Engineer: Think of SceneMosaic like a React component tree with a "state manager." You have a base "Scene" component rendered quickly using static props (the image-based prior). Then, you have "Agent" components that manage specific "local state" (a table, a chair). These agents can swap props or change positions independently. Finally, the system renders the final page by combining the states of all components via a product of states. You get the initial render instantly (efficiency), and the agents can iterate on specific parts without breaking the whole app (locality/diversity).
Key Results & Benchmarks
On the benchmark SceneEval-100, SceneMosaic demonstrates a remarkable trifecta of performance, speed, and physicality.
- Layout Quality: It matches the strongest agentic baseline in semantic layout quality. This means the AI understands where things go just as well as the heavy-duty agentic models.
- The Speedup: This is the headline number. SceneMosaic achieves a 24x speedup compared to the strongest agentic baseline. If the previous state-of-the-art took a full day to generate a set of scenes, SceneMosaic does it in an hour. That is a game-changer for interactive applications.
- Physical Violations: It "substantially reduces physical violations." In plain language, this means fewer floating tables and levitating chairs. The scenes are much more likely to be something a physics engine in a game or robot simulator could actually handle without glitching.
- Human Ratings: Perhaps most importantly, human evaluators preferred the output of SceneMosaic. When shown scenes side-by-side, people rated SceneMosaic highest. This suggests that the "hybrid" approach doesn't just score well on metrics; it actually looks and feels right to people.
Why It Matters (Key Takeaways)
To wrap up, here are the four big takeaways from SceneMosaic:
- The "Best of Both Worlds" Workflow: We no longer have to choose between slow, high-quality agentic generation and fast, physically questionable parametric generation. SceneMosaic proves you can have efficiency and validity.
- Diversity via Composition: The real innovation here isn't just making one good scene, but making many good scenes. By leveraging the Cartesian product of local units, the system generates diverse layouts for a single input. This is crucial for simulation training, where a robot needs to see many variations of a room to learn robust navigation skills.
- The Power of "Locality": The insight that natural scenes are composed of semi-independent local units is key. It’s a recognition of how humans actually design and think about spaces—we don't redesign the whole house every time we move a lamp.
- The Speed Ceiling: That 24x speedup is significant. It brings simulation-ready scene generation into the realm of real-time or near-real-time application. We should watch how this framework integrates into existing game engines and robot simulators; it could drastically accelerate the "world-building" phase of development.
SceneMosaic isn't just a incremental improvement; it's a shift in how we conceptualize 3D scene generation. By respecting the structure of real spaces and combining the strengths of different AI paradigms, it clears a major hurdle for interactive entertainment and embodied AI.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →