arXiv:2609.05416nvidia/nemotron-3.5-lightning-30b-a3bSeptember 4, 2026

WorldSculpt: Generating Compositional Worlds from Grounded VideosExplained for Beginners

Muyao Niu, Jixuan He, Ruihan Yu +9 more

Computer Vision and Pattern Recognition

Abstract

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.

Here is the structured explanation of the WorldSculpt paper.

1. The Problem: From Undifferentiated Geometry to Editable Assets

Imagine you are building a video game or an AR experience. You need a 3D scene, but simply having a "blob" of geometry isn't enough. You need individual objects—a chair, a vase, a coffee cup—so that the player can pick up the cup, move the chair, or drop the vase into a physics simulator.

Current generative AI models, like Marble or HY-World 2.0, are incredible at synthesizing entire worlds from a text prompt or a single image. However, they typically output the scene as a unified mesh or a cloud of points (like 3D Gaussian Splatting). In this format, the individual objects are fused together. You can’t select "the vase" and move it without moving everything else. This creates a fundamental mismatch with the needs of modern applications.

WorldSculpt addresses this gap. The paper asks: Can we generate a complex scene containing hundreds of objects, but represent each one as a separate, independent mesh? This is difficult because in cluttered scenes, objects block each other. If you only see the front of a table, the legs hiding behind it are invisible to a single camera. Traditional geometry-based methods often leave those hidden parts as holes. Previous generative methods are usually limited to simple scenes, like a few objects on a table.

The core problem WorldSculpt solves is: How can we generate a compositional 3D scene (hundreds of separate meshes) from multi-view videos of densely cluttered environments, where objects heavily occlude one another?

2. How It Works: The "3D Jigsaw Puzzle" Approach

WorldSculpt doesn't try to reconstruct the scene all at once. Instead, it tackles one object at a time, using a clever "divide and conquer" strategy.

The Core Analogy: The Anchor and the Cube

Think of each object in the scene as a present inside a box. The system needs to figure out the shape of the present, even if the box is tilted or partially covered.

  1. Anchor View & Canonical Frame: For every object, the system picks one "best" view (the anchor)—the one where the object is most visible. It then constructs a virtual, standardized cube (a "canonical frame") around that object. Regardless of how the object was actually positioned in the real world, the system temporarily "reorients" it to fit inside this standard cube. This normalization is the key that allows the AI to use its training.
  2. Crop-Aware Projections: The system takes the other video frames showing that object and digitally "crops" them. It removes the background and warps the image so the object fits perfectly onto the front of the canonical cube. This ensures that when the AI looks at View 1 or View 2, it is always looking at the same canonical orientation.
  3. Feature Lifting: The system feeds these cropped images into a visual encoder (DINOv3), extracting "features"—digital descriptions of what the object looks like from each angle.

The "IBR" Fusion: Weighing Multiple Opinions

Now the system has several descriptions of the same object from different angles. How does it combine them?

WorldSculpt uses a module inspired by Image-Based Rendering (IBR). Imagine you are trying to figure out what a hidden object looks like. You ask five friends who have seen it. One friend says "it’s pointy on top," another says "it has a blue handle," and a third says "it’s mostly round."

  • Arithmetic Mean: The system averages everyone's opinion. If three say "round" and two say "pointy," it goes with "round."
  • Learned IBR Aggregator: WorldSculpt has a learned "judge" that doesn't just average. It figures out who is reliable. Maybe the friend who saw the object in good light is weighted more heavily than the friend who saw it in the dark.

This "learned IBR aggregator" is the secret sauce. It fuses the features from multiple views, giving more importance to views where the object is clearly visible and less to views where it is hidden or noisy.

The Generator: Pixal3D

The fused 3D features are then fed into Pixal3D, a powerful 3D generative AI model. Importantly, Pixal3D was originally trained on single objects in isolation. WorldSculpt "freezes" Pixal3D's core knowledge but adds LoRA (Low-Rank Adaptation)—tiny, efficient adjustments—to let the model understand the multi-view context.

The result? The AI generates a complete 3D mesh of the object inside the canonical cube, filling in the parts that were occluded in the videos, because it has been "told" (via the multi-view features) where those parts likely belong.

Placement: Canonical to World

Finally, the generated mesh is transformed back into the real world. Using the original box localization data, the system places the mesh back into the scene's coordinate frame. Now you have a collection of individual, separate meshes sitting in a shared world space.

3. Key Results & Benchmarks: Better Than the Baselines

The paper evaluates WorldSculpt across three settings, and the results show a clear trend: the more cluttered the scene, the bigger the advantage.

Single-Object Stress Tests (Toys4k)

In controlled tests where the model generates single objects from 1 to 16 views, with occlusion (blocking) ranging from 0% to 75%:

  • The Trend: With just one view, WorldSculpt is competitive with existing single-view models. But as more views are added, it pulls ahead significantly.
  • Occlusion Robustness: Even with 75% of the object hidden per view, WorldSculpt remains stable if enough views are provided. Competing baselines (like TRELLIS) degrade rapidly as occlusion increases. The paper notes that with 16 views, WorldSculpt’s distance errors remain stable, while others plummet.

Controlled Multi-Object Scenes (Toys4k-Scene & HouseCat6D)

  • Toys4k-Scene: This synthetic dataset places many objects in cluttered layouts. WorldSculpt outperforms methods like ShapeR and SceneGen, particularly in recovering fine details (like thin wires or intricate shapes) that other methods smooth over or miss.
  • HouseCat6D: This real-world dataset is relatively "sparse" (fewer objects, mild occlusion). Here, WorldSculpt still wins, demonstrating that the multi-view training generalizes well to real camera footage, not just synthetic renders.

The Big Test: UE-MeshyScene

This is the paper's flagship benchmark. It contains six large-scale Unreal Engine scenes (like a Japanese school or a desert town) with hundreds of objects (up to 701 per scene) and complex mutual occlusion.

  • The Metric: The paper uses Chamfer Distance (lower is better) and F-Score (higher is better).
  • The Win: WorldSculpt outperforms the previous best method, ShapeR, across the board.
    • Example Impact: On the UE-MeshyScene benchmark, WorldSculpt reduces the Chamfer distance (a measure of geometry error) by roughly 12% compared to the previous best method.
    • F-Score: WorldSculpt achieves an F-Score of 0.951, meaning roughly 95% of the generated surface is within a reasonable distance of the ground truth, compared to 0.944 for the runner-up.
  • Median Performance: The improvements are consistent. It’s not just fixing one or two tricky objects; the median object in the scene is significantly better recovered.

4. Why It Matters: Key Takeaways

Here are the four most important takeaways from this work:

  1. From "Lumps" to "Legos": WorldSculpt enables the creation of compositional 3D scenes. For the first time, large, cluttered environments can be generated as individual meshes. This means downstream apps—like robotics simulators or level editors in games—can actually use the individual objects, rather than just looking at a static picture.
  2. The Power of Multi-View: The paper demonstrates that a model trained on single objects can generalize to hundreds of objects without any scene-level training. By simply providing multiple video angles, the system can recover geometry that would be invisible in a single frame. This is a major scalability win.
  3. The UE-MeshyScene Benchmark: The authors have released a new, rigorous benchmark (UE-MeshyScene) for the community. Because it provides ground-truth meshes for hundreds of objects in complex clutter, it will serve as a standard testbed for future research in scene compositional generation.
  4. Limitations to Watch: The method relies on "upstream" estimates of where objects are and what their masks are. If the input video tracking is poor (e.g., objects are heavily camouflaged or moving fast), the output geometry will suffer. Additionally, the current version generates geometry only—appearance (textures/materials) is not yet generated, though the architecture supports it for future work.

In summary: WorldSculpt bridges the gap between impressive text-to-world generative AI and the practical needs of developers and researchers. It proves that we can take a smart single-object AI, give it a few video angles, and get back a whole, editable, cluttered scene.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →