arXiv:2609.07398nvidia/nemotron-3.5-lightning-30b-a3bSeptember 7, 2026

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model PretrainingExplained for Beginners

Yuran Wang, Siqiao Huang, Mingleyang Li +21 more

Robotics

Abstract

World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

1. The Problem

Most of us take for granted how effortlessly a child learns to navigate the world. A toddler watches a ball roll, realizes it will hit the wall, and learns to nudge it with a toy hammer. They are building a "World-Action Model"—an internal simulation of cause and effect. In robotics and artificial intelligence, we want machines to develop this same intuition. We want robots that can look at a messy kitchen counter and instinctively know which tool to pick up, or virtual agents that can solve complex tasks without needing a manual for every single move.

However, getting AI to develop this "common sense" has been surprisingly difficult. The current state of the art often relies on monolithic systems. Imagine a Swiss Army knife where the blade, the corkscrew, and the tweezers are all welded together into one solid piece of metal. If you want to swap out the tweezers for a better pair, you have to destroy the whole knife. That is how many World-Action Models (WAMs) are built today. The AI model that generates videos, the way it sees pixels, the way it remembers past events, and the way it decides on an action are all glued together tightly. Researchers can’t tell if a success came from a clever new algorithm or just from the fact that they used a specific, expensive video generator.

This lack of modularity slows down progress. It makes it hard to answer fundamental questions: Does the model need to see hours and hours of video, or just a specific kind of video? Does it need to "actually" move to learn, or can it just watch? And crucially, once we build a great model, can we actually put it on a real robot, or is it stuck in a simulation?

2. How It Works (The Technical Mechanics)

Enter OpenWAM, an open-source "stack" designed to treat building these models like building with LEGO bricks rather than carving a statue from a single block of marble.

The authors break the problem down into modular components housed within OpenWAM-Infra. Think of this as the chassis and the electrical system. It standardizes how different parts talk to each other. It unifies the training loop, the way the model is run (inference), how it’s deployed on a device, and how we measure its success. This means a researcher can swap out the "brain" (the generative model) or the "senses" (the visual representation) without having to rewrite the entire program from scratch.

Now, the interesting part is what happens inside these modules. The paper investigates three core questions through OpenWAM-Study.

  • What to inherit? The model needs "world knowledge"—an understanding of physics, object permanence, and geometry. The authors find this transfers best from a "sufficiently capable generative backbone" (a strong video generator) and a "compact, information-rich latent space." In plain English: you need a good eye to understand the world, and you need to compress that visual information efficiently so the brain isn't overwhelmed by pixels.

  • How do world and action learning interact? This is the crux of the paper. The authors argue that "world-action synergy" doesn't happen by accident. It requires three specific things:

    1. Dedicated action capacity: The model needs a specific channel or "budget" to learn how to move, separate from just watching.
    2. Explicit world-to-action information flow: The model needs a direct line of communication from what it sees to what it does. It can't just magically know how to act; the architecture must route that information.
    3. Synchronized joint denoising: This is the clever technical bit. The model is trained to remove noise from both images and actions simultaneously. It’s like teaching a student to draw a picture and solve a puzzle at the same time, reinforcing both skills.
  • Embodiment and generalization: Does training a robot in a simulator make it good at real-world tasks? The paper suggests embodied pretraining primarily improves "out-of-domain generalization." This means if you train on a mix of human egocentric videos (like GoPro footage) and robot data, the model becomes much better at handling situations it hasn't seen before, especially across different types of robots.

Putting these principles together, they constructed OpenWAM-α. This isn't just a theoretical exercise; it’s a real, trained model. It was trained on roughly 6,400 hours of video data—mostly egocentric (first-person) footage from humans and robots. This diverse diet of data is key; it gives the model wide "world coverage" and grounds its actions in real physics.

3. Key Results & Benchmarks

So, does OpenWAM-α actually work? The authors put it through its paces across eight different simulation benchmarks and real-world robot experiments. The results are compelling.

Across these diverse tests—ranging from simple single-arm reaching to complex bimanual manipulation and even dexterous hand tasks—OpenWAM-α consistently performs at a top-tier level. The authors report that it sustains its "top-tier standing" as it moves from the simulated world into the physical robot world. This is significant because usually, a model that wins in simulation struggles when faced with the messy, unpredictable physics of the real world.

While the paper provides specific numbers (like success rates improving by X% over baselines), the most striking takeaway is the consistency. The model doesn't just perform well in one specific test; it is a generalist. It shows that the modular design and the principles of world-action synergy actually pay off. The "one-stage co-training" over both human and robot data proved to be the secret sauce for transferring skills across different embodiments.

4. Why It Matters (Key Takeaways)

  • A New Toolkit for Researchers: By open-sourcing the infrastructure and evaluation protocols, the authors are lowering the barrier to entry. Future researchers can now experiment with "designer" WAMs—mixing and matching components to see what works—without building everything from scratch. It democratizes access to cutting-edge world-model research.
  • The Path to Generalist Robots: The principle that "embodied pretraining improves out-of-domain generalization" is a major step toward robots that don't need to be retrained from scratch for every new task or environment. A robot equipped with a model like OpenWAM-α could potentially transfer skills from a simulated kitchen to a real kitchen, or from human demonstrations to robot execution, more seamlessly.
  • Watching the "Action Capacity": Readers should watch how the field handles the balance between "world knowledge" (watching videos) and "action grounding" (learning to move). The paper suggests that simply watching more video isn't enough; the architecture must explicitly dedicate resources to learning how to act. Future models will likely focus on optimizing this specific bandwidth.
  • The Data Mix: The finding that co-training on both egocentric human data and robot data yields the best results highlights a trend. To make truly robust AI, we need a diverse diet of data. Relying solely on internet text or standard video datasets may leave gaps that real-world robot interaction fills.

Summary

OpenWAM addresses the "black box" nature of current World-Action Models by providing a modular, open framework. Its core insight is that world knowledge and action learning must be deliberately synchronized through specific architectural choices and a diverse dataset. The result, OpenWAM-α, is a versatile model that bridges the gap between simulation and reality, offering a promising foundation for the next generation of generalist robots.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →