arXiv:2608.23070nvidia/nemotron-3.5-lightning-30b-a3bAugust 24, 2026

From Generation to Simulation: How Far Are World Models from Being True Simulators?Explained for Beginners

Tong Wang, Huan Deng, Mucheng Yang +3 more

Artificial IntelligenceComputer Vision and Pattern Recognition

Abstract

With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engines, and reinforcement-learning environments. Yet the remaining distance from generation to simulation lacks a systematic assessment. We present a capability-based study using an external yardstick: eight capabilities of a traditional simulator, namely asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, and evaluation metrics. We trace three main technical routes--latent dynamics, video generation, and joint-embedding prediction--and map exactly 200 representative works published from 2018 to June 2026 onto these capabilities. Our analysis shows that world models have achieved functional substitution in interaction and controllability for specific scenarios, but remain short of traditional simulators in formal guarantees of physical laws, structured state feedback, and reproducible long-horizon evolution. State feedback is the most neglected cross-route shortcoming: only 6 of 163 implementation papers expose a runtime interface for querying entity states or physical parameters. We identify six research directions: formalized physics, a unified action interface, first-class state feedback, long-horizon stability, downstream-utility evaluation, and cross-route hybridization. Project page: https://github.com/AtongWang/world-model-simulators

From Generation to Simulation: How Far Are World Models from True Simulators?

1. The Problem

Imagine you are building a self-driving car. Before those vehicles ever hit the road, engineers need to test them in a virtual environment—a simulator—that can faithfully reproduce how a car handles on icy roads, how traffic reacts to a sudden brake, or how a pedestrian might step out from behind a bus. Traditional simulators, built on physics engines like PyBullet or game engines like Unreal, have served this role for decades. They are the "digital twins" of the physical world, allowing us to run millions of safety tests without risk.

Recently, a new contender has emerged: generative world models. Powered by diffusion models and large-scale video generators (think of systems like Sora or Google's Genie 2), these AI models can create stunningly realistic videos of cars driving down a street or robots manipulating objects. The allure is obvious. A traditional simulator requires hand-coding physics laws and building 3D assets—an expensive, time-consuming engineering effort. A generative model, once trained, can seemingly "hallucinate" a functioning world from scratch, offering the promise of data-driven, general-purpose simulation that is cheaper and more flexible.

But can they actually replace the real thing?

This paper sets out to answer that question with systematic rigor. The authors argue that while generative world models have made incredible strides in looking like simulations, there hasn't been a standard way to measure how well they actually function as simulators. Existing surveys tend to organize papers by architecture (e.g., "this one uses diffusion") or application (e.g., "this one is for robotics"), but they fail to ask the fundamental question: Does this model actually behave like a simulator?

To answer this, the authors deploy an "external yardstick." They look at eight core capabilities that define a traditional simulator and check which of those capabilities the world models possess. These eight capabilities are:

  1. Asset Construction: Can the model build the world (maps, objects, characters)?
  2. Physics Engine: Does it obey the laws of physics (conservation of energy, collision rules)?
  3. Interaction: Can agents interact with the environment in real-time?
  4. Controllability: Can a user precisely control the environment's configuration?
  5. Stability: Is the simulation reproducible? (Give the same input, get the same output.)
  6. State Feedback: Can the model provide rich, structured information about what's happening inside the "black box" at any moment? (e.g., the exact position of a car, the force of a collision).
  7. Diversity: Can it generate varied scenarios?
  8. Evaluation Metrics: Does it have tools to measure performance?

The paper's central problem is a gap analysis: Where do these generative models succeed, and where do they crumble compared to the rigorous demands of a traditional simulator?

2. How It Works (The Technical Mechanics)

The authors trace the evolution of world models along three distinct "technical routes." Think of these as three different architectural philosophies for building a virtual world.

Route 1: Latent Dynamics (The "Internal Map" Approach)

This is the earliest approach, pioneered around 2018. The model compresses the high-resolution video of the world into a compact, low-dimensional "latent space"—a sort of compressed ZIP file of the environment's state. It then learns a dynamics model (often using RNNs or Transformers) to predict how this latent state evolves over time.

  • Analogy: Imagine a video game that doesn't store every pixel, but instead stores a simplified "state sheet" (e.g., "player at X,Y, health 80"). Instead of rendering the full scene every frame, the game predicts the next state sheet. This saves massive computing power but risks losing fine details.

Route 2: Video Generation (The "Pixel Painter" Approach)

This route, which surged forward with diffusion models around 2023, predicts the future directly in pixel or token space. If latent dynamics is a state sheet, video generation is painting the next frame on a canvas. The model takes the current frame and predicts what the next frame should look like.

  • Analogy: This is like a video compression algorithm in reverse. Given a starting frame, it "guesses" the next frame by looking at patterns it learned from millions of videos. It can produce incredibly realistic visuals, but it is essentially guessing pixel by pixel.

Route 3: Joint-Embedding Predictive Architecture (JEPA) (The "Abstract Thinker" Approach)

This is the newest kid on the block. Instead of predicting pixels or even latent states directly, JEPA models predict "abstract embeddings"—high-level representations of what should happen. It answers the question: "Given I took action A, what is the abstract concept of the next state?" This allows for much more efficient planning because the model doesn't have to deal with the messiness of raw pixels.

  • Analogy: Imagine a flight simulator where, instead of rendering the outside world in high definition, the system just tracks abstract instruments: "Altitude: 10,000 ft," "Speed: 500 knots," "Heading: North." The pilot (or another AI) makes decisions based on these instruments. It's incredibly efficient and precise, but you can't see the clouds or the runway unless you have a separate renderer.

The Convergence Trend

A key finding of the paper is that these routes are not isolated silos. Since 2024, we've seen "hybrid" models. For example, a model might use a JEPA-style abstract planner to decide what should happen, but then use a diffusion model to render the actual pixels of that outcome. The authors argue this convergence is necessary for world models to truly become simulators.

3. Key Results & Benchmarks

The authors mapped 200 representative papers onto the eight simulator capabilities. The results, visualized in Figure 1 of the paper, paint a picture of selective competence.

Where they shine (Strengths)

The paper highlights three capabilities where world models have made significant inroads:

  • Controllability (62.5% coverage): Many models allow users to control the simulation via text prompts, keyboard/mouse inputs, or latent actions. For instance, models like UniPi or DrivingGPT let you steer a virtual car simply by typing "turn left."
  • Interaction (40% coverage): There is meaningful progress here, particularly at the "camera level" or for specific object interactions. GameNGen, for example, can generate interactive moments in a Doom-style game.
  • Stability (40% coverage): Some models, particularly the latent-dynamics ones like DreamerV3, are praised for being reproducible. If you set the random seed, you get the same rollout every time.

Where they fall short (Structural Gaps)

The most critical gaps are in three areas, listed from most to least covered:

  • Asset Construction (19% coverage): Most world models generate a scene "from nowhere." They aren't systematically building a structured world of editable assets. You usually get a fixed video loop, not a editable 3D scene where you can pick up and move individual objects.
  • Physics Engine (17% coverage): This is a stark finding. While models can look like they are obeying physics (a ball falls down), they often fail at "hard physics." The paper notes that collision behavior often violates physical intuition, and energy/momentum conservation is not guaranteed. The models learn the distribution of visual patterns, not the invariant laws of physics.
  • State Feedback (22.5% coverage): This is the most neglected area, and the one the authors flag as the most critical failure.

The State Feedback Crisis

The authors performed a deep audit of 163 implementation papers. Their definition of "state feedback" is essentially: Can you ask the model, "What is the current position of the agent?" or "What is the friction coefficient of this surface?" and get a reliable answer?

The result was sobering: Only 6 out of 163 papers (about 3.7%) expose a runtime interface for querying entity states or physical parameters.

To put this in plain language: If you are building an AI agent that needs to know "Is the robot arm blocked?" to make a decision, a world model with poor state feedback would have to guess based on the pixel output. A traditional simulator would instantly give you the exact coordinate and collision status. The paper notes that this lack of feedback means world models are "conditional-distribution samplers rather than physical-evolution solvers." They can generate a plausible video of a crash, but they can't tell you the exact kinematics of the crash.

4. Why It Matters (Key Takeaways)

The paper concludes with six research directions, but we can distill the broader significance into four key takeaways:

  • The "Realism" Trap: We must stop confusing visual realism with simulation fidelity. A world model can generate a video of a car driving off a cliff that looks absolutely real, but if it doesn't faithfully simulate the physics of the fall, it is useless for safety testing. The paper warns that current models are "conditional-distribution samplers," not "physical-evolution solvers."
  • The Transparency Problem: The lack of state feedback is a major blocker for real-world deployment. Autonomous systems need to know the exact state of their environment to make safe decisions. Without a way to query "What is the pose of the robot?" or "Is the trajectory collision-free?", these models are "black boxes" that are risky to deploy in critical scenarios.
  • The Long-Horizon Instability: While models are good at predicting the next second of a video, they fall apart over long rollouts. Objects disappear, gravity switches off, or the scene morphs unpredictably. For a simulator, we need stability—the guarantee that if you press "accelerate" for 10 seconds, the car moves forward 100 meters, not vanishes or teleports.
  • The Road Ahead: The authors identify six critical research directions to watch. The most actionable seem to be formalized physics (embedding hard physical laws into the model architecture, not just learning them from data) and first-class state feedback (building the ability to query internal states into the model's core interface).

In summary, generative world models have crossed a threshold where they can mimic simulators for specific, narrow tasks—particularly in generating visually plausible interactions and allowing high-level control. However, they remain a significant step away from being "true simulators" capable of providing the rigorous physical guarantees, structured feedback, and long-horizon stability required for safety-critical applications like autonomous driving or robotic manipulation. The path forward isn't just making prettier videos; it's building models that understand the physics of the world well enough to reliably predict its future.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →