arXiv:2608.20430nvidia/nemotron-3.5-lightning-30b-a3bAugust 20, 2026

RISE: Adaptive Imagination for World Action ModelsExplained for Beginners

Hongbo Lu, Liang Yao, Chenghao He +5 more

Computer Vision and Pattern Recognition

Abstract

World Action Models (WAMs) improve planning by incorporating future world evolution into action generation, yet existing methods allocate a fixed imagination budget to every scene. We propose RISE (Refining Imagination through SElective Rollout), a system-level adaptive imagination framework that makes sequential Roll/Stop decisions according to the expected planning benefit of continued rollout. At each step, a Latent Evaluator estimates the risk revealed by the current prefix and how much planning could improve if imagination continues, while a Rollout Gate weighs this expected benefit against additional computation cost. Since factual driving logs expose only one realized future, we further construct CounterDrive, a counterfactual dataset with diverse outcomes and risk levels, to enrich future dynamics and provide localized risk supervision. Each retained sample undergoes expert verification and annotation of trajectory validity, incident onset, and causal category, providing a reusable resource for safety-critical world-modeling research. Experiments on NAVSIM and nuScenes show that RISE achieves the best overall planning performance while reducing unnecessary rollout, with additional transfer results supporting its plug-in generality across WAM architectures.

The Problem: The Fixed-Imagination Bottleneck

Imagine you’re navigating a busy intersection. To decide whether to merge, you mentally simulate a few scenarios: "If that car speeds up, I stop. If it slows down, I go." This mental simulation is valuable, but it takes time and cognitive effort. Now imagine if every driver—from someone cruising down an empty highway to someone creeping through a school zone—spent the exact same amount of time and mental energy simulating the exact same number of future scenarios before acting. It would be wasteful for the highway driver and potentially unsafe for the school zone driver if their limited simulation budget runs out too early.

This is the core problem that the paper RISE: Adaptive Imagination for World Action Models addresses.

In the world of World Action Models (WAMs)—AI systems designed to plan actions by imagining future world states—existing methods suffer from a "fixed imagination budget." Whether the road is simple or complex, the model allocates the same amount of computational "thinking time" to every driving scene.

This creates two opposing failures:

  1. Wasted Compute: On simple, open roads, the model wastes precious milliseconds simulating complex futures that don't actually help the decision.
  2. Insufficient Planning: On complex, ambiguous roads (like dense pedestrian crossings), the model might run out of its fixed simulation budget before it has imagined the critical "what-ifs" needed to avoid a collision.

The paper argues that planning should be adaptive: the model should decide step-by-step whether the current mental simulation is "good enough" to act, or if it needs to keep imagining. This decision should be based on the expected benefit of continuing versus the cost of computing more.

How It Works: The RISE Scheduler (A Software Engineering Analogy)

RISE introduces a lightweight "Scheduler" plugged into a standard WAM architecture. Think of the Scheduler as a vigilant project manager for the model's computation budget.

At every step of the imagination process, the Scheduler runs two specialized checks:

  1. The Latent Evaluator (The Analyst): This component acts like a code reviewer or a risk analyst. It looks at the "prefix"—the partial future the model has already imagined—and produces two key reports:

    • Risk Profile: "How risky is the situation currently? Have we spotted any red flags (like a potential collision course)?" It summarizes the risk already revealed by the current prefix.
    • Future Planning Gain Profile: "If we spend more compute imagining the next moment, how much better will our final plan be? Will this extra step actually change the recommended action, or is the current plan already solid?"
  2. The Rollout Gate (The Decision Maker): This is the gatekeeper. It takes the Analyst's reports and weighs them against a simple resource question: "Is the expected improvement in the plan worth the milliseconds of extra computation required to imagine the next step?"

    • If the Gain > Cost, the Gate says "Roll"—the model generates one more future latent step and the Scheduler re-evaluates.
    • If the Gain ≤ Cost, the Gate says "Stop"—the model takes the current prefix and passes it directly to the Planner to generate the final action.

This process repeats sequentially. The model might stop after 0 steps (planning from the current observation), 2 steps, or the maximum possible steps. The "horizon" emerges naturally from these repeated decisions, rather than being pre-set.

Crucially, RISE needs data to learn when to stop. Since real driving logs only show one realized future (the one that actually happened), the model can't learn about "alternative" futures. To fix this, the authors construct CounterDrive, a counterfactual dataset. For every real driving scenario, they generate multiple plausible "what-if" videos using a video generation model (Wan 2.7), featuring different outcomes (e.g., a crash vs. a safe stop) and different risk levels. Human annoters verify these trajectories and label incidents. This gives the model a much richer education in risk, teaching it not just what happened, but what could have happened and how dangerous those alternatives were.

Key Results & Benchmarks: Beating the Fixed Budget

The results are compelling. RISE is evaluated on two major driving benchmarks, NAVSIM and nuScenes, and it consistently outperforms state-of-the-art methods while being more computationally efficient.

  • NAVSIM V1: RISE achieves a PDMS score of 91.5, surpassing the previous best baseline by 0.8 points. It also improves the previous best EP (Execution Precision) by 2.9 points and TTC (Time to Collision) by 1.9 points.
  • NAVSIM V2: RISE hits 90.8 EPDMS, beating the previous best by 0.9 points. It ranks first or ties for first on seven out of nine component metrics, covering safety, compliance, and planning quality.
  • nuScenes: RISE sets a new state-of-the-art, achieving the lowest average L2 error of 0.31 meters and a collision rate of 0.10. This means RISE drives more accurately and hits fewer things than any previous model.

Perhaps most impressively, RISE achieves this with lower rollout cost. In ablation studies, the RISE Scheduler average only 2.40 rollouts per scene, compared to 2.98 for a "Latent Margin" baseline and a whopping 3.0+ for random or full-horizon rollout. It gets to the right answer faster.

Why It Matters: Key Takeaways

  1. Computational Efficiency: By stopping early on simple scenes, RISE saves significant inference time and energy. This is crucial for real-time autonomous driving, where latency budgets are strict.
  2. Safety Through Adaptation: The system isn't just lazy; it stops when it knows enough. On complex scenes with dense traffic or crossing pedestrians, it will run the full rollout, ensuring the model has imagined enough futures to stay safe.
  3. The Power of Counterfactuals: The creation of the CounterDrive dataset is a significant contribution beyond just the RISE algorithm. By providing diverse, verified "what-if" futures, it improves the model's ability to discriminate risk (AUC scores jumped from ~0.5, near random, to ~0.96). This resource is reusable for other safety-critical AI research.
  4. Plug-in Generality: The RISE Scheduler is architecture-agnostic. In experiments, transferring the Scheduler to a different WAM architecture (DAWN) improved that model's performance without needing to rewrite its core brain (Predictor/Planner). This means the "adaptive imagination" trick can be applied to many different future-modeling systems.

What to Watch For Next

  • The Cost-Lambda Trade-off: The system relies on a parameter λ\lambda that balances planning gain against computation cost. Tuning this for different hardware (e.g., a high-end self-driving car computer vs. a mobile robot) will be important.
  • Edge Cases: While the paper shows RISE handles dense traffic well, real-world "edge cases" (rare, bizarre scenarios) might still catch the model off guard, as no dataset can cover every possibility.
  • Scalability: As WAMs are applied to longer horizons (planning minutes into the future rather than seconds) or different embodiments (not just cars), the dynamics of "when to stop" will evolve.

RISE moves the field from a "one-size-fits-all" approach to a smart, adaptive strategy. It teaches AI to know not just how to imagine the future, but when to stop imagining and start acting—which is, ultimately, the essence of intelligent decision-making.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →