arXiv:2609.11561nvidia/nemotron-3.5-lightning-30b-a3bSeptember 10, 2026

Memory as Plans: World-Action Modeling with Memory-Grounded PlanningExplained for Beginners

Sizhe Zhao, Haozhe Xie, Weiyu Zhao +5 more

Robotics

Abstract

Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.

Memory as Plans: How MaP-WAM Remembers Without Forgetting

1. The Problem: When Memory Gets in the Way

Anyone who has tried to remember a long sequence of steps without writing them down knows the struggle: the longer the list, the easier it is to lose the plot. Modern robotics faces a similar crisis. Many real-world manipulation tasks—like "find the red button under this stack of boxes" or "swap these two T-shaped blocks"—require the robot to remember things that are no longer visible. Perhaps the button is now hidden under a cover, or the blocks have been moved when you weren't looking.

The paper Memory as Plans: World-Action Modeling with Memory-Grounded Planning (MaP-WAM) identifies that current robotic approaches typically fall into one of two failing strategies.

The first is Markovian thinking: the robot acts only on what it sees right now, completely ignoring the past. This works for simple tasks like "pick up the apple," but it fails immediately when the critical information is currently hidden.

The second is growing memory windows: the robot keeps feeding more and more past observations into its "brain" (the neural network). As the list of observations grows, the system gets slower and hungrier for GPU memory. Worse, it often discards the "fine-grained" visual details—the exact shape or color—that are essential for precision tasks.

The authors argue that neither approach is ideal. Remembering everything is expensive and slow; remembering summaries loses the details. They ask: Can we remember just enough, in a way that is both efficient and precise?

2. How It Works: Memory as a set of "Plans"

MaP-WAM introduces a clever decomposition. Instead of forcing the robot's "executor" (the part that decides what move to make next) to read through a growing history of every frame it has ever seen, MaP-WAM splits the job into two distinct roles: The Planner and The Executor.

Here is the mechanical breakdown, explained with a software engineering analogy:

The Analogy: The "Instruction Manual" vs. The "Worker"

Imagine you are managing a complex assembly line (the robot task).

  • The Old Way (Growing Window): You hand the worker a stack of every single instruction sheet from the beginning of time. As the project gets longer, the stack gets heavier, and the worker spends more time flipping pages than building. Eventually, the stack is so thick the system crashes.
  • The Language-Only Summary: You give the worker a one-page summary of "what has happened so far." It’s lightweight, but you lose the specific diagrams and details. The worker might misunderstand the exact angle to turn a bolt.
  • The MaP-WAM Way (Memory as Plans):
    1. The Planner (The Editor): Every time a major "chunk" of the task is finished (a segment), an editor summarizes just that chunk. They write a short language instruction (e.g., "Left arm moved red block to bin A") and snap a few key photos (sparse visual context). They file this away as a "segment record."
    2. The Plan: The editor then looks at the file of all previous summaries and writes a new, compact plan for the next step. This plan consists of two parts: a language instruction ("Pick up the blue block") and a visual guide (a rough sketch of where the blue block should be based on past evidence).
    3. The Executor (The Worker): The worker only reads this single, short plan. They don't care about the 500 frames that came before. Because the plan is short and fixed, the worker operates efficiently, without needing massive memory.

In technical terms, MaP-WAM maintains a structured multimodal episodic context C<kC_{<k}—a history of "segment records" containing language instructions and sparsely sampled visual frames. At each stage, a Language Planner predicts the next language instruction, and a Causal World Model (CWM) generates corresponding visual guidance. Together, these form a "memory-grounded plan" C^k\hat{C}_k.

Crucially, the Executor (πE\pi_E) only conditions on this static plan. Because the plan doesn't grow with the task history, the executor's context length remains fixed, keeping latency constant regardless of how long the task goes on.

The World-Action-Progress (WAP) Model

Here is where it gets interesting. Even with a great plan, how long should the executor follow that plan before checking in with the planner again? The execution duration is initially unknown.

Enter the World-Action-Progress (WAP) model. Think of this as the executor having a built-in "odometer" and "GPS progress bar."

  • Joint Prediction: WAP doesn't just predict what action to take; it simultaneously predicts how far along the current plan the robot is. It outputs "action chunks" and a "progress score" (from 0% to 100%) at the same time.
  • The GPS Calibration: As the robot executes actions, it looks at the visual plan (the sketch drawn by the Planner). If the robot's current view looks different from the sketch, the WAP model "calibrates" its progress bar. It effectively says, "My progress prediction was 70%, but looking at the plan, the scene looks like it's only at 50%. Let me adjust."
  • Segment Completion: When the predicted progress crosses a threshold (τ = 0.95), the current segment is deemed "complete." The robot pauses, takes a quick look at the real world, files away a new segment record (language + sparse visual context), and the whole cycle restarts with a new plan.

This progress-aware mechanism is vital. It prevents the robot from drifting into nonsense. Because the progress is anchored to the visual plan, the robot can distinguish between two states that look identical but are actually at different semantic stages of the task.

3. Key Results & Benchmarks: The Numbers

The proof is in the performance. MaP-WAM was tested on two main fronts: simulation (RMBench) and real-world robotics.

Simulation (RMBench)

In the RMBench simulation benchmark, which is specifically designed to test long-horizon memory, MaP-WAM achieved an 83.3% success rate. This puts it head and shoulders above the previous best baseline, which scored 79% on average across the tasks.

Translation: On the standard set of memory-dependent manipulation tests used by researchers, MaP-WAM answers roughly 12% more questions correctly than the previous best system. It solves tasks that require remembering observations from minutes ago, even when those observations are no longer visible.

Real-World Robotics

On physical robot arms (Franka robot), MaP-WAM attained a 78.0% success rate across 50 trials per task for two specific memory-dependent tasks: "Find Button" and "Press Buttons."

Translation: This is a significant jump from baseline methods. For instance, on the "Press Buttons" task—which requires pressing multiple buttons in sequence—the best previous method managed only a 19% success rate with 50 demonstrations. MaP-WAM jumped that to 68%. It shows that the robot can reliably execute long-horizon sequences (like "press, press, confirm") without getting lost.

Efficiency: The "Magic" of Constant Latency

Perhaps the most impressive technical result is the latency measurement. As the task history grows (more and more frames are added to the memory), the authors measured the executor's inference time.

  • The Baseline (Full Context): Inference time skyrocketed. With 1,500 historical frames, latency was roughly 4x higher. With 1,700 frames, the system hit an Out-Of-Memory (OOM) error and crashed.
  • MaP-WAM: Maintained an approximately constant latency of ~827ms per action chunk, regardless of whether the task had 500 frames or 1,500 frames.

Translation: This means a factory deploying these robots can scale up to very long, complex tasks without needing to buy exponentially more expensive GPUs. The system remains responsive.

4. Why It Matters: Key Takeaways

This research matters because it solves the "Memory Scalability Problem" for robots. Here are the four key takeaways:

  1. Decoupling Planning from Execution is Key: By separating the "memory" function (planning) from the "action" function (execution), the system avoids the trade-off between remembering everything and remembering nothing. It achieves the best of both worlds.
  2. Sparse Visual Context is Sufficient: The paper shows that you don't need to feed the robot every single frame from the past. Retaining just 8 key frames per segment (sparse visual context) is enough to ground the plans and maintain precision. This drastically reduces storage and computational costs.
  3. Progress Awareness Prevents Drift: The introduction of the WAP model's progress prediction, coupled with plan-observation alignment, is a significant step toward robust long-horizon manipulation. It gives the robot an internal "counters" to ensure it doesn't lose its place in a multi-step task.
  4. Real-World Viability: Achieving 78% success on real robot tasks with only 50 demonstrations per task is a strong signal that this approach is moving out of the lab and toward practical deployment in complex environments.

What to watch for next: The authors acknowledge a limitation: the current system relies on pre-defined segment boundaries (i.e., knowing when one "chunk" of the task ends and another begins). A natural next step is automatic segment discovery—teaching the robot to figure out its own "milestones" in a task without human guidance. Additionally, while the plan-observation alignment is effective, using learned similarity metrics rather than simple pixel differences could make the system even more robust in cluttered, real-world scenes.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →