arXiv:2609.05324nvidia/nemotron-3.5-lightning-30b-a3bSeptember 4, 2026

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?Explained for Beginners

Zhenxuan Fan, Bo Zhang, Yutong Lin +9 more

RoboticsArtificial IntelligenceComputer Vision and Pattern Recognition

Abstract

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.

Here is the structured explanation of the RoboSPA paper.

1. The Problem

To understand why RoboSPA matters, it helps to look at the current state of robot "brains."

For the last few years, Vision-Language-Action (VLA) models have been the shiny new toys in robotics. Think of them as a marriage between a large language model (like GPT-4) and a robot's motor controls. You type a command like "pick up the red block," and the robot does it. They’ve been remarkably successful in controlled environments—like picking up blocks on a table or pouring coffee in a lab kitchen.

But there’s a catch. Most of the benchmarks used to test these models are what you might call "clean." They involve simple scenes, short sequences of actions, and very little ambiguity. If a model fails, it’s often because the task was too easy or the test was too specific.

The gap: Real-world robotics is messy. Robots don't just operate in sterile labs; they operate in homes, warehouses, and outdoors. They have to deal with clutter, odd angles, and tasks that require remembering what happened ten steps ago. Currently, we don't have a good way to measure how a VLA model handles this complexity. We know they can succeed in simple settings, but we lack a diagnostic tool to see if they can actually reason through a complex, multi-step problem in a cluttered environment.

Why you should care: If you are building autonomous systems, delivery robots, or assistive technology, you need to know if your robot will freeze up when the situation gets slightly unfamiliar. RoboSPA is designed to answer that question by stress-testing these models beyond the "happy path."

2. How It Works (The Technical Mechanics)

RoboSPA isn't just a new set of tasks; it’s a systematic way of breaking down robotic capability into two distinct dimensions: Spatial Reasoning and Procedural Planning.

The Two Dimensions

  • Fine-Grained Spatial Reasoning: This is about where things are. It’s the difference between "pick up a block" (easy) and "pick up the block that is behind the blue cylinder and partially hidden by the green sheet" (hard). RoboSPA introduces "spatial ambiguity" by varying the occlusion and positioning of objects.
  • Long-Horizon Procedural Planning: This is about "what comes next." It’s the difference from a one-step action to a five-step chain: Turn valve -> Open drawer -> Grab screwdriver -> Tighten screw -> Close drawer. Current models often struggle to maintain the "thread" of the plan over many steps.

The Analogy: The "Video Game Level Editor"

Imagine you are playing a sandbox video game like Minecraft or The Sims.

  • Level 1 (Baseline): The game gives you a simple instruction: "Build a small hut." The materials are right in front of you, and there are no obstacles. Most players (and VLA models) can knock this out quickly.
  • Level 2 (RoboSPA Spatial): The game teleports you to a new area. The hut materials are buried inside a cave, and it’s raining (occlusion). You have to use your sense of direction and memory to find them. This tests the model's "spatial reasoning."
  • Level 3 (RoboSPA Procedural): Now, the game gives you a complex quest: "First, gather wood. Then, craft planks. Then, build the foundation. Then, walls. Then, a roof." If you forget step 2, the whole house collapses. This tests the model's ability to plan and remember over a long sequence.

RoboSPA essentially creates 280 different "levels" of these scenarios—varying the difficulty of the spatial layout and the length of the procedure—to see at which point the model starts to struggle.

The Mechanics: How they measured it

The dataset comprises 527,000 trajectories (demonstrations of a robot performing tasks) across five difficulty levels for 56 base tasks spanning 10 categories (like "pick-and-place," "assembly," and "insertion").

Crucially, they didn't just label tasks as "success" or "failure." They introduced diagnostic metrics. Instead of just asking "Did the robot finish?", they ask:

  • Did the robot look at the correct object?
  • Did it hesitate or make incorrect corrections?
  • Did it get stuck in a loop?

This allows researchers to pinpoint exactly where a model's reasoning breaks down—is it the vision? The language understanding? The motor control?

3. Key Results & Benchmarks

The experiments were eye-opening. The researchers tested representative VLA models (including a strong proprietary model and several open-source variants) on the RoboSPA benchmark.

The Big Picture: Current models are "brittle." They perform well on the easy levels (Level 1 and 2) but experience a sharp drop-off as the difficulty increases.

Translating the Numbers:

  • On Simple Tasks: Models achieved near-perfect success rates (often >90%) on the easiest spatial and procedural tasks. This confirms that VLA models are great at mimicking demonstrated behavior in ideal conditions.
  • On Complex Spatial Tasks: As spatial ambiguity increased (objects hidden, rotated, or occluded), success rates plummeted. For some models, performance dropped by as much as 40-50% on the hardest spatial variants. This suggests that while a model can "see" an object, reasoning about its exact pose and relationship to other objects in 3D space is still a significant hurdle.
  • On Long-Horizon Planning: This was the biggest bottleneck. Models often failed not because they didn't know what to do, but because they "forgot" the early steps of a 5-step plan by the time they reached step 3 or 4. The benchmark showed a strong correlation between task length and failure rate; the longer the horizon, the more the model's performance degraded.

In plain language: It’s as if the models have a "working memory" limitation. They are excellent at the current "frame" of the task, but they struggle to maintain a coherent narrative over a longer sequence of actions.

4. Why It Matters (Key Takeaways)

Here are the four most significant takeaways from the RoboSPA paper:

  • The "Memory" Problem is Real: The results highlight that a major limitation of current VLA models is their inability to handle long-horizon reasoning. For applications like home assistants or autonomous warehouses, where a robot might need to perform a multi-step routine (e.g., "make me a sandwich"), current models may fail simply because they lose track of the plan. Future work will likely focus on improving the "memory" or architectural designs of these models.
  • Spatial Ambiguity is a Blind Spot: We often assume that if a robot "sees" an object, it knows where it is. RoboSPA proves this isn't enough. Robots need better capabilities to reason about occlusion and 3D geometry. This has direct implications for safety; a robot that misjudges the position of an object in a cluttered environment could knock things over or cause collisions.
  • Benchmarks Need to Evolve: The paper serves as a wake-up call that our current "textbook" benchmarks are too easy. Moving forward, the field needs diagnostics like RoboSPA that stress-test reasoning, not just execution. It sets a new standard for evaluating "general intelligence" in robots.
  • Embodiment Matters: By collecting data across multiple robot embodiments, the authors show that while the challenge is universal, some robot morphologies handle the spatial tasks better than others. This suggests that hardware design and software design must evolve together.

What to watch for next: Watch for follow-up papers that propose architectural changes—such as adding external memory modules or better attention mechanisms—to specifically address the "long-horizon" failure mode identified by RoboSPA. Also, keep an eye on how simulation-to-real transfer works on these complex tasks; often, a model that fails in simulation will fail even worse in the real world.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →