arXiv:2608.15875nvidia/nemotron-3.5-lightning-30b-a3bAugust 16, 2026

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System ArchitectureExplained for Beginners

GigaBrain Team, Angen Ye, Axiang Sun +56 more

Robotics

Abstract

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.

The Problem: Scaling Generalization in Embodied AI

Imagine teaching a robot to clean a kitchen. You show it videos of humans washing dishes, stacking groceries, and organizing pantries. The robot gets remarkably good at those specific tasks. But then you ask it to fold a crumpled shirt or handle a uniquely shaped mug it’s never seen before, and it stumbles. The problem isn’t just a lack of data—it’s that today’s Vision-Language-Action (VLA) models are often brittle when faced with new embodiments, new object types, or instructions that deviate from their training distribution.

This is the gap GigaBrain-0.7 addresses. The paper identifies that despite the success of models like π₀.₅ and OpenVLA, current VLA systems struggle with three interconnected challenges:

  1. Architectural rigidity: Most VLA models couple vision-language understanding and action generation in a way that doesn’t scale well when the robot’s body changes. A model trained on a 7-degree-of-freedom arm often fails catastrophically on a humanoid with 32 degrees of freedom.

  2. Data heterogeneity: Collecting thousands of hours of robot data is expensive and often limited to a single robot type or environment. Prior work has primarily focused on scaling within one embodiment, making it hard to transfer learning across different morphologies.

  3. Partial coordination: Many systems treat understanding, prediction, and action as separate modules trained sequentially. This “pipeline” approach means errors compound—if the world model mispredicts, the policy receives bad signals; if the policy makes a mistake, the world model never learns from it.

The paper’s core thesis is that we need a unified system that can reason about tasks, predict how the world will evolve, and act across diverse robots—all while preserving the general vision-language capabilities that let the model understand novel instructions.

How It Works: The Three-System Architecture

GigaBrain-0.7 organizes embodied intelligence into three distinct but tightly coupled systems. Think of it not as a single monolithic model, but as a collaborative team of specialists.

System 1: Action and Control (The “Hands”)

System 1 is the engine that turns language and vision into movement. It’s built on a PaliGemma2 (3B) vision-language backbone coupled with a 0.5B Action Expert in a Mixture-of-Transformers (MoT) design.

The key innovation here is the dual-stream architecture. In standard VLA models, the action expert either piggybacks on the VLM’s tokens or sits at the very end, getting only the final representation. GigaBrain-0.7’s dual-stream approach interleaves the two streams layer-by-layer. The vision-language stream uses causal self-attention (preserving its language understanding), while the action expert attends bidirectionally to both streams. This allows the action expert to be grounded in rich semantic context without “polluting” the VLM’s language capabilities.

Embodiment-aware control: A major hurdle in scaling to multiple robots is that different robots have different numbers of joints, different end-effectors, and different coordinate systems. GigaBrain-0.7 solves this with an “embodiment-aware state interface.” It provides the model with both proprioceptive state (joint positions, gripper state) and a Robot ID. The proprioceptive state is masked by a per-embodiment validity vector—distinguishing “this dimension doesn’t exist on this robot” from “this dimension exists but is zero.” This means the core manipulation knowledge is shared across robots, while embodiment-specific details are handled by lightweight projection layers.

Discrete and continuous pathways: System 1 supports two action representations that share the same visual-language context:

  • Discrete pathway: Generates action tokens autoregressively via the VLM head. This is useful for training and symbolic reasoning.
  • Continuous pathway: Generates action chunks through conditional flow matching. This is what actually moves the robot at deployment time.

A software engineer might analogize this to a compiler that takes high-level source code (the language instruction), optimizes it through multiple intermediate representations (the dual-stream MoT), and outputs machine code (continuous action chunks) for different target architectures (robot embodiments).

System 2: Understanding and Planning (The “Brain”)

Before System 1 can act, it needs to know what to do. System 2 is a PaliGemma2 (3B) vision-language model that interprets the current scene and decomposes tasks.

Given an instruction like “organize the desk,” System 2 performs chain-of-thought reasoning: it identifies relevant objects (“laptop, notebook, pen”), their spatial arrangement, and then generates a subtask instruction—for example, “clear the left side of the desk.” This subtask is passed to System 1 as an intermediate goal.

This hierarchical decomposition is crucial for long-horizon tasks. Instead of trying to map “organize the desk” directly to a sequence of motor commands (which might span 50+ action steps), the model first reduces the problem to a manageable subtask. System 2 also provides task context for System 3, allowing the prediction and evaluation system to operate on a more explicit representation of the current task stage.

System 3: Prediction and Evaluation (The “Crystal Ball”)

System 3 is an independent GigaWorld-1 based world model (5B parameters) that provides two complementary signals to System 1:

  1. Future observation (subgoal image): The world model generates a short video prediction of what the scene will look like after the current subtask is executed. The last frame of this video is extracted as a subgoal image—a visual representation of the anticipated physical state. This is provided to System 1 as an additional visual condition, effectively letting the policy “see” the future it should aim for.

  2. Value-based evaluation: The model estimates a scalar value VtV_t representing how far the current execution state has advanced toward subtask completion. This is converted into a binary advantage condition At∈{0,1}A_t \in \{0, 1\}, indicating whether the predicted task progress is increasing. At inference time, the policy is conditioned on positive progress, biasing action generation toward behaviors that advance the task state.

Why this matters: In complex manipulation, the robot often finds itself in visually similar states at different stages of a task (e.g., the gripper is near the cup at the start and at the end). System 3’s value signal helps the policy distinguish “I’m making progress” from “I’m spinning my wheels,” leading to more efficient and robust execution.

Key Results & Benchmarks: From Emergence to Excellence

The paper’s experimental results are extensive, but several trends stand out when translated into plain language.

Scaling with Data Size

The paper systematically varied the amount of pretraining data. The results, visualized in Figure 8a, show a clear positive trend: larger training sets consistently reach lower validation loss, and this separation remains visible throughout optimization. In plain terms: the more diverse embodied experience the model sees during training, the more reliable it becomes at real-world tasks.

When looking at real-robot success rates (Figure 8b), the pattern is even more compelling. Increasing robot-data scale improves performance, but the effect depends on task complexity. Fruit manipulation improves smoothly with more data, but clothes folding—a complex long-horizon skill—shows a much sharper dependence on data scale. This suggests that scaling does more than improve a training objective; it makes previously unreliable behaviors increasingly reliable. This is what the authors call “capability emergence with scale.”

The Value of Human Data (UMI and EGO)

The paper investigates whether adding human demonstration data helps. Starting from 3,000 hours of robot data, they added equal amounts of UMI (near-hand human demonstrations) or EGO (first-person human videos). Both improved performance over robot-only pretraining, and the combination of robot, UMI, and EGO data provided the strongest results across all settings. The improvement persisted after task-specific post-training, indicating that the additional human experience changes the quality of the learned pretraining prior, not just the base policy.

System 3 Ablations: Prediction and Evaluation

The paper includes detailed ablations of System 3’s two conditioning signals, evaluated on clothes folding, gift wrapping, and cube sorting.

  • On clothes folding: All configurations achieved 100% task success, but System 3 signals improved efficiency. Adding the subgoal image reduced average completion time from 107s to 83s. Adding the value signal reduced it to 79s. Using both together brought it down to 75s.
  • On gift wrapping: This is where System 3 truly shines. The base model failed to complete the task. Adding subgoal images got to 20% success; adding value signals got to 60%; using both together achieved 80% success. Completion time for successful episodes dropped from 65s to 55s.
  • On cube sorting: A mixed pattern emerged. The combined model (subgoal + value) achieved the highest average task score (increasing from 55.0% to 87.5%), while value conditioning alone yielded the shortest completion time (80s).

The takeaway: System 3 doesn’t just increase the chance of success; it makes successful episodes more efficient and robust, especially on challenging tasks.

Out-of-the-Box Generalization

Perhaps the most impressive results are from evaluating the pretrained policy without any task-specific adaptation.

  • Language following: On both AgileX PiPER and Maker H01, the same pretrained GigaBrain-0.7 policy could follow multiple language instructions out of the box. It executed diverse instruction-conditioned behaviors involving object selection, scene rearrangement, and multi-step manipulation.
  • Out-of-distribution generalization: The paper evaluated settings where objects, scenes, or task compositions were completely unseen during training. For example, on AgileX PiPER, the target “pepper” instance was absent from training data, and the two target containers were never collected together in the same task configuration. Despite this, the pretrained policy successfully identified the language-specified target, localized the relevant container in the current arrangement, and composed the appropriate pick-and-place behavior.
  • Deformable object generalization: The paper evaluated garment folding on unseen garment instances from diverse, unstructured initial states. The pretrained policy could reorganize the garment and sustain the folding behavior across both AgileX PiPER and Maker H01 without task-specific adaptation. This suggests that heterogeneous pretraining transfers reusable interaction structure across unseen objects, rather than merely memorizing specific appearances.

Post-Training Performance

After task-specific post-training, GigaBrain-0.7 consistently improved over preceding GigaBrain models and matched or exceeded π₀.₅ across reported tasks. On the Maker H01 platform, GigaBrain-0.7 improved the average success rate from 69.6% (GigaBrain-0.1) to 84.2%, with particularly large gains on button push and spoon grasping.

On the AgileX PiPER platform, GigaBrain-0.7 achieved an average success rate of 85.0% on the six language-following tasks, compared to 88.8% for π₀.₅ (with GigaBrain-0.1 at 76.1%). While π₀.₅ edged out GigaBrain-0.7 on this specific benchmark, the paper notes that GigaBrain-0.7’s gains are “particularly visible when language must be grounded into target-specific or directional physical behavior.”

In complex manipulation (Table 7), GigaBrain-0.7 showed consistent improvements. On AgileX PiPER, it achieved an average complex manipulation success rate of 84.9%, compared to 64.8% for GigaBrain-0.1 and 37.8% for G0.5. On Maker H01, the average was 74.1%, compared to 43.2% for GigaBrain-0.1 and a mere 8.3% for G0.5. Notably, GigaBrain-0.7 achieved 100% success on Food Prep. & Heating and Sweeping on Maker H01, while other models struggled below 50%.

Simulation Benchmarks

In simulation, GigaBrain-0.7 achieved state-of-the-art results:

  • RoboTwin 2.0: Ranked first under the challenging Hard (domain-randomized) setting, with a 67.3% overall success rate, outperforming π₀.₅ (58.3%) and significantly surpassing earlier GigaBrain models.
  • EBench: Achieved the strongest performance in both overall success rate (33.3%) and aggregate task score (46.1), outperforming π₀.₅ (28.0 / 37) and π₀.₅ (28.0 / 42).
  • RoboColiseum: Achieved the strongest performance across all four dimensions (instruction following, spatial reasoning, robustness, and general manipulation), with an 81.66 score in instruction following.

Why It Matters: Key Takeaways

GigaBrain-0.7 isn’t just a incremental improvement; it points toward a new paradigm for embodied AI. Here are the four most important takeaways:

  1. Three-system architectures are the way forward. The paper convincingly demonstrates that separating concerns—understanding (System 2), prediction/evaluation (System 3), and action (System 1)—allows each component to specialize without breaking the others. The “pipeline” approach of many prior VLA models (where understanding feeds action, which feeds into a world model trained later) is giving way to tightly coupled, jointly trained systems. Watch for more architectures that explicitly coordinate planning, prediction, and execution in a single learning loop.

  2. Heterogeneous data scales emergently. The finding that 37,000+ hours of data spanning 16 robot morphologies and multiple data sources (real robot, UMI, EGO, simulation, world-model-generated) yields emergent generalization is significant. It suggests that the diversity of embodied experience matters as much as the quantity. In practical terms, this means robot developers should prioritize collecting data across varied morphologies and viewpoints, not just more trajectories of the same robot doing the same tasks.

  3. System 3 provides a powerful “progress awareness” signal. The ablations showing that value-based evaluation and future-state prediction dramatically improve completion times and success rates on complex tasks (especially gift wrapping and clothes folding) indicate that simply telling a policy “are you making progress?” is a highly effective training signal. This has implications for offline and online RL: world models that predict future states and estimate progress could be the key to sample-efficient robot learning.

  4. Generalization beyond memorization. The out-of-distribution results are the paper’s most striking contribution. The model doesn’t just work on tasks it’s seen before; it generalizes to new objects, new scenes, and new garment configurations. This is the difference between a robot that can “do the tasks it was trained on” and a generalist foundation model that can follow novel instructions in novel environments. For product managers and robotics enthusiasts, this is the moment where VLA models stop being clever demos and start being deployable generalists.

What to watch for next: The paper explicitly outlines future directions—scaling to even larger and more diverse datasets, strengthening long-horizon predictive modeling, and developing more scalable closed-loop learning from autonomous experience and human correction. The release of training code and pretrained weights (mentioned throughout the paper) will likely accelerate follow-up work. If the pattern holds, we can expect to see subsequent models that build on GigaBrain-0.7’s three-system architecture to push the boundaries of what generalist robots can do in the real world.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →