GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic ManipulationExplained for Beginners
AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42 more
Abstract
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.
The Problem
Robots that can manipulate objects in unstructured environments—like a cluttered kitchen or a dynamic warehouse—have traditionally been trained for very specific tasks. If you wanted a robot to pick up a mug, it needed to be programmed or meticulously trained for that exact mug in that exact lighting. This "narrow" approach is slow, expensive, and fails when the environment changes.
The dream is a general-purpose robotic "brain" that understands the world like we do: by watching and acting. Recent research has looked at "World-Action Models" (WAMs). The idea is simple but powerful: if a robot can predict what the world will look like after it moves, it can plan its actions backward from that prediction, much like a human mentally rehearses a move before making it. Most current WAMs, however, are just adapted from video-generation models pretrained on massive amounts of general internet video. They haven't been properly "scaled up" or trained from scratch specifically for the messy, complex data of robot manipulation. The authors of GE-Act 2.0 argue that to truly get general robot intelligence, we need to build these models from the ground up, specifically for how robots see and interact.
The Problem in Plain English
Imagine you want to teach a robot to clean a table. Old approaches would show the robot thousands of videos of people cleaning tables, then fine-tune it for your specific table. GE-Act 2.0 takes a different approach: it’s like teaching the robot to "dream" about the future first. By learning to predict what the table will look like after a wipe or a pick, the robot builds an internal understanding of physics and object permanence. Once it knows how the world works, it can figure out how to act to make that predicted future happen. The key advance here is that they built the model from scratch on robot data, rather than borrowing a model designed for human YouTube videos.
How It Works (The Technical Mechanics)
At the heart of GE-Act 2.0 is a clever architecture with three main components, which we can think of as the robot's "imagination," its "GPS," and its "muscle memory."
- The Control-Oriented Autoencoder (CoAE): Think of this as a aggressive data compressor. Raw video from a robot's camera is huge. The CoAE compresses this video into a tiny, compact "latent" representation—essentially a stripped-down version of the scene. Crucially, it is designed to keep only the information relevant to the action and the user's instructions, throwing away the rest. It’s like taking a photo of a messy room and compressing it down to just the list of objects on the floor, discarding the wallpaper pattern.
- The Single-Step Visual Planner (SVP): This is the robot's "mental time machine." Normally, planning requires stepping forward in time one tiny step at a time, which is slow. The SVP does it all at once. It takes the current compressed scene and predicts the complete future state in a single differentiable pass. For a software engineer, think of this like a function that takes "now" and outputs "what the world looks like after I do X," rather than a loop that updates "now" 60 times a second.
- The Inverse Dynamics Model (IDM): This answers the question: "What move did I just make?" Given a starting state and an ending state, the IDM predicts the action that likely caused the transition. It’s the rewind button.
These three components are trained separately at first—like training a student on theory, then on practice—but they are brought together later via a process called Knowledge-Aligned Selective Optimization (KASO). This is the secret sauce. During training, the model predicts what the future will look like, and KASO acts as a strict gatekeeper: it only keeps the training examples where the predicted future is "behaviorally compatible" with the action actually taken by the robot. If the model daydreams about something wild that isn't what happened, KASO filters it out. This ensures the model learns from reality, not fantasies.
Key Results & Benchmarks
The results are striking because they show that scale matters, even when starting from scratch. The researchers tested the model on 100 tasks across 20 manipulation skill groups, using scenes and objects the model had never seen before (zero-shot, out-of-distribution testing).
- The Power of Scale: They started with 300 hours of training data and scaled up to 30,000 hours. The impact was dramatic. On the G1-OP task group, success rates jumped from 17.1% to 44.1%. On G2-90D, they went from 13.4% to 31.1%. To put that in perspective, that’s roughly a 2.5x and 2.3x improvement respectively just by showing the model more examples.
- The "Hidden Gem" Dataset: Perhaps the most interesting finding involves G2-90D. This dataset comprises less than 2% of the total training data, yet it drove the success rate up by 17.7 points. The authors suggest this is evidence of "cross-embodiment transfer"—learning generalizable skills from diverse data that apply even when the specific robot embodiment changes.
- Broad Success: The model performed well across the board. Gains were observed in 19 out of 20 skill groups on one metric, and 18 out of 20 on another. This suggests the model isn't just overfitting to specific quirks but learning a robust understanding of manipulation.
- Correlation with Coverage: There was a very strong correlation (Pearson r=0.80, Spearman rho=0.85) between how much "skill-specific coverage" the model had during training and its success on out-of-distribution tasks. Essentially, if the model saw a diverse set of skills during training, it was much better at handling new, unseen scenarios.
Why It Matters
- Generalization Without Fine-Tuning: The most impressive takeaway is that these results were achieved without any "per-task fine-tuning." The model was tested directly on 100 tasks it had never seen. This moves us closer to the vision of a robot that can walk into a new kitchen and immediately start helping, rather than needing a software update for every new environment.
- Instruction Following: The model demonstrates a remarkable ability to follow explicit instructions, even when those instructions conflict with its current behavior or common sense. If told to move a specific red block, it will do so, ignoring other blocks or "conventional" ways the scene "usually" looks. This is a critical step toward natural human-robot interaction.
- Efficient Data Usage: The finding that a tiny subset of data (G2-90D) yields significant improvements hints at a future where we don't need 30,000 hours of video to get a capable robot. We might need high-quality, diverse "seed" data, and the model does the rest.
Limitations and Caveats
While the results are promising, it is wise to view them through a realistic lens. The model is still operating within a simulation or controlled real-world setup; transferring these "zero-shot" capabilities to a physical robot in a bustling, unpredictable home environment involves engineering challenges (sensor noise, hardware latency) that the paper doesn't fully address. Additionally, while the model follows instructions, it still operates within the constraints of its physical embodiment and the physics engine it uses to predict the future. If the "world model" it built is slightly wrong, the robot might execute an action that fails in reality.
What to Watch For Next
The most exciting direction arising from this work is the potential for sim-to-real transfer. If a model can learn to "dream" and plan entirely in the latent space of compressed video, we might see robots that spend less time physically bumbling around and more time "thinking" in simulation before acting. Watch for follow-up research that pushes these scaling laws further—can we get to 100,000 hours? And can we integrate language instructions more deeply so that "Pick up the thingamajig" becomes as natural as asking a friend?
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →