arXiv:2608.24680nvidia/nemotron-3.5-lightning-30b-a3bAugust 25, 2026

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model TrainingExplained for Beginners

Wenxuan Shen, Dongna Jin, Dongping Chen

Computer Vision and Pattern Recognition

Abstract

Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in-the-wild gameplay videos. However, raw gameplay footage entangles the game world with screen-space interfaces, introducing game-specific biases and irrelevant dynamics that hinder world-model training. To address this problem, we introduce GameUI-Taxonomy and G2WEngine, a full-stack framework that formalizes gameplay UI grounding and removal. G2WEngine automatically extracts reusable UI assets from real gameplay videos and synthesizes temporally coherent UI overlays on clean footage. Using this engine, we construct Game2World, comprising 96K synthetic paired videos with precise reconstruction targets and 1,079 in-the-wild clips from 303 games for realistic evaluation. Its asset library contains 5,132 verified UI elements across 21 taxonomy categories, collected from 1,010 representative gameplay frames. Based on Game2World, we propose GameCleaner, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities. Unlike mask-based methods, GameCleaner directly identifies and removes diverse HUD elements while preserving the underlying scene content and temporal dynamics. In a controlled pilot, world models trained on UI-free gameplay improve overall VideoReward by 6.83% over those trained on UI-overlaid data. On UI-removal evaluation, GameCleaner achieves an average AAR of 95.36 on synthetic videos, outperforming the strongest temporal mask baseline by 57.3%, and obtains the best in-the-wild AAR of 80.05 with 99.8 background preservation. These results demonstrate the scalable potential of transforming Internet gameplay videos into high-quality world-model training data. Code, dataset, and model will be available at https://github.com/Dongping-Chen/Game2World.

Here is the structured explanation of the Game2World paper.

1. The Problem: The "UI Tax" on Gameplay Data

The core problem the paper addresses is the contamination of gameplay videos by Heads-Up Displays (HUDs) and user interfaces.

In the real world, anyone capturing gameplay footage gets a video that is a composite of two things: the actual 3D game world and the 2D screen-space interfaces that sit on top of it. These interfaces include health bars, maps, ammo counters, dialogue boxes, and streaming overlays.

While these are useful for human players, they are "noise" for AI training. Raw gameplay footage "entangles the game world with screen-space interfaces," introducing game-specific biases. A model trained on UI-overlaid data might learn to predict the static position of a health bar rather than learning the physics of the game world. This hinders the development of "world models"—AI systems that learn how environments evolve over time. Essentially, the paper argues that to treat internet gameplay as a scalable training source, we first need to strip away the interfaces to get to the "clean" world underneath.

2. How It Works: G2WEngine and GameCleaner

The authors introduce two key components to solve this: G2WEngine (the data pipeline) and GameCleaner (the AI model).

The Engine (G2WEngine): A Factory for Clean Data G2WEngine is a "full-stack framework" that automates the process of cleaning gameplay videos. It works in four stages:

  1. Taxonomy: It defines a unified vocabulary (GameUI-Taxonomy) to categorize UI elements (e.g., radar, inventory, chat).
  2. Asset Extraction: It scans real gameplay videos, identifies UI elements, and extracts them as reusable "assets" (think of cutting out a logo from a video frame).
  3. Clean Curation: It selects segments of gameplay that are naturally UI-free.
  4. Synthesis: It composites the extracted UI assets onto the clean footage. This creates "paired data"—videos with the UI on, and the same videos with the UI off. This paired data is the ground truth needed to train an AI model to remove UI.

The Model (GameCleaner): The Mask-Free Eraser GameCleaner is the model that does the actual removal. Unlike many video editing tools that require you to paint a mask around the thing you want to delete, GameCleaner is "mask-free." It uses a Multimodal Large Language Model (MLLM) to "see" and understand what a UI element is based on its appearance and context.

It identifies HUD elements directly from the video frames and removes them. Crucially, it doesn't just delete the pixels; it has to "hallucinate" or reconstruct what was behind the UI (the game world) and ensure the video remains temporally consistent (things don't jitter or shift incorrectly as the player moves).

3. Key Results & Benchmarks

The results demonstrate that this pipeline works remarkably well.

  • The Pilot Study: In a controlled test, researchers trained a video-generation model on clean vs. UI-overlaid data. Models trained on UI-free data improved "VideoReward" (a quality score) by 6.83%. This included significant gains in motion quality (+18.8%) and aesthetics (+8.59%). This proves that removing UI makes the data significantly better for training.
  • GameCleaner’s Performance:
    • On Synthetic Data (Game2World-S): GameCleaner achieved an Average Artificial Removal (AAR) score of 95.36. To put this in context, the strongest mask-free competitor scored roughly 38 points lower. Furthermore, GameCleaner preserved the background (BG score) at 99.00, meaning almost none of the game world was accidentally destroyed.
    • On Real-World Data (Game2World-W): This is the harder test. GameCleaner achieved the best in-the-wild AAR of 80.05, outperforming the strongest baseline (Aurora) by nearly 29 percentage points. It also maintained a background preservation score of 99.80, meaning the game world stayed intact despite the complex, varying interfaces found in real games.

4. Why It Matters

The implications of this work are significant for the future of AI and game development:

  • Unlocking Internet Data: The internet is full of gameplay videos. This framework transforms them from "noisy" footage into high-quality training data. We are no longer limited to expensive, custom-built simulator data; we can use the millions of hours of gameplay already on YouTube and Twitch.
  • Better World Models: By cleaning the data, AI models can focus on learning the actual dynamics of the game world (physics, object permanence, agent navigation) rather than learning where the health bar sits on the screen.
  • Generalization: Because GameCleaner is mask-free and taxonomy-based, it can handle UI from games it has never seen before. It doesn't need a custom mask for each new game; it applies its understanding of UI categories universally.
  • What to Watch For: The paper notes limitations. Very large or opaque interfaces that blend into the game world might still pose challenges. Additionally, while the data is clean, getting the AI to understand actions (what the player is doing) within these cleaned videos remains a future challenge, as the current work focuses on the visual content.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →