arXiv:2608.21833nvidia/nemotron-3.5-lightning-30b-a3bAugust 22, 2026

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?Explained for Beginners

Kun Chen, Haorong Hong, Peizhong Gao +7 more

Artificial IntelligenceComputation and Language

Abstract

Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development trajectories identifies three stages that together span the lifecycle of game development with a coding agent: initial game generation, bug diagnosis and repair, and optimization over multiple turns. Therefore, we introduce GameXpert-Bench, which operationalizes the three lifecycle stages as three complementary benchmark tracks. GameGen evaluates complete game creation from a single request in an empty workspace. GameFix evaluates diagnosis and repair when defects are reported or left for the agent to discover. GameOpt evaluates cumulative optimization through request chains seeded by real development trajectories between users and agents. We evaluate each track using live game interaction, deterministic behavioral tests, or final product criteria with regression checks. The suite contains 97 generation tasks across 11 genres; 100 repair tasks from 50 game levels verified by humans, each with 19-27 injected bugs; and 17 optimization chains with six turns and 102 requests. Across the three tracks, current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.

Here is the technical explanation, written in a conversational yet precise tone, based on the provided abstract and full text of the GameXpert-Bench paper.

1. The Problem

To a software engineer or product manager, the allure of "build me a game" is obvious, but the engineering reality is brutal. Game development is a high-stakes coordination problem. You aren't just writing a function; you are synthesizing program logic, visual assets, audio, user interfaces, and physics—all of which must coexist in a single executable artifact.

Existing benchmarks for Large Language Models (LLMs) typically suffer from a narrow scope. Some evaluate the final product alone, checking if a game compiles or runs. Others isolate a single stage, like testing code generation in a vacuum. But real-world development with an AI agent is rarely a one-shot deal. It is a lifecycle.

The authors of GameXpert-Bench identified a critical gap: current evaluation methods fail to capture the full trajectory of building a game with an AI. They observed that human-agent interactions naturally fall into three distinct phases: first, generating the game; second, diagnosing and fixing bugs; and third, iteratively optimizing the experience based on feedback. Most benchmarks only measure the first phase, or measure it poorly. The problem the paper addresses is the need for a unified yardstick that measures an agent's ability to not just create a game, but to heal it and improve it over time.

2. How It Works (The Technical Mechanics)

GameXpert-Bench operationalizes the game development lifecycle into three complementary "tracks," each acting as a specific test for a specific phase of the AI’s capability.

The Three Tracks

  1. GameGen (Generation): This is the "blank slate" test. The agent is dropped into an empty workspace with only a natural-language brief. It must generate a complete, browser-native game from scratch, choosing its own engine and assets. It contains 97 tasks across 11 genres (including 44 3D games). The evaluation here is rigorous: it doesn't just check if the code compiles; it checks if the game is "playable." The authors built a Shared Rubric by analyzing the events (like "player jumps" or "enemy shoots") produced by various models. Human annotators then curated a unified checklist. Every model’s output is scored against this same checklist, ensuring a fair comparison of functional completeness and richness.
  2. GameFix (Repair): If GameGen is the creation phase, GameFix is the hospital phase. The benchmark takes 50 verified "Gold Games" and injects 19–27 bugs into each level using reversible mutations. This creates 100 repair tasks. The critical twist is that agents are evaluated in two modes: "Explicit Issue," where the bug is reported to the agent, and "Self-Discovery," where the agent must find the bug based on limited symptoms. A repair is only considered successful if it fixes the broken behavior and preserves the behavior that was working before (regression prevention).
  3. GameOpt (Optimization): This tracks the iterative refinement loop. It contains 17 optimization chains, each with six turns and 102 total requests. These chains are seeded from real human-agent development histories. The agent starts with a playable snapshot and receives successive requests spanning gameplay, level design, balance, art, interface, and audio. The evaluation checks if the agent can incorporate these requests while preserving the core game loop and maintaining balance across these different "product dimensions."

An Analogy for Engineers

Think of the agent like a junior developer starting a new project in a foreign codebase.

  • GameGen is like assigning that junior developer a "Build a platformer from scratch" ticket. They have never seen the project before. The benchmark measures if they can set up the repository, write the core player movement code, and get a playable window on screen.
  • GameFix is like finding out that the platformer’s jumping mechanic is broken after the first playtest. In the "Explicit Issue" mode, a QA tester hands the developer a bug report. In "Self-Discovery" mode, the developer has to boot up the game, notice the jump feels "floaty," and track down the bug in the code. The benchmark measures if the developer fixes the jump without breaking the walking mechanic or the enemy AI.
  • GameOpt is the feedback loop. After the first playtest, the product manager says, "The controls feel great, but the level is too easy," and later, "The visuals are bland, make them neon." The benchmark measures if the developer can implement these successive changes while keeping the game stable and the core loop intact.

3. Key Results & Benchmarks

The results paint a picture of a capability gap that will be familiar to anyone who has coded with an LLM.

  • GameGen (Generation): The results are decent but reveal a strong bias toward the "skeleton" over the "flesh." The top model, Claude-Opus-5, scored a 79.7 out of 100. It excelled at Completeness (94.4), meaning it reliably built the core mechanics requested. However, it scored much lower on Richness (72.0). This confirms the paper's thesis: agents are great at implementing explicit requirements but struggle to add "bonus" content or diverse mechanics beyond the prompt.
  • The 3D vs. 2D Divide: There is a significant drop in performance for 3D games. Across 15 models, the average score dropped from 65.9 on 2D games to 60.1 on 3D games—a gap of 5.8 points. The drop was most severe in Completeness (-8.8 points), suggesting that 3D reasoning, spatial logic, and engine integration are currently much harder for these agents.
  • GameFix (Repair): This is where the agents struggle most. The paper notes that current agents are "more reliable at producing playable foundations... than at discovering defects." In the self-discovery mode, where the agent isn't told what is broken, performance plummeted. Furthermore, when games contained multiple bugs, "near-complete repair remained uncommon." This highlights a major weakness: agents can follow a spec, but they struggle with autonomous debugging and regression safety.
  • GameOpt (Optimization): Here, leading agents often retained requested functionality across the six turns. However, preserving the "core game loop" and achieving "balanced improvement" across art, audio, and gameplay was inconsistent. The agents could follow the instructions, but they often broke things they weren't explicitly asked to change.

Translation into plain language: If these agents were students, they’d ace the "build the basic prototype" exam but fail the "find and fix the hidden bugs" and "polish based on feedback" exams. A model might generate a working racing game from a prompt, but if you ask it to "make the cars handle better" and then "add a soundtrack," there’s a significant chance it will break the steering physics or delete the car models in the process.

4. Why It Matters (Key Takeaways)

  • Initial generation is not enough. You cannot judge an agent's game development skill solely by how good the first build is. An agent might generate a gorgeous, functional game that falls apart the moment you ask for a single tweak. The benchmark proves that the "create" and "maintain" skills are distinct.
  • The "Self-Discovery" gap is the biggest hurdle. The most striking finding is how poorly agents perform when they have to find the bugs themselves without a prompt. For product managers, this means: don't expect your AI game dev to act as a QA tester out of the box. You will likely need to provide explicit bug reports or rely on automated test suites to catch regressions.
  • 3D is the current frontier of difficulty. The significant performance drop in 3D generation indicates that the industry should not expect these agents to replace human 3D developers anytime soon. The complexity of 3D space, physics, and asset management currently outpaces the capabilities of general-purpose coding agents.
  • Watch for regression safety. As the optimization track shows, the real danger in using these agents for live development is "bit rot"—where a requested feature improves one part of the game but silently breaks another. The benchmark's focus on "Pass-to-Pass" checks (ensuring old behavior is preserved) sets the standard for how we should evaluate future agents: not just on what they add, but on what they don't break.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →