A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?
Abstract
Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at https://a2z-gamespec-bench.github.io.
Community
A2Z GameSpec-Bench evaluates how faithfully coding agents build complete games from 100 long-form Game Design Documents. Games may compile and execute successfully but fail to adhere to the intended design, especially when multiple requirements interact.
We introduce Dependency-Aware Contracts that capture requirements and their prerequisite relationships, and agent-generated Test Policies that exercise them through scenario-based replay and adaptive playtesting. Source-code inspection, replay, and playtesting collectively assess the same requirements across implementation, rendered behavior, and player interactions.
Our evaluations show that current agents struggle to satisfy interdependent requirements across code and actual play. The fixed contracts enable consistent comparison across agents, while requirement-level evidence pinpoints design violations that source-code checks alone may overlook.
- Project page: https://a2z-gamespec-bench.github.io/
- GitHub: https://github.com/krafton-ai/a2z-gamespec-bench
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper