Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at https://a2z-gamespec-bench.github.io.
Figures & tables
Figure 1 : Overview of A2Z GameSpec-Bench. Dependency-Aware Contracts guide evaluation and targeted revision through source-code inspection, replay, and playtesting. Our evaluation results demonstrate that High compilability and runnability do not guarantee high GDD Fidelity.
Figure 2 : Challenges in specification-driven game development. 2(a) In our game Siege Deck 2D , the deck-limit requirement gets full source-code credit, but playtesting reveals a reward gain without the card removal. 2(b) Across three game-design corpora with 353 documents, CiRA ( Fischbach et al., 2021 ) classifies 38.4% of sentences as causal on average, compared with 28% for general documents ( Frattini et al., 2023 ) (Appendix B.1 ). 2(c) Source-code pass rates decrease when all upstream rules must also pass (Appendix B.7 ).
Benchmark Setup
Requirement Evaluation
Evaluation Channels
Benchmark
Full-Game Generation
Input Type
# of Spec. Tokens
Evaluation Target
Relations
Predefined
Cross-Axis
Source Code
Rendered Behavior
Adaptive Play
GameDevBench
✗
Task + project
185
Task-specific tests
✗
✓
✗
✓
✗
✗
GameEngineBench
✗
Spec. + project
–
Tests + LLM judge
✗
✓
✗
✓
✗
✗
OpenGame-Bench
✓
Game spec.
1,830
Build / visual / intent
✗
✓
✗
✗
✓
✗
WebGameBench
✓
Structured spec.
–
Runtime quality
✗
✓
✗
✗
✓
✓
GameCraft-Bench
✓
Game spec.
1,547
Predefined rubric
✗
✓
✗
✗
✓
✗
Table 1 : Comparison of agentic game development benchmarks. A2Z GameSpec-Bench fixes a dependency-aware evaluation contract from long-form GDDs independently of generated outputs and links judgments from complementary evidence channels to the same requirements.
Figure 3 : Overall evaluation and revision pipeline of A2Z GameSpec-Bench.
Model
Overall GDD Fidelity
Evaluation Axis
Verifiable Rate
Source
Replay
Playtest
All
Small
Big
Small
Big
Small
Big
Small
Big
Small
Big
Claude-Fable-5.1
77.0
82.8
71.1
89.7
83.5
73.5
55.1
85.2
74.7
98.7%
98.7%
Claude-Opus-5
73.9
80.9
66.8
87.9
78.4
71.2
54.4
83.6
67.7
100.0%
99.3 %
Claude-Opus-4.8
56.8
67.9
45.6
72.5
49.6
62.5
41.4
68.7
46.0
99.3%
100.0%
GPT-6-Astra
71.6
78.5
64.6
84.2
74.4
69.9
52.1
81.6
67.4
99.3%
97.3%
Table 2 : Main results on A2Z GameSpec-Bench. Overall GDD Fidelity, per-axis scores, and verifiable rates on 100 GDDs. Bold represents the best result, and underline indicates the second-best.
Figure 4 : Complementarity of source-code evaluation, scenario-based replay, and adaptive playtesting. 4(a) Examples of specification violations identified through different evidence traces. 4(b) Playtest judgment outcomes grouped by source-code score. 4(c) Ratio of reversed GDD fidelity rank orderings across all 100 evaluated games built by GPT-5.6-Sol upon omission of individual evaluation axes.
Figure 7
Figure 7 : Requirement-level repairs beyond self-revision. Each panel pairs a GDD requirement with the initial build, the self-revision baseline, and Source + Replay + Playtest ; both revision columns show second-round builds. The examples compare 7 body recoloring, 7 estimate-class icons and confidence, 7 a held fragment following the pointer, and 7 the Hull update required by oxygen depletion. In each illustrated case, the violation remains in the initial and self-revision builds and is repaired by Source + Replay + Playtest . Detail crops enlarge the highlighted regions in 7 – 7 ; crosshairs in 7 mark the actual pointer position. In 7 , all three builds display a drowned-run screen, but only the feedback revision sets Hull to 0. The state cards report values recorded during execution. In 7 , R is the reveal radius and d is Manhattan distance from the ship.
Table 4 : Component analysis. 4(a) Mean and maximum absolute changes in source-code score from uniform weighting (Appendix F.1.2 ). 4(b) Evidence support for fixed-rate versus adaptive frame selection (Appendix G.3 ). 4(c) Judgment coverage (%) after a second normal or adversarial pass under matched budgets. The first normal pass reaches 40.5%; gains are percentage points relative to this shared first pass (Appendix F.3 ). Coverage includes both satisfied and violated requirements.
Figure 8 : Extension of contract-based evaluation to 3D games. GDD requirement summaries accompany original gameplay captures from two Three.js builds: 8 vehicle-based ball play in Rocket League and 8 momentum-based platforming in Sonic: Cascade Coast . Both builds are assessed through source-code inspection, scenario-based replay, and adaptive playtesting against their fixed contracts. Appendix H reports the scores and requirement-level diagnoses.
Appendix figures & tables46 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1 : One game traced from brief to GDD. The brief fixes what to build, the CV adds how the game should feel and why (without numbers), and the GDD specifies every value and rule. The highlighted boxes follow one requirement, the cost of letting a unicorn flee, through all three stages: a rough rule in the brief, a design intent in the CV, and a cited rule, constants, and a worked example in the GDD.
Setting
Value
CVs used in the search ( N )
13
Candidate harnesses per iteration ( K )
3
Repeats per quality judgment / pairwise comparison
3 / 3
Proposer and generator
Claude-Opus-4.8
Evaluator (scoring and pairwise judging)
GPT-5.5
Maximum iterations
20
Appendix
Table S1: Harness search configuration.
Criterion
Requirement
Failure if violated
Example violation
G1 Structural Conformance
• Standalone, all sections present • No empty or [TBD] section • No dangling identifier
The requirement cannot be checked.
A section cited elsewhere had been removed.
G2 Referential Consistency
• One canonical definition ( DBT- n ) • Every citation resolves • Repeated values agree
Contradictions; no implementation can satisfy all.
bool[56] declared for 55 species; 1 wave in App. A vs. 3 in App. B.
G3 Behavioral Soundness
• Total, deterministic rule tables • No unreachable or dead mechanic • Difficulty never decreases
Faithful implementations are penalized.
A 90 s boss timer never fires; HP decay kills the patient at 40 s.
G4 Executable Formulas
• Variables bound to constants • Fixed evaluation order, bounds • Worked example that recomputes
Agent and evaluator can compute different answers.
Kill time ignores the boss’s DEF; ceil(50/4) written as 12.
Appendix
Table S2: Criteria for golden GDDs. Each criterion excludes one way a GDD can invalidate conformance-based evaluation. Example violations are taken from our validation logs.
Statistic
Small
Big
Documents
50
50
Lines (mean)
629
1,153
Tokens (total)
704,270
1,314,862
Tokens (mean, range)
14,085 (5,240–24,022)
26,297 (12,489–51,433)
Outcome requirements (total)
2,710
4,202
Outcome requirements (mean, range)
54.2 (9–151)
84.0 (10–246)
Appendix
Table S3 : GDD dataset statistics. Document length and requirement composition of the Small and Big splits. Outcome requirements count individual promises in GDD rule-table rows, including presentation outcomes.
Table S5 : Genre distribution of the Big GDD split.
Figure S2 : A Small golden GDD example. Excerpt of traces_left ; omitted sections are marked.
Figure S3 : A Big golden GDD example. Excerpt of chameleon_slide GDD; omitted sections are marked.
Document
Sentences
Causal (%)
The Sky Above, the Sky Below
2,069
38.3
Frontier Pharmacist
1,996
42.5
Leisure Suit Larry 5
1,701
38.3
Love for Sail
1,551
37.5
Shape Up or Slip Out
1,444
44.6
Claw
1,428
29.6
Appendix
Table S6 : Studio game-design documents used in the causality analysis. Sentence counts are computed after text normalization and non-prose filtering; causal rates are CiRA predictions over the retained sentences.
Corpus
Documents
Sentences
Causal (%)
Studio GDDs
10
11,741
37.9
Game benchmark specifications
246
11,441
32.0
Board game rulebooks
97
50,518
45.3
Pooled
353
73,700
42.0
Macro average over corpora
–
–
38.4
Appendix
Table S7: Causal-sentence estimates in external game-design corpora. Rates are CiRA predictions on retained prose. The pooled rate weights corpora by sentence count, while the macro average gives equal weight to the three corpora.
Figure S4 : Contract construction from a GDD. The pipeline fixes vocabulary and rule entries before completing rule definitions and validating the contract. Coral boxes denote model-based stages, while dark-red boxes denote deterministic stages without model calls. The ×3 labels indicate independent repetitions: row facts are consolidated by two-out-of-three voting, whereas generated contracts are compared for structural consistency before one is frozen for evaluation. The dashed line indicates that the vocabulary Vi remains fixed throughout subsequent stages.
Figure S5 : Contract-linked execution of traces_left . Nodes show rule occurrences in a recorded normal playtest, and labeled links indicate the state attributes connecting them. Three sensor placements and one removal precede P4 , which commits the run. The fixed contract supplies the rule identities and dependencies; the trace supplies their observed order. Source GDD rows appear in Figure S6 .
Figure S6 : The GDD rows the entries in Figure S5 were derived from, in the four sections of traces_left that the run touched, each labelled with its line in the framed GDD. A row states a situation and its required effects ( → ); the pipeline formalizes the situation as the rule’s trigger and preconditions, and the required effects as state updates or event emissions.
Agreement Measure
Contract
Matched Items
100%
Dependency Graph
≥0.999
Trigger Name
0.94
Predicate Structure
0.70
Appendix
Table S8: Consistency of evaluation targets across repeated generations. Three rule generations per GDD on 50 Small and 50 Big GDDs with fixed vocabulary and rule entries, averaged over the two splits; items are matched by their source entries.
Game & Rule
GDD statement (abridged)
Generation A
Generation B
Ratio of fi,rsrc ( A/B )
untangled_cords R8 ( Small )
All other frames: keep state, keep updating t in Play_Active
timer INCREMENT ⟨ dt ⟩
timer SET ⟨ advanced by dt ⟩
1.00 / 1.00
magma_arc STAR2 ( Big )
Not STAR3, and hit events ≤3 this attempt: 2 stars
event <= 3 STAR SET 2
event IN [1, 2, 3] STAR SET 2
1.00 / 1.00
Appendix
Table S9 : Examples of predicate-structure variation across contract generations. Rules whose predicate structure differs between two contract generations, with the GDD row they encode and the source-code score fi,rsrc each generation received on the same build.
Split
games
rules
Predicate structure
pairs
∣Δfi,rsrc∣≥0.25
Small
50
869
identical
1,773
12.7%
different
834
17.7%
Big
50
1,860
identical
4,045
8.8%
different
1,535
11.1%
Appendix
Table S10: Effect of predicate-structure variation on the source-code score. Rule pairs across the three generations, three per rule, judged on the same build; the last column is the share of pairs whose per-rule scores differ by at least 0.25.
Figure S7 : Excerpt of the source-code judge system prompt. Text is verbatim, and omitted passages are marked [...] .
Figure S8 : Excerpts of the two replay-judge system prompts. Frame selection over the recorded replay and scoring from the selected frames. Text is verbatim, and omitted passages are marked [...] .
Figure S9 : Excerpts of the three adaptive-playtest system prompts. Text is verbatim, and omitted passages are marked [...] .
Figure S10 : The playtest runtime interface and two recorded bots for magnet_dock . (a) Every bot receives the same ctx . (b) The normal playtest bot reads state, issues inputs, and records labelled snapshots naming the target rule. (c) The adversarial playtest bot enters through a declared scenario, pins the precondition of rule D1 with setState , and observes the outcome.
Figure S11 : Pairwise model comparisons and Elo ratings: Pooled GDDs. Matrix entries give the row model's win rate against each column model, with ties counted as one half. The rightmost column reports the model's Elo rating for that panel. Values above 50% favor the row model; Elo ratings are relative and centered at 1,000 within each panel.
Figure S12 : Pairwise model comparisons and Elo ratings on Small GDDs. The comparison protocol, relative model order, and win-rate color scale follow Figure S11 . Each panel includes an Elo column alongside the empirical pairwise matrix.
Figure S13 : Pairwise model comparisons and Elo ratings on Big GDDs. The comparison protocol, relative model order, and win-rate color scale follow Figure S11 . Each panel includes an Elo column alongside the empirical pairwise matrix.
Figure S14 : Requirement consistency across repeated source-code judgments. The naive judge evaluates each of 100 unchanged builds five times. Top: the minimum, maximum, and mean number of extracted requirements per game, with rule counts shown for comparison. Bottom: the distribution of the fraction of one run’s requirements matched in another run at a Jaccard threshold of 0.5.
Figure S15 : Dependency weighting changes game-level comparisons. Each cell represents one GDD. Purple cells indicate a strict reversal in the source-code score ordering of at least one pair of the six proprietary agents between uniform weighting ( α=1 ) and downstream-reach weighting ( α=0 ), with rule judgments and invariant scores fixed. Such reversals occur on 23 of the 100 GDDs.
Figure S16 : Effect of dependency weighting on rule scores. Each bar is Δα=Firule(α)−Firule(1) at α=0.5 . Left: root-defective games meeting the joint source-code and playtest criteria. Right: comparison games without an observed playtest violation among their assessed root rules. Negative values indicate lower weighted scores; the sign is unchanged at α=0 .
Judge, effort
Item r
Mean reward
Δ vs. ref.
s / call
$ / call
sol, xhigh (reference)
–
0.639
–
78
0.048
sol, high
0.88
0.667
+0.027
71
0.034
sol, medium
0.88
0.669
+0.030
63
0.034
sol, low
0.89
0.680
+0.041
52
0.045
luna, xhigh
0.85
0.634
−0.005
63
0.0074
luna, high
0.85
0.645
+0.006
41
0.0072
Appendix
Table S11 : Agent-as-a-judge on six builds with identical scenario replays and rubric. Item r is the Pearson correlation of rubric-item scores with the original sol extra-high judgments of the same builds (re-run agreement ≈0.86 ); Δ is the mean reward difference from them. Cost is per judge call at list prices. sol = GPT-5.6-Sol, luna = GPT-5.6-Luna.
Coding agent
Normal (%)
Adversarial gain (pp)
Combined (%)
Sat.
Viol.
Cov.
Sat.
Viol.
Score
Cov.
All
Claude-Fable-5.1
28.6
0.4
29.0
51.4
7.6
80.0
88.0
Claude-Opus-5
26.7
0.4
27.1
49.0
9.0
75.7
85.1
Claude-Opus-4.8
19.2
0.6
19.7
38.2
13.6
57.4
71.5
GPT-6-Astra
25.8
0.5
26.3
48.7
8.5
74.5
83.4
Appendix
Table S12 : Normal and adversarial playtest decomposition across agents. Each split uses the same score-reporting cohort as Table 2 ; All is the mean of Small and Big . Sat., Viol., and Cov. denote satisfied, violated, and conclusively judged requirements, respectively. Adversarial gain reports additional judgments for previously unverified requirements in percentage points (pp). Combined Score reproduces the Small and Big Playtest scores in Table 2 using the same archived per-game scores; phase contributions and coverage use exact verdict counts. All values are reported to one decimal place.
First pass
Second pass
Coverage
Δ
Normal
–
40.5
–
Normal
Normal
41.0
+0.5
Normal
Adversarial
48.5
+8.0
Adversarial
–
45.0
–
Appendix
Table S13 : Controlled playtest comparison. Mean judgment coverage (%) on 86 target rules from 25 Small and 25 Big games. A second pass keeps the first pass’s conclusive verdicts and re-tests only the targets it left unverified.
Agent
Equal
Source emphasis
Replay emphasis
Adaptive emphasis
(31,31,31)
(21,41,41)
(41,21,41)
(41,41,21)
Claude-Fable-5.1
77.0 (1)
79.4 (1)
73.8 (1)
77.7 (1)
Claude-Opus-5
73.9 (2)
76.2 (2)
71.1 (2)
74.3 (2)
GPT-6-Astra
71.6 (3)
73.5 (3)
68.9 (3)
72.3 (3)
GPT-5.5
63.2 (4)
64.5 (4)
60.3 (5)
64.8 (4)
GPT-5.6-Sol
61.8 (5)
62.0 (5)
60.4 (4)
63.1 (5)
Appendix
Table S14: Sensitivity to three-axis aggregation weights. Each cell reports the aggregate score (0–100) and rank among nine agents. The weight order is source, replay, and adaptive playtest. Source dependency weighting remains fixed at α=0.5 . All-game axis scores equally average the Small and Big means. Bold cells indicate a rank change from equal weighting.
Figure S18 : Full source-code credit with a runtime violation. Each rule passes the source-code judge (item score 1.00) and is judged violated in the playtest. (a) Rainwright : with 8% water, below the 25% minimum cost, the held build preview stays teal instead of turning red, because the denial flag is set only on release. (b) Strata Keepers : a two-day zone entry consumes four days because the cost is applied twice. (c) Chroma Bastion : the finished run scores 4,869 but a constant 2,760 is staged for submission. Screenshots are from asset-integrated builds whose relevant logic is unchanged from the evaluated source; boxes, callouts, and crops are annotations.
Figure 42
Measure
Small
Big
All
Games
50
50
100
Full-credit rules
738
896
1,634
Eligible full-credit targets
377
531
908
Additional inspection targets
97
311
408
Games with additional targets
17
36
53
Additional targets / eligible targets
25.7%
58.6%
44.9%
Appendix
Table S16: Dependency context identifies additional source-code inspection targets. A target has full enumerated source credit and a lower-scored direct predecessor. Percentages are pooled over eligible targets; the enumerated judgments remain fixed.
Judge
Effort
Rule-Level Correlation r
Invariant Agreement
Mean Rule Score
Tokens/Call (103)
Relative Cost
GPT-5.6-Sol
Extra high
0.74
0.83
0.76
213
1.00
High
0.70
0.83
0.76
214
1.00
Medium
0.75
0.88
0.78
176
0.83
Low
0.80
0.86
0.79
160
0.75
GPT-5.6-Luna
Extra high
0.72
0.77
0.79
198
0.15
High
0.72
0.76
0.79
187
0.15
Appendix
Table S17 : Agreement and cost across source-code judge configurations. Rule-level correlation and invariant agreement are measured against a fixed set of reference judgments from GPT-5.6-Sol at extra-high reasoning effort. The first row reports an independent re-run of that configuration. Mean Rule Score reports the average rule implementation score in each run. Relative cost is estimated from token usage per call and the prices used in the experiment, normalized to the first row.
Figure S20 : Tokens vs. frame width . Left: tokens per frame, fixed prompt excluded. Right: item-level agreement with the next-larger width.
Native
640 px
480 px
1920 × 1080
1280 × 720
Appendix
Figure S21 : A sampled frame at different resolutions. The same frame region as the judge receives it at native resolution, 640 px and 480 px (Lanczos downscaling as in the pipeline). Top: 1920 × 1080 game; bottom: 1280 × 720 game.
Judge, effort
Small
Big
Tokens / vote
Sol, xhigh
0.959 / 2
0.930 / 4
485–753K
Sol, high (re-run)
0.959 / 3
0.959 / 1
516–688K
Sol, low
0.892 / 7
0.897 / 3
217–315K
Luna, xhigh
0.919 / 3
0.893 / 1
57–86K
Luna, high
0.910 / 6
0.888 / 3
51–71K
Luna, low
0.833 / 9
0.810 / 10
46–64K
Appendix
Table S18 : Playtest judge comparison on fixed traces. Each cell reports verdict agreement with a GPT-5.6-Sol high reference, followed by the number of rules whose binary satisfaction score changes ( Small : 222 rules; Big : 242 rules). Each rule receives three judgments. An independent Sol high re-run provides a reference for run-to-run variation.
Audit Outcome
Assessments
Share (%)
Fixed-rate preferred
458
11.2
Adaptive preferred
701
17.1
Equivalent
2,848
69.4
Insufficient evidence
95
2.3
Total
4,102
100.0
Appendix
Table S19 : Rationale-Support Outcomes. Counts and percentages over all 4,102 paired rubric-item assessments.
Figure S22 : Visual evidence from adaptive frame selection. Each panel compares the same replay of a Small game. Adaptive selection captures (a) Confirm activation, (b) the path before removal, (c) holding and defense feedback, and (d) tile convergence and overlap. These cues are absent from the displayed uniform samples. Each row shows three cropped frames with timestamps.
Gameplay Agent
Adaptive Playtest
Judgment Coverage (%)
Game
Progress (%)
Requirements N
Normal
Normal + Adv.
Abyssal Chain
6.7±0.0
110
6.4
89.1
Alias Alchemy Shop
13.3±0.0
133
12.0
91.0
Beat Reroute
26.7±0.0
84
19.1
86.9
Chameleon Slide
13.3±0.0
87
0.0
67.8
Fogfall Delivery
15.6±10.2
119
5.0
88.2
Appendix
Table S20 : Recorded results for all ten games. Gameplay agent progress is milestone attainment (mean ± sample SD over three seeds); Judgment coverage is (S+V)/N over the archived GDD-derived requirements. The percentages use different target sets. Coverage is calculated from the saved judgments, with unverified requirements retained in the denominator.
Game
Source
Replay
Playtest
Sonic: Cascade Coast
99.71
53.50
80.00
Diablo Cathedral
86.74
55.00
66.55
Rocket League
83.64
57.40
77.54
Appendix
Table S21 : Evaluation of three additional 3D games. Scores are reported on a 0–100 scale. Playtest combines normal and adversarial judgments, using the full fixed rule set as the denominator.
Figure S23 : Selected GDD-specific repairs beyond self-revision. Initial builds are compared with second-round self-revision and Source + Replay + Playtest under matched inputs and preconditions. The cases test S23 confirmation before discarding a run, S23 CODE BLUE feedback on boss timeout, S23 opening the shop after alias selection, and S23 a diagonal attack against a target inside the specified hit area. Each displayed requirement remains unresolved after self-revision and is satisfied by Source + Replay + Playtest in the probe. Life and vigor values are recorded during execution; all targets in S23 start at vigor 1. Crops retain the archived builds’ native art; arrows indicate attack aim. These selected outcomes do not establish complete GDD compliance.
Figure S24 : Sampled gameplay in generated Small games (01). Four games from the Small split illustrate (a) selective clicking, (b) visual discrimination, (c) ring switching and collision feedback, and (d) the observation and input phases of a memory task. Each row shows three states in the order of the supplied capture sequence. The full game screens are retained, including their prompts and feedback.
Figure S25 : Sampled gameplays in generated Small games (02) Additional four frames illustrate (a) tracking a hidden key through a cup shuffle, (b) sorting curved and angular shapes, (c) accumulating and banking dice points, and (d) rotating a combination dial with persistent progress cues. Each row contains three full screenshots from one replay scenario, ordered from left to right.
Figure S26 : Sampled gameplays in generated Big games (1). Four games illustrate (a) fishing controls and catch rewards, (b) traversal tools and base defense, (c) unit deployment and tactical cards, and (d) delivery route and fleet management. Each row groups three scenario-specific screenshots from one game. Full screenshots retain the gameplay context and interface feedback.
Figure S27 : Sampled gameplays in generated Big games (2). Four games illustrate (a) light and shadow traversal, (b) soul encounters and cleansing, (c) production and persistent shop progress, and (d) excavation, restoration, and research. Each row groups three scenario-specific screenshots from one game. Full screenshots retain the gameplay context, controls, and state information.
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.
Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22
1The Chinese University of Hong Kong, Shenzhen · 2Shenzhen Loop Area Institute · 4USTB +5
Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A game can end in a valid state even after violating its rules during the run. Current game-development benchmarks replay fixed examples, score videos, or ask another model to judge the result. However, no existing benchmark checks game rules throughout execution across varied evaluator-selected scenarios while ensuring exactly reproducible verdicts. We introduce GameLogicBench, a benchmark of 72 gameplay-logic tasks in Godot projects. An automated evaluator checks each game's rules at every simulation tick. Across 403 hand-designed scenarios, seeded parameter variations produce 1,451 test cases. To ensure that the evaluator measures behavior rather than implementation choice, it must accept different correct implementations for each task while rejecting mutants, implementations with one required capability removed. The tasks span isolated mechanics, multi-system interactions, and repository-scale features. Across 20 combinations of language models and scaffolds, the best observed run solves 52.78% of tasks. Under Claude Code, all twelve models solve fewer tasks as task scope expands from isolated mechanics, through interacting systems, to repository-scale features. Agents inspect code more often and make more tool calls on repository-scale tasks than on isolated-mechanic tasks. Most unsuccessful submissions are runnable, but implement some required game behavior incorrectly. We compared versions of our benchmark evaluator built with and without validation using mutants. Without this validation, incorrect agent submissions passed. A separate analysis finds agents copying code from public repositories when network access is open. Reliable evaluation thus depends both on what the tests reject and on what external code agents can access.
Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate traces rather than the delivered application. We introduce WebGameBench, a requirement-to-application benchmark that evaluates whether coding agents can turn a frozen Structured WebGame Specification into a browser-accessible game. Browser-native games provide a compact but behavior-dense testbed: even simple games require coordinated input handling, spatial mapping, rule execution, state transitions, terminal conditions, restart behavior, and visible feedback. In WebGameBench, each generated artifact is built, served, and exposed as a browser-accessible application under a unified deployment protocol. A runtime evaluator then interacts with the delivered game in a real browser and assigns a three-way label: EXCELLENT, USABLE, or UNUSABLE. On a human-reviewed subset, the runtime label is broadly aligned with human gameplay review under the Usable-rate criterion. Across 111 tasks, 12 coding agents, and 14 evaluation configurations, WebGameBench separates current systems: the best configuration reaches a 76.9% usable rate but only a 20.2% excellent rate. This gap shows that crossing the minimum playable-delivery threshold is still far from complete requirement satisfaction. To our knowledge, WebGameBench is the first requirement-to-application benchmark for browser-native game delivery that validates delivered-application runtime labels against independent human gameplay review under the Usable-rate criterion.
Wenyu Zhang, Guoliang You, Tianlun +8
1Baidu · University of Science and Technology of China