SWE-Game: Can Coding Agents Build the Games We Want?
Authors: Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, +3 more
Organizations: Shanghai Jiao Tong University · Zhongguancun Academy · Shenzhen University · Elbetech Technology · Beijing University of Posts and Telecommunications · Shanghai Innovation Institute
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.
Figures & tables
Figure 1: SWE-Game: from reference gameplay to playable games. The center illustrates an agent using supplied assets and gameplay demonstrations to recreate the target appearance and behavior through coding and play-testing. Five tasks vary the requirements and starting project, covering construction, repair, and cross-engine porting.
Figure 2: Overview of SWE-Game. Top: agent development from reference materials through coding and execution, followed by assessment using functional tests, visual rubrics, and human validation. Visual rubric scoring applies to construction tasks. Bottom: reference-game construction, five development tasks and their evaluation, and game coverage. The gameplay panels use constructed examples to illustrate functional failure, visual failure, and correct reference behavior.
Figure 3: Reference game statistics. (a) Median numbers of Godot script and scene files (bars) and GDScript code lines (line) within each gameplay category. (b) Distributions of reference-video duration (inner ring, minutes) and frame count (outer ring). (c) Supplied asset-file counts by gameplay category, with one point per game on a logarithmic scale.
Benchmark
Runtime
Input
Tasks
Evaluation
Scoring unit
Det. runtime checks
OpenGame ( 2026 )
Web
T
G
VLM judge
Game, dimension
×
WebGameBench ( 2026b )
Web
T
G
Agent judge
Requirement
×
V-GameGym ( 2026a )
Pygame
T
G
VLM judge
Game, dimension
×
VERIGAME ( 2026 )
Web
T
G
State inj., VLM
Keypoint
×
GameXpert-Bench ( 2026 )
Web
T, S
G, O, R
Runtime tests, human
Event, assertion
∘
GameCraft-Bench ( 2026 )
Godot
T, A
G
VLM judge
Rubric item
×
Table 1: Comparison of Game-Development Benchmark Protocols
Task
Agent inputs
Required artifacts
Evaluation components
Brief-to-Game
Brief, assets, reference video
GDD, Godot project, feature demos
Mechanics, content, playability, design, VLM
GDD-to-Game
GDD, assets, reference video
Godot project, feature demos
Mechanics, content, playability, VLM
Skeleton Completion
Code skeleton, requirements, assets, reference video
Table 2: Task modes in SWE-Game: agent inputs, required artifacts, and evaluation components.
Task
Model
Mechanics
Content
Playability
Design
VLM
Total
Brief-to-Game
Qwen3.8 Flash
21.10
23.32
28.38
65.85
40.40
27.00
Grok4.6
25.66
35.34
61.21
70.00
48.39
39.01
GPT-5.6 Luna
24.17
21.29
55.51
92.68
60.05
34.18
Opus5
27.55
48.26
80.99
53.85
69.23
50.38
GLM5.3 Flash
17.25
33.07
44.51
61.33
33.03
31.05
Minimax M3
25.30
27.20
42.65
67.50
32.09
30.47
Table 3: Component and overall scores on SWE-Game (0 to 100; higher is better).
Figure 4: Resource use, task performance, and problem composition on SWE-Game. Panels (a) and (b) compare task scores against mean input and output tokens, respectively; colors identify models and shapes identify tasks. Dashed lines connect the selected within-task nondominated configurations. Panel (c) summarizes total scores across the five tasks. Panel (d) shows the distribution of problem categories within each model.
Video
Game- play
Content
Feedback
Visual
VLM Total
GPT-5.6 Luna
On
68.87
60.69
63.24
55.92
60.05
Off
58.57
66.70
53.10
49.44
54.45
Opus5
On
73.54
70.55
78.06
62.45
69.23
Off
77.48
80.48
76.73
53.45
67.00
Table 4: Reference-video ablation in brief-to-game generation. Scores are on a 0 to 100 scale.
Functional judgments (%)
Metric
Executable
Video VLM
Correct acceptance
90.93
81.41
Defect detection
94.25
75.40
Balanced accuracy
92.59
≈78.41
Reviewed games/samples
100/1200
100/1200
Table 5: Human validation of executable and visual evaluation.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Brief-to- Game
GDD-to- Game
Skeleton Completion
Godot-to-Unity Porting
Mechanics
26.71
27.50
31.57
35.00
Content
42.50
43.75
34.00
–
Playability
13.36
13.75
8.50
25.00
Design
2.43
–
–
–
Scaffold
–
–
10.93
–
VLM
15.00
15.00
15.00
–
Appendix
Table 9: Component weights (%) for construction and porting. A dash indicates that the component does not apply. Bug Repair uses multiplicative aggregation.
Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A game can end in a valid state even after violating its rules during the run. Current game-development benchmarks replay fixed examples, score videos, or ask another model to judge the result. However, no existing benchmark checks game rules throughout execution across varied evaluator-selected scenarios while ensuring exactly reproducible verdicts. We introduce GameLogicBench, a benchmark of 72 gameplay-logic tasks in Godot projects. An automated evaluator checks each game's rules at every simulation tick. Across 403 hand-designed scenarios, seeded parameter variations produce 1,451 test cases. To ensure that the evaluator measures behavior rather than implementation choice, it must accept different correct implementations for each task while rejecting mutants, implementations with one required capability removed. The tasks span isolated mechanics, multi-system interactions, and repository-scale features. Across 20 combinations of language models and scaffolds, the best observed run solves 52.78% of tasks. Under Claude Code, all twelve models solve fewer tasks as task scope expands from isolated mechanics, through interacting systems, to repository-scale features. Agents inspect code more often and make more tool calls on repository-scale tasks than on isolated-mechanic tasks. Most unsuccessful submissions are runnable, but implement some required game behavior incorrectly. We compared versions of our benchmark evaluator built with and without validation using mutants. Without this validation, incorrect agent submissions passed. A separate analysis finds agents copying code from public repositories when network access is open. Reliable evaluation thus depends both on what the tests reject and on what external code agents can access.
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.
Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22
1The Chinese University of Hong Kong, Shenzhen · 2Shenzhen Loop Area Institute · 4USTB +5
Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at https://a2z-gamespec-bench.github.io.
Seonho Lee, Wonryeol Jeong, Alberto Cereser +4
KRAFTON · KAIST · Korea National University of Arts