SWE-Game: Can Coding Agents Build the Games We Want?
Authors: Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, +3 more
Organizations: Shanghai Jiao Tong University · Zhongguancun Academy · Shenzhen University · Elbetech Technology · Beijing University of Posts and Telecommunications · Shanghai Innovation Institute
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.
Figures & tables
Figure 1: SWE-Game: from reference gameplay to playable games. The center illustrates an agent using supplied assets and gameplay demonstrations to recreate the target appearance and behavior through coding and play-testing. Five tasks vary the requirements and starting project, covering construction, repair, and cross-engine porting.
Figure 2: Overview of SWE-Game. Top: agent development from reference materials through coding and execution, followed by assessment using functional tests, visual rubrics, and human validation. Visual rubric scoring applies to construction tasks. Bottom: reference-game construction, five development tasks and their evaluation, and game coverage. The gameplay panels use constructed examples to illustrate functional failure, visual failure, and correct reference behavior.
Figure 3: Reference game statistics. (a) Median numbers of Godot script and scene files (bars) and GDScript code lines (line) within each gameplay category. (b) Distributions of reference-video duration (inner ring, minutes) and frame count (outer ring). (c) Supplied asset-file counts by gameplay category, with one point per game on a logarithmic scale.
Benchmark
Runtime
Input
Tasks
Evaluation
Scoring unit
Det. runtime checks
OpenGame ( 2026 )
Web
T
G
VLM judge
Game, dimension
×
WebGameBench ( 2026b )
Web
T
G
Agent judge
Requirement
×
V-GameGym ( 2026a )
Pygame
T
G
VLM judge
Game, dimension
×
VERIGAME ( 2026 )
Web
T
G
State inj., VLM
Keypoint
×
GameXpert-Bench ( 2026 )
Web
T, S
G, O, R
Runtime tests, human
Event, assertion
∘
GameCraft-Bench ( 2026 )
Godot
T, A
G
VLM judge
Rubric item
×
Table 1: Comparison of Game-Development Benchmark Protocols
Task
Agent inputs
Required artifacts
Evaluation components
Brief-to-Game
Brief, assets, reference video
GDD, Godot project, feature demos
Mechanics, content, playability, design, VLM
GDD-to-Game
GDD, assets, reference video
Godot project, feature demos
Mechanics, content, playability, VLM
Skeleton Completion
Code skeleton, requirements, assets, reference video
Table 2: Task modes in SWE-Game: agent inputs, required artifacts, and evaluation components.
Task
Model
Mechanics
Content
Playability
Design
VLM
Total
Brief-to-Game
Qwen3.8 Flash
21.10
23.32
28.38
65.85
40.40
27.00
Grok4.6
25.66
35.34
61.21
70.00
48.39
39.01
GPT-5.6 Luna
24.17
21.29
55.51
92.68
60.05
34.18
Opus5
27.55
48.26
80.99
53.85
69.23
50.38
GLM5.3 Flash
17.25
33.07
44.51
61.33
33.03
31.05
Minimax M3
25.30
27.20
42.65
67.50
32.09
30.47
Table 3: Component and overall scores on SWE-Game (0 to 100; higher is better).
Figure 4: Resource use, task performance, and problem composition on SWE-Game. Panels (a) and (b) compare task scores against mean input and output tokens, respectively; colors identify models and shapes identify tasks. Dashed lines connect the selected within-task nondominated configurations. Panel (c) summarizes total scores across the five tasks. Panel (d) shows the distribution of problem categories within each model.
Video
Game- play
Content
Feedback
Visual
VLM Total
GPT-5.6 Luna
On
68.87
60.69
63.24
55.92
60.05
Off
58.57
66.70
53.10
49.44
54.45
Opus5
On
73.54
70.55
78.06
62.45
69.23
Off
77.48
80.48
76.73
53.45
67.00
Table 4: Reference-video ablation in brief-to-game generation. Scores are on a 0 to 100 scale.
Functional judgments (%)
Metric
Executable
Video VLM
Correct acceptance
90.93
81.41
Defect detection
94.25
75.40
Balanced accuracy
92.59
≈78.41
Reviewed games/samples
100/1200
100/1200
Table 5: Human validation of executable and visual evaluation.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Brief-to- Game
GDD-to- Game
Skeleton Completion
Godot-to-Unity Porting
Mechanics
26.71
27.50
31.57
35.00
Content
42.50
43.75
34.00
–
Playability
13.36
13.75
8.50
25.00
Design
2.43
–
–
–
Scaffold
–
–
10.93
–
VLM
15.00
15.00
15.00
–
Appendix
Table 9: Component weights (%) for construction and porting. A dash indicates that the component does not apply. Bug Repair uses multiplicative aggregation.