Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
Figures & tables
Figure 1: RSIGame turns agentic game development into autonomous recursive self-improvement. It combines a local explore-diagnose-improve loop with global progress monitoring and control, enabling Qwen3.8-27B approach GPT-5.5-level performance.
Figure 2: Overview of RSIGame . The local loop autonomously evolves the game through direction decision, active exploration, evidence-grounded editing, and agentic verification. The global loop monitors global quality, preserves the best checkpoint, and introduces sparse high-level guidance when progress saturates, enabling progressive multi-stage game evolution.
Method
Mechanics
Depth
Visuals
Art
Overall ↑
Tok. ↓
Cost ↓
Godot ( Godot Engine contributors, 2026 )
Generator: Codex + GPT-5.5 (high)
Base (frozen P0 )
58.9
51.7
51.4
44.6
50.26
0.26M
–
+ Play2Code
59.2
51.3
52.2
45.8
50.74
3.20M
$1.06
+ RSIGame
72.8
60.3
65.9
64.5
64.53
3.12M
$0.88
Generator: Codex + Kimi-K2.6
Table 1: Main results on GameCraft-Bench (top: Godot; bottom: Phaser). Methods are compared from matched base games under one backbone and budget, averaged over 140 tasks; Tok. and Cost are mean billable tokens and cost per task (Appendix D.1 ). Bold marks the best value per column within a group; deeper rows internalize development experience.
Figure 3: Development-time scaling on GameCraft-Bench (40 tasks). Across Godot and Phaser with strong (GPT-5.5) and weak (Qwen3.8-27B) initial generators, iterative development can plateau or regress, while RSIGame achieves sustained improvement across development budgets. The Global Quality Monitor preserves the best checkpoint reached so far, improving the quality–cost tradeoff over returning the last checkpoint. Shaded regions denote ±1 standard error.
Figure 4: Global quality control and experience transfer. ( a ) Best-checkpoint tracking protects earlier gains from later regressions. ( b ) Saturation-aware stopping achieves comparable quality with fewer development rounds. ( c ) Internalizing verified development experience improves one-shot generation across all quality dimensions.
Figure 5: Adaptive development follows the game state and improves final quality. ( a, b ) Adaptive development reallocates effort according to the current bottleneck, shifting toward improvement after build failures and toward art when visual quality lags. ( c, d ) Blind pairwise evaluation consistently favors adaptive development over round-robin scheduling.
Table 2: Agentic verification improves both target grounding and post-edit reliability. Before editing, it filters unsupported improvement targets; after editing, replay-based verification detects failed changes and regressions that compilation alone cannot reveal.
Appendix figures & tables38 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Contents / specification
Location
Released artifacts
Implementation
Local/global loops, Global Quality Monitor, director-review tool, and scoring harness
Code, live
Prompts and output contracts
Complete prompts and structured outputs used by all agents
Code, live; App. C.4
Playable evolution
Six representative games with retained checkpoints
Project page, live
Frozen base projects P0
Starting project for each task and generator
On release
Evaluation artifacts
Per-task rubric scores, judge outputs, and replay reports (51,644 files)
HF, live †
Appendix
Table 3: Reproducibility index. Artifacts, release status, and locations of the corresponding implementation and protocol details.
Figure 6: The director review page, mid-session, on the Alley Brawlers build that had survived six monitor checkpoints: the live game, the budget, the task specification, and the development summary — champions at R03 and R09, no update since, and the split of the eighteen rounds between functional and presentational work. The frame is a real moment of play and shows the defect both directors went on to describe in different words: the player’s fighter has come apart into a band of scrambled sprite tiles. The brief form below it is Figure 7 .
Figure 7: The brief form, the lower half of the page in Figure 6 . The model director submits the same four fields as JSON through the endpoint this form posts to.
Human director — 60/60 actions, 0 resets.
Objective
The art needs work: the title and game-over screens are cluttered and have no proper background, the character art is ugly and should be regenerated in a comic style, and attacks have no effects. The game also needs to teach how it is played — right now I do not know what to press.
Why now
The characters move and there are title and game-over screens, but it is still too rough to be appealing.
Model director (Claude Opus 5) — 25/60 actions, 0 resets.
Objective
Turn the match from a scripted beatdown into a real two-sided fight: player inputs must visibly and reliably land, block and escape, against an AI that can be beaten, with two fighters that genuinely look and play differently.
Why now
Across two full matches (one as each fighter) the AI never lost a pixel of health. Every round ended in about 18 seconds with me pinned in the left corner in a hit, knockdown, hit loop. Light, heavy, the special input, holding back, crouching and jumping all failed to trade, block or get me out. Nobody can learn to play when nothing they press matters. […] Seven rounds of visual polish have gone into the frame around the fight (title, background, banners, results screen), but the fight itself is still missing.
Appendix
Table 4: The two briefs written for the same Alley Brawlers build (tree c12fe90b30b6c176 ), abridged to their first two fields. Priorities and preserve items are omitted for space.
Figure 8: The case list a director opens: six games, the build under review for each, and the budget remaining. Case identifiers are internal labels.
Stage
Module
Output contract
Consumed by
Direction Decision
Controller
One development_question ; action in { reuse , explore , escalate }; an exploration_plan naming the cheapest mode that can answer it.
Sets the round’s objective; escalate reports the direction exhausted rather than inventing work.
Agentic Exploration
Explorer
The 1–3 requirements a session should settle, the scenario to boot, and one actionable instruction for the play agent.
Drives the test agent’s interaction with Pt , producing the trace τt .
Evidence-Grounded Editing
Editor
Edits to the project under a build constraint, plus a report of what was changed.
Produces Pt+1 .
Agentic Verification
Verifier
goal_achieved in {yes, no, unclear}; broken_now ; next_goal when that list is non-empty.
Accepts or rejects the round; anything broken re-enters the state as pending work.
Global Quality Monitoring
Global Quality Monitor
Per-criterion readings for both builds, changes with direction and magnitude, and regressions — never a score.
Aggregated into the proxy value that decides the champion (Appendix C.2 ).
Appendix
Table 5: The five LLM-facing modules of the RSIGame loop, by the stage of Section 3 they implement.
Role
Model
Served by
Decoding
Generator
one per row group of
harness
harness
Table 1
default
default
Controller
GLM-5.3-Flash
OpenRouter
T=0 , top- p 1
Explorer
Qwen3.8-27B
local vLLM (TP = 2)
T=0 , top- p 1
Verifier
GLM-5.3-Flash
OpenRouter
T=0 , top- p 1
Editor
GLM-5.3-Flash
OpenRouter
T=0 , top- p 1
Appendix
Table 6: Models by role. The generator differs per row group; every other role is the same model in every row, so that a comparison between rows is a comparison of development methods and not of the models inside them. Decoding is greedy wherever the answer is a claim about the game: a coordinate, a verdict, a rubric item. Served by matters for cost and for reproducibility — one model on OpenRouter is sold by a dozen providers, so the judge and the monitor pin theirs.
Parameter
Value
Round
Tool calls the Editor may make per round
26
Wall-clock budget per improvement round
1800 s
Improvement attempts per round before the round is given up
3
Verification samples per claimed fix
3
Probe steps taken before an edit is proposed
8
Run
Generated assets per run
20
Appendix
Table 7: Loop, monitor and scoring parameters. One value per knob, the same in every run reported here.
Model
Provider
Input
Cache read
Output
GPT-5.5
OpenAI
5.000
0.500
30.00
Kimi-K2.6
Baidu
0.408
0.069
0 1.72
Qwen3.8-27B
DeepInfra
0.150
0.037
0 1.88
GLM-5.3-Flash
DeepInfra
0.075
0.015
0 0.25
Appendix
Table 8: OpenRouter list prices ($ per million tokens), retrieved 2026-09-16. Each model is priced at its cheapest standard endpoint for the cache mix of the runs reported here; the serving provider is named because prices for one model differ by up to a factor of two across providers.
Engine
Generator
Billable
Total
Cache
Output
Cost
Godot
Codex + GPT-5.5 (high)
0.26M
0 3.83M
94%
0 28k
$ 0 3.77
Godot
GLM-5.3-Flash
0.44M
11.43M
96%
0 36k
$ 0 0.20
Godot
Kimi-K2.6
3.22M
0 8.98M
65%
123k
$ 0 1.87
Godot
Qwen3.8-27B
6.41M
11.62M
46%
179k
$ 0 1.47
Godot
Qwen3.8-27B (SFT)
0.57M
0 8.21M
95%
140k
–
Phaser
OpenGame + GPT-5.5
5.36M
0 9.21M
42%
0 97k
$31.17
Appendix
Table 9: Generation of the frozen base P0 , mean per task. Billable is the Tok. of the Base rows; Total adds cache reads; Cost applies Table 8 to all three token counts. The two GPT-5.5 rows are the same model under two harnesses and differ in price: Codex re-reads a cached prefix, OpenGame does not. Over all 140 tasks.
Subset
n
p25
Median
p75
Identical 3/3
All artifacts
529
0.00
0.73
3.33
225
Base P0
266
0.00
0.36
2.33
131
Improved
263
0.00
0.87
3.50
0 94
Appendix
Table 10: Score variation across three independent end-to-end evaluations of the same frozen artifact. Each run uses a fresh gameplay replay and fresh judging by the Qwen3.8-27B judge. Range denotes the difference between the highest and lowest Overall scores across the three runs; the columns are its quartiles.
Family
n
Median
Family
n
Median
sports
14
4.66
shooter
24
0.44
horror
19
3.37
puzzle
32
0.42
openworld
62
1.41
simulation
24
0.28
visualnovel
32
1.17
strategy
63
0.00
platformer
77
0.87
rhythm
21
0.00
tycoon
58
0.73
idle
16
0.00
Appendix
Table 11: Replay variation by task family, for the families with at least twelve scored artifacts, ordered by median range. Same 529 artifacts and same definition of range as Table 10 .
Overall
Method
Qwen3.8-27B
GPT-5.5
Base (frozen P0 )
49.07
50.34
+ Play2Code
49.63
50.62
+ RSIGame
63.10
62.38
Comparison
Δ Qwen
Δ GPT-5.5
95% CI Qwen
95% CI GPT-5.5
Agreement
RSIGame − Base
+14.04
+12.04
[+9.84,+18.59]
[+9.53,+14.60]
37/40
Appendix
Table 12: Cross-judge robustness on 40 family-stratified Godot tasks. Qwen3.8-27B and GPT-5.5 score the same replay recordings with the same rubric. Both judges preserve the method ordering and yield similar pairwise improvement margins. Confidence intervals are 95% task-level bootstrap intervals.
Comparison
W/L/T
p
Agreement
RSIGame – Base
17/3/0
0.003
45/59
RSIGame – Play2Code
18/2/0
<0.001
Play2Code – Base
9/11/0
0.824
Appendix
Table 13: Free-play evaluation on 20 family-stratified Godot tasks. Blind agents independently play each build, and a separate judge compares their reports using the same quality dimensions and weighting as the main evaluation. p is a two-sided sign test over decided comparisons.
Comparison
Δ
95% CI
p
W/L
Godot
GPT-5.5: RSIGame – Base
+14.26
[11.49, 17.06]
<10−4
112/16
GPT-5.5: RSIGame – Play2Code
+13.79
[10.89, 16.65]
<10−4
118/19
Qwen3.8-27B: RSIGame – Base
+10.70
[7.75, 13.77]
<10−4
100/31
Qwen3.8-27B: RSIGame – Play2Code
+7.24
[4.49, 10.04]
<10−4
92/41
Phaser
Appendix
Table 14: Paired significance of the main results. Δ is the mean task-level Overall difference. Confidence intervals are obtained from 20,000 paired bootstrap resamples of benchmark tasks; p is from a Wilcoxon signed-rank test. W/L counts tasks on which the first method scores higher/lower.
Figure 9: Blind pairwise win rate of the later stage in each pair, with ±1 s.e.
Stage
Unit
Teacher
Godot
Phaser
Total
S1 generation
trajectory
GPT-5.5
1,929
284
2,213
S1 generation
decision row
GPT-5.5
66,312
28,097
94,409
S2 planning
plan
distilled
1,979
129
2,108
S3 improvement
round
GLM-5.3-Flash
3,730
273
4,003
of which verified
round
—
1,869
144
2,013
S3 improvement agent
decision row
GLM-5.3-Flash
111,302
9,930
121,232
Appendix
Table 15: The released training corpus, counted from the published tables. A decision row is one assistant turn together with the history that preceded it, so decision rows repeat context and their token counts are not independent; the trajectory rows are the same recordings un-windowed and are what to count tokens over. Verified improvement rounds are those an independent post-improvement verification judged to have achieved the stated goal; the remainder are released as negatives rather than dropped.
Stage added
Unit
Godot
Phaser
Rows added
Cumulative
S1 generation
agent trajectory
876
418
1,294
1,294
S2 planning
brief → file plan
216
0 92
0, 308
1,602
S3 improvement
improvement round
509
00 0
0, 509
2,111
Total
1,601
510
2,111
Appendix
Table 16: The training corpus of each variant in Table 18 . Each stage adds rows to the previous corpus; the cumulative column is what that variant was trained on. Tasks in the held-out brief list are filtered from every stage.
Configuration
Value
Base model
Qwen3.8-27B
Adaptation
LoRA ( Hu et al., 2022 )
LoRA rank r
16
LoRA α
32
LoRA dropout
0.05
LoRA targets
All linear projections
Appendix
Table 17: Fine-tuning configuration used for experience internalization. These are the settings of the released adapter; every variant in Table 18 is trained with them and differs only in the corpus of Table 16 .
Variant
Mechanics
Depth
Visuals
Art
Overall ↑
Base
41.2
33.4
38.6
38.3
37.07
+Gen
49.0
41.4
42.0
38.0
41.45
+Gen+Plan
48.5
40.8
42.6
39.5
41.77
+Gen+Plan+Improve
56.1
47.5
49.8
44.8
48.22
Appendix
Table 18: One-shot Godot generation by training variant. Every row is over the 140 tasks.
the 12 shared tasks
VibeGame’s own 40
VibeGame
23.3
30.5
OpenGame + GPT-5.5, one shot
50.1
46.0
+ Play2Code
56.4
—
+ RSIGame
62.4
—
Appendix
Table 19: VibeGame against one-shot generation and against RSIGame . Left: the twelve tasks shared by VibeGame’s forty and our forty-task Phaser development subset, so every column is the same twelve games. Right: VibeGame’s own forty tasks, against the one-shot reference its release quotes on them.
Figure 10: Lawn Guardians , from the generated game to the build we shipped. The score band gives the official score at every third round of the 30-round autonomous run, the version the Global Quality Monitor holds, and the star where the saturation stop delivers it at round 21, nine rounds early. The round band gives one cell per round — improvement or art — with the post-improvement verdict above it. Past the axis break are the passes the loop did not run: the stage brief written once the champion had survived three checkpoints, and the three iterations after it, which are development passes rather than loop rounds and so carry no round numbers. Each build below is opened up as a card in Figures 11 – 12 .
Figure 11: Lawn Guardians from the generated build to the stage brief. Each card carries what that round or pass found, what it changed, and the frames it is judged on; boxes over the screenshots are placed by eye, green where something arrived and red where it is still wrong, and every build is driven through the same script so the scenes are comparable. Base game: the generator has the rules right and the presentation wrong — the art it shipped is on disk and unused. Autonomous: twelve rounds wire that art in and add hit feedback, and the champion that emerges is never beaten. Stage 1: with the champion held for three checkpoints the loop has run out of its own evidence, and a brief written from outside — by Claude Opus 5, alongside the human’s — puts sun on the lawn and zombies inside their lanes.
Figure 12: The three development passes that follow the stage brief. Iteration 1 is a fix no still of the board can show, captured instead as two before/after pairs under the same script: at a stretched canvas the click was transformed twice, so a plant landed a cell from the cursor, and peas stayed on the lawn after connecting. Iteration 2 replaces the title screen and the walk cycle, the one step a still frame does carry. Iteration 3 is set dressing and a rewritten soundtrack; the seed bar and verge are in the frames, and the audio is the half no figure can show.
Figure 13: Alley Brawlers , from the generated game to the build we shipped. Score band, round band and frames as in Figure 10 . The saturation stop ends the autonomous run at round 18 and delivers round 9. Past the axis break are the passes the loop did not run: two stage briefs, each written by a director who played the build in front of it, and three development passes. Stage 2 and the iterations both continue from the stage-1 champion, so they are alternatives rather than a sequence.
Figure 14: Alley Brawlers from the generated build to the first brief. Cards, boxes and script as in Figure 11 . Base game: a working match on a flat strip under falling squares. Autonomous: the arena arrives and the fighters break, and the delivered build scores exactly what the base game scored. Stage 1: a director who played two matches asks for a fair exchange and for a fighter that stays one intact figure; the score moves for the first time in the run.
Figure 15: The second brief and the development passes, both continuing from the stage-1 champion. Stage 2: a second director names spacing as the reason a player cannot act, and the exchange becomes legible from both sides. Iterations 1–2 rebuild the front end — the title becomes steps rather than a wall of text, and the HUD names and shows both fighters — and add hit sparks and damage numbers. Iteration 3 is the feel of contact: hitstop, screen shake, K.O. slow motion, a combo counter and a soundtrack.
Figure 16: Block Cascade , from the generated game to the build we shipped. The generated game already implements every mechanic in the spec, so the improvement rounds have almost nothing to improve: through nine checkpoints the Global Quality Monitor never prefers a later build, the saturation stop ends the run at round 9, and what it delivers is round 0 itself. Everything past the axis break follows a restart from that build.
Figure 17: Block Cascade from the generated build through both briefs. Cards, boxes and script as in Figure 11 . Base game: every mechanic in the spec works first try, and nine rounds later the loop delivers this same build. Stage 1: the well gets a painted ground, faceted blocks and a danger line — and a control branch with no brief reaches the same score. Stage 2: a named mode and a live speed readout arrive, printed on top of the readouts already there.
Figure 18: The two development passes that close the run. Iteration 1 gives every number a place of its own and a landing its feedback. Iteration 2 adds a centred result card over a dimmed board, a HOLD box that says which key fills it, and a written soundtrack in place of a 1.8-second fragment.
Family
N
Base ↑
+P2C
+ RSIGame
idle
4
81.9
86.8
87.5
sports
4
72.8
70.5
81.2
horror
5
67.2
71.4
75.1
openworld
15
59.2
57.9
62.2
tycoon
16
53.3
54.4
73.2
visualnovel
11
52.0
55.8
71.5
Appendix
Table 20: Per-family breakdown on Godot for the Codex + GPT-5.5 (high) generator (Table 1 , first row group), by mean Overall ( ↑ ), Qwen3.8-27B judge. N is the number of tasks in the family (scored/total while scoring is in progress); families are ordered by base score; a task with no generated project scores 0. Abbreviations: +P2C Play2Code; –: not yet available on all tasks.
Family
N
Base ↑
+P2C
+ RSIGame
idle
4
53.5
50.1
60.0
sports
4
45.9
57.3
64.4
horror
5
43.1
56.6
63.4
tycoon
16
36.1
40.7
46.4
rhythm
5
33.7
37.1
45.4
shooter
7
33.4
35.9
39.1
Appendix
Table 21: Per-family breakdown on Godot for the Kimi-K2.6 generator (Table 1 , second row group).
Family
N
Base ↑
+P2C
+ RSIGame
sports
4
49.6
61.5
72.6
idle
4
39.0
59.2
68.6
tycoon
16
38.8
43.1
56.1
racing
4
34.9
36.7
53.0
horror
5
34.1
43.6
49.9
visualnovel
11
32.7
38.8
59.0
Appendix
Table 22: Per-family breakdown on Godot for the GLM-5.3-Flash generator (Table 1 , third row group).
Family
N
Base ↑
+P2C
+ RSIGame
horror
5
58.7
56.6
57.0
idle
4
48.8
53.2
73.9
cardgame
5
48.3
53.1
59.3
sports
4
44.8
51.2
55.6
shooter
7
42.3
44.7
61.5
visualnovel
11
42.0
46.0
44.9
Appendix
Table 23: Per-family breakdown on Godot for the Qwen3.8-27B generator (Table 1 , fourth row group).
Family
N
Base ↑
Play2Code ↑
RSIGame ↑
idle
4
80.4
78.6
84.1
sports
4
73.3
68.7
76.1
horror
5
60.6
62.4
71.5
tycoon
16
54.7
56.4
66.3
strategy
17
52.1
51.0
62.0
visualnovel
11
50.2
48.9
58.4
Appendix
Table 24: Per-family breakdown on Godot for the Qwen3.8-27B (SFT) generator (Table 1 , fifth row group).
Family
N
Base ↑
+P2C
+ RSIGame
idle
4
72.7
73.8
78.2
cardgame
5
68.9
68.9
61.8
visualnovel
11
68.0
68.6
73.5
tycoon
16
59.9
62.9
74.3
sports
4
57.2
59.8
64.8
roguelike
14
56.5
58.2
66.7
Appendix
Table 25: Per-family breakdown on Phaser for the OpenGame + GPT-5.5 generator (Table 1 , Phaser block, first row group).
Family
N
Base ↑
+P2C
+ RSIGame
sports
4
56.1
68.8
65.4
horror
5
55.6
52.9
57.6
simulation
6
55.3
53.3
67.8
visualnovel
11
53.3
54.2
52.4
rhythm
5
50.9
58.4
61.4
tycoon
16
47.5
58.9
56.7
Appendix
Table 26: Per-family breakdown on Phaser for the Qwen3.8-27B generator (Table 1 , Phaser block, second row group).
Family
N
Base ↑
Play2Code ↑
RSIGame ↑
idle
4
73.0
75.2
78.5
horror
5
56.0
60.7
67.5
tycoon
16
53.3
62.0
66.5
sports
4
51.6
49.9
56.8
shooter
7
41.8
51.4
55.4
rhythm
5
46.8
61.3
66.2
Appendix
Table 27: Per-family breakdown on Phaser for the Qwen3.8-27B (SFT) generator (Table 1 , Phaser block, third row group).
Large Language Models (LLMs) have shown great ability in generating executable code from natural language, opening the possibility of automatically constructing environments for AI agents. Recent work on Code World Models (CWMs) demonstrates that LLMs can translate game rules into Python implementations compatible with solvers like Monte Carlo Tree Search. We study this problem in game settings, where generated environments must implement rules, legal actions, state transitions, observations, and rewards. We refer to these game-specific executable models as Game Code World Models (GameCWMs). However, current approaches to generating code world models rely on frontier models and inference-time refinement loops, limiting accessibility and scalability. This work investigates whether GameCWM generation capabilities can be distilled into smaller models through post-training. We introduce: (1) a curated dataset of 30 games spanning perfect and imperfect information games, (2) a verification framework that evaluates generated code against structural and semantic game properties, and (3) a post-training pipeline combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR). We experiment with Qwen2.5-3B-Instruct and find that SFT can increase syntactic correctness, while RLVR can improve execution-level adherence to game rules, thereby improving Qwen's ability to generate valid GameCWMs in both perfect and imperfect information games. Overall, our pipeline makes Qwen2.5-3B-Instruct more capable of generating valid GameCWMs, thereby offering a scalable path toward automatic environment generation from natural language.
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.
Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22
1The Chinese University of Hong Kong, Shenzhen · 2Shenzhen Loop Area Institute · 4USTB +5
Large language models can generate plausible game code, but turning this capability into \emph{iterative creative improvement} remains difficult. In practice, single-shot generation often produces brittle runtime behavior, weak accumulation of experience across versions, and creativity scores that are too subjective to serve as reliable optimization signals. A further limitation is that mechanics are frequently treated only as post-hoc descriptions, rather than as explicit objects that can be planned, tracked, preserved, and evaluated during generation. This report presents \textbf{CreativeGame}, a multi-agent system for iterative HTML5 game generation that addresses these issues through four coupled ideas: a proxy reward centered on programmatic signals rather than pure LLM judgment; lineage-scoped memory for cross-version experience accumulation; runtime validation integrated into both repair and reward; and a mechanic-guided planning loop in which retrieved mechanic knowledge is converted into an explicit mechanic plan before code generation begins. The goal is not merely to produce a playable artifact in one step, but to support interpretable version-to-version evolution. The current system contains 71 stored lineages, 88 saved nodes, and a 774-entry global mechanic archive, implemented in 6{,}181 lines of Python together with inspection and visualization tooling. The system is therefore substantial enough to support architectural analysis, reward inspection, and real lineage-level case studies rather than only prompt-level demos. A real 4-generation lineage shows that mechanic-level innovation can emerge in later versions and can be inspected directly through version-to-version records. The central contribution is therefore not only game generation, but a concrete pipeline for observing progressive evolution through explicit mechanic change.
Hongnan Ma, Han Wang, Shenglin Wang +6
University of Bristol · Shanghai Jiao Tong University · Shandong University +1