Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.
Figures & tables
Figure 2: Recursive Game Creator’s multi-round refinement workflow. The Designer plans revisions and the Builder implements code and assets. The coding-native Player collects gameplay trajectories. The Reviewer combines behavioral and visual evidence with shared and game-specific criteria to produce preference-informed revision feedback. The central circular arrow denotes repeated refinement with a configurable number of rounds. Agentic-player refinement improves game quality across rounds, while human-involved refinement improves the play experience and its alignment with inferred player preferences.
Method
Action
Timing
Strategy
Simulation
Adventure
Overall
Codex + GPT-5.5 (high)
Vanilla
48.74
48.80
44.06
53.63
52.68
49.58
HoH@1 ( Yan et al., 2026 )
59.73
53.79
57.01
66.75
61.24
59.71
HoH@2
64.34
62.03
59.97
71.77
66.11
64.84
HoH@3
71.02
70.26
66.13
78.42
71.76
71.52
OpenCode + DeepSeek-V4-Pro
Table 1: GameCraft-Bench results (0–100, higher is better). @ k denotes refinement round k .
Model
Harness
Task success
L1
L2 Mean
L2 P0
L2 P1
L2 P2
GPT-6-Astra ultra
Codex CLI
26/47 (55.3%)
98.0
93.2
99.0
92.7
92.4
Claude-Opus-5
Claude Code
24/47 (51.1%)
99.6
90.4
91.2
89.9
91.1
GPT-5.6-Sol
Codex CLI
21/47 (44.7%)
98.9
91.3
100.0
89.7
91.9
DeepSeek-V4-Flash
Claude Code
18/47 (38.3%)
99.4
89.3
98.0
86.7
90.2
DeepSeek-V4-Pro
Claude Code
15/47 (31.9%)
99.6
85.2
85.3
83.9
88.9
Kimi-K3
Claude Code
15/47 (31.9%)
97.7
88.5
98.0
86.3
89.7
Table 2: GameASG-Bench results. Baselines are from Zhang et al. (2026c) , Table 3. All scores are percentages. Task success also reports the count. Bold marks column maxima. Green parenthesized gains in the final row are percentage-point improvements over GPT-6-Astra high with Codex CLI.
Figure 3: Qualitative analysis of multi-round RSI. The three panels compare development versions, organized as early, intermediate, and later snapshots (V1–V3). (a) Action-game combat, encounters, interfaces, and upgrade feedback. (b) Visual-novel dialogue, choices, illustrated events, and reading support. (c) Meme Arena scenes, skills, finishers, and control guidance.
Figure 4: Spatial exploration in an FPS test scene. The left panel shows three illustrative GUI-style routes from a shared spawn. The right panel shows visitation from 48 recorded coding-native policy rollouts. Color shows normalized visitation density. The source reports 97.1% grid coverage, including 11 interiors.
Round
Controls
Playability
Depth
Art
Avg. playtime (min)
Round 1
2
1
3
2
1.50
Round 2
5
4
4
6
4.87
Round 3
8
7
6
8
12.41
Table 3: Expert scores and average playtime across refinement rounds. Controls, playability, depth, and art are rated on a 1–10 scale. Average playtime is reported in minutes.
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
Wenyi Wu, Minghao Fu, Jieyu You +10
University of California San Diego · ByteDance Inc. · Carnegie Mellon University
Generating a game is not the same as making one that can be played. Despite advances in code generation, existing approaches treat game generation as one-shot translation from prompt to artifact, leaving interaction-level failures undetected. We argue that evaluating and improving game generation requires a player, and study two roles for graphical user interface (GUI) agents in this process: (1) as an objective evaluator, for which we introduce PlaytestArena, a new evaluation environment that pairs 200 browser-based game generation tasks across eight genres with rubrics of expected in-play behaviors, adjudicated by a GUI agent that loads each build in a browser and plays it; and (2) as a subjective playtester, for which we propose Play2Code, where a game agent and a GUI agent operate in a sustained loop with shared memory, turning game generation into a dialogue between coding and playing. Our experiments show that even frontier models struggle to generate playable games directly, while Play2Code achieves a 66.8% rubric pass-rate, improving over single-pass and agentic-coding baselines by 37.1 and 14.6 points respectively. Further analysis shows that GUI playtester feedback is more traceable than a human report, yet idiosyncratic in ways reminiscent of human testers, establishing game playtesting as a critical testbed for interactive code generation. Our project website is available at https://continual-game-generation.vercel.app/.
Yixu Huang, Bo Li, Na Li +8
αFudan University · βXiaohongshu Inc. · γTongji University +2
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.
Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22
1The Chinese University of Hong Kong, Shenzhen · 2Shenzhen Loop Area Institute · 4USTB +5