Generating a high-quality gaming world from a natural-language game intent requires joint reasoning about scene structure, spatial layout, gameplay objectives, interactive entities, and executable gameplay logic. Existing coding agents can generate individual assets, scenes, or scripts, but often struggle to maintain consistency across these components. We propose Code2Games, an agentic framework that builds a structured gaming world upon a base Blender world generated from the same game intent. Code2Games coordinates scene analysis, gameplay planning, constrained gaming-world generation, and gaming-engine customization through a shared scene-gameplay representation with persistent element correspondence. After world generation, Code2Games adapts the generated world to Unreal Engine 5 and employs an execution-guided reconstruction process that uses compilation diagnostics, runtime feedback, and gameplay test results to resolve inconsistencies arising during engine adaptation. To systematically evaluate gaming-world generation, we introduce the GameCode4D benchmark, which comprises ten fixed game prompts spanning different levels of scene and gameplay complexity. We evaluate the generated results across four dimensions: visual quality, interactive fidelity, multimodal artifact quality, and playable-game quality. Experiments demonstrate that, compared with direct gaming-world generation by coding agents and existing baseline methods, Code2Games consistently improves the visual quality and interactive fidelity of generated gaming worlds, as well as the quality of the resulting games after engine adaptation.
Figures & tables
Figure 1: Overview of the gaming world scenarios generated by Code2Games.
Figure 2: Complexity of gaming worlds generated by Code2Games. The horizontal axis reports the number of scene entities, whereas the vertical axis reports the number of executable behaviors; both axes use logarithmic scales. Bubble size denotes the number of gameplay elements, and color indicates the game category.
Figure 3: Overview of Code2Games, which transforms a game intent and its paired base world into a gaming world and adapts it to a game engine.
Method
Visual Quality
Interactive Fidelity
IQ ↑
Dynamic ↑
Motion ↑
Setting ↑
Inter. ↑
Physics ↑
OpenGame + Qwen3.8-Max
62.8
80.5
94.1
78.5
65.0
67.0
AutoUE + Qwen3.8-Max
63.1
82.2
93.8
80.3
65.8
68.7
Claude Code + Claude Opus 5
63.6
81.8
95.1
80.0
66.0
68.4
Codex + GPT-5.6 Sol
63.3
83.2
94.9
80.7
67.8
69.2
Qwen Code + Qwen3.8-Max
61.7
77.8
93.5
75.2
60.8
64.8
Table 1: Results on the GameCode4D Benchmark using the same ten prompts. The top panel reports visual quality and interactive fidelity, while the bottom panel reports multimodal artifact quality and playable-game quality. Bold values indicate the best results in each column.
Figure 4: Qualitative results of gaming worlds generated by Code2Games.
Figure 5: Visual quality and interactive fidelity on the GameCode4D benchmark.
Variant
Setting ↑
Inter. ↑
Physics ↑
M ↑
D ↑
V ↑
w/o Scene Representation
78.2
71.5
72.8
60.5
49.8
55.0
w/o Gameplay Specification
82.1
69.8
73.5
62.0
51.5
56.2
w/o Scene Constraints
83.4
74.1
66.2
64.5
53.0
57.8
Full Code2Games
87.5
75.8
73.9
68.8
56.7
61.0
Table 2: Ablation of Scene-grounding.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
M ↑
D ↑
V ↑
A ↑
Unconstrained Realization
58.0
45.0
52.0
45.0
+ Region Compatibility
62.0
50.0
55.5
50.0
+ Geometric Feasibility Checks
65.0
53.0
58.0
53.0
+ Full Scene-Constrained Selection
68.8
56.7
61.0
57.0
Appendix
Table 4: Fine-grained ablation of scene-constrained game realization.
ID
Prompt Content
Game Scenario
Core Interactions
01
Create a first-person sci-fi assault on a red desert ridgeline. Eliminate three hostile sentries and secure the remote signal relay.
FPS
Shoot; survive; secure relay
02
Create a first-person sci-fi combat sweep on a red-desert ridgeline. Track and eliminate three hostile sentries, then secure the remote signal uplink.
FPS
Track; eliminate; secure uplink
03
Create a third-person monster hunt in a dark forest. Track the Thornback, break its neck armor, defeat it, and claim the trophy.
Monster hunt
Track; break armor; claim trophy
04
Create a storm-weather Thornback hunt through a wet forest. Track the creature in low visibility, defeat it, and retrieve the trophy.
Monster hunt
Track; fight; retrieve trophy
05
Create a high-speed off-road rally through a desert oasis. Drive across uneven terrain, use nitro, pass eight checkpoints, and reach the finish.
Racing
Drive; boost; pass checkpoints
06
Create a desert-storm rally across layered dunes. Control the vehicle over changing slopes, pass eight rally gates, and complete the stage.
Racing
Steer; boost; pass gates
Appendix
Table 5: Game-intent prompts for the qualitative examples. IDs match the qualitative figure sequence.
Dimension
Rating statement
Core mechanics ( M )
The central player action and its consequence are clear and functional.
Content depth ( D )
The world, objectives, and interactions provide meaningful progression.
Functional visuals ( V )
Visual assets, interface, and feedback support readable gameplay.
Art and presentation ( A )
The game appears coherent, complete, and engaging as a demonstration.
Appendix
Table 6: User-study rubric. Each dimension is scored from 0 to 100, using the same playable-game criteria as the automatic evaluation.
Figure 6: FPS gameplay example 1.
Figure 7: FPS gameplay example 2.
Figure 8: Monster-hunt gameplay example 1.
Figure 9: Monster-hunt gameplay example 2.
Figure 10: Racing gameplay example 1.
Figure 11: Racing gameplay example 2.
Figure 12: Skiing gameplay example 1.
Figure 13: Skiing gameplay example 2.
Figure 14: Temple-run gameplay example 1.
Figure 15: Temple-run gameplay example 2.
Figure 16: Third-person shooter gameplay example 1.
Figure 17: Third-person shooter gameplay example 2.
Figure 18: Underwater exploration gameplay example 1.
Figure 19: Underwater exploration gameplay example 2.
Figure 20: Wingsuit gameplay example 1.
Figure 21: Wingsuit gameplay example 2.
Figure 22: Assassin’s Creed-style gameplay example 1.
Figure 23: Assassin’s Creed-style gameplay example 2.
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.
Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22
1The Chinese University of Hong Kong, Shenzhen · 2Shenzhen Loop Area Institute · 4USTB +5
Large Language Models (LLMs) have shown great ability in generating executable code from natural language, opening the possibility of automatically constructing environments for AI agents. Recent work on Code World Models (CWMs) demonstrates that LLMs can translate game rules into Python implementations compatible with solvers like Monte Carlo Tree Search. We study this problem in game settings, where generated environments must implement rules, legal actions, state transitions, observations, and rewards. We refer to these game-specific executable models as Game Code World Models (GameCWMs). However, current approaches to generating code world models rely on frontier models and inference-time refinement loops, limiting accessibility and scalability. This work investigates whether GameCWM generation capabilities can be distilled into smaller models through post-training. We introduce: (1) a curated dataset of 30 games spanning perfect and imperfect information games, (2) a verification framework that evaluates generated code against structural and semantic game properties, and (3) a post-training pipeline combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR). We experiment with Qwen2.5-3B-Instruct and find that SFT can increase syntactic correctness, while RLVR can improve execution-level adherence to game rules, thereby improving Qwen's ability to generate valid GameCWMs in both perfect and imperfect information games. Overall, our pipeline makes Qwen2.5-3B-Instruct more capable of generating valid GameCWMs, thereby offering a scalable path toward automatic environment generation from natural language.
Generating a game is not the same as making one that can be played. Despite advances in code generation, existing approaches treat game generation as one-shot translation from prompt to artifact, leaving interaction-level failures undetected. We argue that evaluating and improving game generation requires a player, and study two roles for graphical user interface (GUI) agents in this process: (1) as an objective evaluator, for which we introduce PlaytestArena, a new evaluation environment that pairs 200 browser-based game generation tasks across eight genres with rubrics of expected in-play behaviors, adjudicated by a GUI agent that loads each build in a browser and plays it; and (2) as a subjective playtester, for which we propose Play2Code, where a game agent and a GUI agent operate in a sustained loop with shared memory, turning game generation into a dialogue between coding and playing. Our experiments show that even frontier models struggle to generate playable games directly, while Play2Code achieves a 66.8% rubric pass-rate, improving over single-pass and agentic-coding baselines by 37.1 and 14.6 points respectively. Further analysis shows that GUI playtester feedback is more traceable than a human report, yet idiosyncratic in ways reminiscent of human testers, establishing game playtesting as a critical testbed for interactive code generation. Our project website is available at https://continual-game-generation.vercel.app/.
Yixu Huang, Bo Li, Na Li +8
αFudan University · βXiaohongshu Inc. · γTongji University +2