RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
Organizations: University of California San Diego · ByteDance Inc. · Carnegie Mellon University
Abstract
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
Figures & tables
| Method | Mechanics | Depth | Visuals | Art | Overall | Tok. | Cost |
| Godot ( Godot Engine contributors, 2026 ) | |||||||
| Generator: Codex GPT-5.5 (high) | |||||||
| Base (frozen ) | 58.9 | 51.7 | 51.4 | 44.6 | 50.26 | 0.26M | – |
| + Play2Code | 59.2 | 51.3 | 52.2 | 45.8 | 50.74 | 3.20M | $1.06 |
| + RSIGame | 72.8 | 60.3 | 65.9 | 64.5 | 64.53 | 3.12M | $0.88 |
| Generator: Codex Kimi-K2.6 | |||||||
Appendix figures & tables38 assets
Supplementary material from the paper’s appendix.
Appendix
| Item | Contents / specification | Location |
| Released artifacts | ||
| Implementation | Local/global loops, Global Quality Monitor, director-review tool, and scoring harness | Code, live |
| Prompts and output contracts | Complete prompts and structured outputs used by all agents | Code, live; App. C.4 |
| Playable evolution | Six representative games with retained checkpoints | Project page, live |
| Frozen base projects | Starting project for each task and generator | On release |
| Evaluation artifacts | Per-task rubric scores, judge outputs, and replay reports (51,644 files) | HF, live † |
| Human director — 60/60 actions, 0 resets. | |
|---|---|
| Objective | The art needs work: the title and game-over screens are cluttered and have no proper background, the character art is ugly and should be regenerated in a comic style, and attacks have no effects. The game also needs to teach how it is played — right now I do not know what to press. |
| Why now | The characters move and there are title and game-over screens, but it is still too rough to be appealing. |
| Model director (Claude Opus 5) — 25/60 actions, 0 resets. | |
| Objective | Turn the match from a scripted beatdown into a real two-sided fight: player inputs must visibly and reliably land, block and escape, against an AI that can be beaten, with two fighters that genuinely look and play differently. |
| Why now | Across two full matches (one as each fighter) the AI never lost a pixel of health. Every round ended in about 18 seconds with me pinned in the left corner in a hit, knockdown, hit loop. Light, heavy, the special input, holding back, crouching and jumping all failed to trade, block or get me out. Nobody can learn to play when nothing they press matters. […] Seven rounds of visual polish have gone into the frame around the fight (title, background, banners, results screen), but the fight itself is still missing. |
| Stage | Module | Output contract | Consumed by |
|---|---|---|---|
| Direction Decision | Controller | One development_question ; action in { reuse , explore , escalate }; an exploration_plan naming the cheapest mode that can answer it. | Sets the round’s objective; escalate reports the direction exhausted rather than inventing work. |
| Agentic Exploration | Explorer | The 1–3 requirements a session should settle, the scenario to boot, and one actionable instruction for the play agent. | Drives the test agent’s interaction with , producing the trace . |
| Evidence-Grounded Editing | Editor | Edits to the project under a build constraint, plus a report of what was changed. | Produces . |
| Agentic Verification | Verifier | goal_achieved in {yes, no, unclear}; broken_now ; next_goal when that list is non-empty. | Accepts or rejects the round; anything broken re-enters the state as pending work. |
| Global Quality Monitoring | Global Quality Monitor | Per-criterion readings for both builds, changes with direction and magnitude, and regressions — never a score. | Aggregated into the proxy value that decides the champion (Appendix C.2 ). |
| Role | Model | Served by | Decoding |
| Generator | one per row group of | harness | harness |
| Table 1 | default | default | |
| Controller | GLM-5.3-Flash | OpenRouter | , top- 1 |
| Explorer | Qwen3.8-27B | local vLLM (TP 2) | , top- 1 |
| Verifier | GLM-5.3-Flash | OpenRouter | , top- 1 |
| Editor | GLM-5.3-Flash | OpenRouter | , top- 1 |
| Parameter | Value | |
| Round | Tool calls the Editor may make per round | 26 |
| Wall-clock budget per improvement round | 1800 s | |
| Improvement attempts per round before the round is given up | 3 | |
| Verification samples per claimed fix | 3 | |
| Probe steps taken before an edit is proposed | 8 | |
| Run | Generated assets per run | 20 |
| Model | Provider | Input | Cache read | Output |
|---|---|---|---|---|
| GPT-5.5 | OpenAI | 5.000 | 0.500 | 30.00 |
| Kimi-K2.6 | Baidu | 0.408 | 0.069 | 0 1.72 |
| Qwen3.8-27B | DeepInfra | 0.150 | 0.037 | 0 1.88 |
| GLM-5.3-Flash | DeepInfra | 0.075 | 0.015 | 0 0.25 |
| Engine | Generator | Billable | Total | Cache | Output | Cost |
| Godot | Codex GPT-5.5 (high) | 0.26M | 0 3.83M | 94% | 0 28k | $ 0 3.77 |
| Godot | GLM-5.3-Flash | 0.44M | 11.43M | 96% | 0 36k | $ 0 0.20 |
| Godot | Kimi-K2.6 | 3.22M | 0 8.98M | 65% | 123k | $ 0 1.87 |
| Godot | Qwen3.8-27B | 6.41M | 11.62M | 46% | 179k | $ 0 1.47 |
| Godot | Qwen3.8-27B (SFT) | 0.57M | 0 8.21M | 95% | 140k | – |
| Phaser | OpenGame GPT-5.5 | 5.36M | 0 9.21M | 42% | 0 97k | $31.17 |
| Subset | p25 | Median | p75 | Identical 3/3 | |
|---|---|---|---|---|---|
| All artifacts | 529 | 0.00 | 0.73 | 3.33 | 225 |
| Base | 266 | 0.00 | 0.36 | 2.33 | 131 |
| Improved | 263 | 0.00 | 0.87 | 3.50 | 0 94 |
| Family | Median | Family | Median | ||
|---|---|---|---|---|---|
| sports | 14 | 4.66 | shooter | 24 | 0.44 |
| horror | 19 | 3.37 | puzzle | 32 | 0.42 |
| openworld | 62 | 1.41 | simulation | 24 | 0.28 |
| visualnovel | 32 | 1.17 | strategy | 63 | 0.00 |
| platformer | 77 | 0.87 | rhythm | 21 | 0.00 |
| tycoon | 58 | 0.73 | idle | 16 | 0.00 |
| Overall | |||||
| Method | Qwen3.8-27B | GPT-5.5 | |||
| Base (frozen ) | 49.07 | 50.34 | |||
| Play2Code | 49.63 | 50.62 | |||
| RSIGame | 63.10 | 62.38 | |||
| Comparison | Qwen | GPT-5.5 | 95% CI Qwen | 95% CI GPT-5.5 | Agreement |
| RSIGame Base | 37/40 | ||||
| Comparison | W/L/T | Agreement | |
|---|---|---|---|
| RSIGame – Base | 17/3/0 | 0.003 | 45/59 |
| RSIGame – Play2Code | 18/2/0 | ||
| Play2Code – Base | 9/11/0 | 0.824 |
| Comparison | 95% CI | W/L | ||
|---|---|---|---|---|
| Godot | ||||
| GPT-5.5: RSIGame – Base | +14.26 | [11.49, 17.06] | 112/16 | |
| GPT-5.5: RSIGame – Play2Code | +13.79 | [10.89, 16.65] | 118/19 | |
| Qwen3.8-27B: RSIGame – Base | +10.70 | [7.75, 13.77] | 100/31 | |
| Qwen3.8-27B: RSIGame – Play2Code | +7.24 | [4.49, 10.04] | 92/41 | |
| Phaser | ||||
| Stage | Unit | Teacher | Godot | Phaser | Total |
| S1 generation | trajectory | GPT-5.5 | 1,929 | 284 | 2,213 |
| S1 generation | decision row | GPT-5.5 | 66,312 | 28,097 | 94,409 |
| S2 planning | plan | distilled | 1,979 | 129 | 2,108 |
| S3 improvement | round | GLM-5.3-Flash | 3,730 | 273 | 4,003 |
| of which verified | round | — | 1,869 | 144 | 2,013 |
| S3 improvement agent | decision row | GLM-5.3-Flash | 111,302 | 9,930 | 121,232 |
| Stage added | Unit | Godot | Phaser | Rows added | Cumulative |
| S1 generation | agent trajectory | 876 | 418 | 1,294 | 1,294 |
| S2 planning | brief file plan | 216 | 0 92 | 0, 308 | 1,602 |
| S3 improvement | improvement round | 509 | 00 0 | 0, 509 | 2,111 |
| Total | 1,601 | 510 | 2,111 | ||
| Configuration | Value |
|---|---|
| Base model | Qwen3.8-27B |
| Adaptation | LoRA ( Hu et al., 2022 ) |
| LoRA rank | 16 |
| LoRA | 32 |
| LoRA dropout | 0.05 |
| LoRA targets | All linear projections |
| Variant | Mechanics | Depth | Visuals | Art | Overall |
|---|---|---|---|---|---|
| Base | 41.2 | 33.4 | 38.6 | 38.3 | 37.07 |
| +Gen | 49.0 | 41.4 | 42.0 | 38.0 | 41.45 |
| +Gen+Plan | 48.5 | 40.8 | 42.6 | 39.5 | 41.77 |
| +Gen+Plan+Improve | 56.1 | 47.5 | 49.8 | 44.8 | 48.22 |
| the 12 shared tasks | VibeGame’s own 40 | |
|---|---|---|
| VibeGame | 23.3 | 30.5 |
| OpenGame GPT-5.5, one shot | 50.1 | 46.0 |
| Play2Code | 56.4 | — |
| RSIGame | 62.4 | — |
| Family | Base | +P2C | + RSIGame | |
|---|---|---|---|---|
| idle | 4 | 81.9 | 86.8 | 87.5 |
| sports | 4 | 72.8 | 70.5 | 81.2 |
| horror | 5 | 67.2 | 71.4 | 75.1 |
| openworld | 15 | 59.2 | 57.9 | 62.2 |
| tycoon | 16 | 53.3 | 54.4 | 73.2 |
| visualnovel | 11 | 52.0 | 55.8 | 71.5 |
| Family | Base | +P2C | + RSIGame | |
|---|---|---|---|---|
| idle | 4 | 53.5 | 50.1 | 60.0 |
| sports | 4 | 45.9 | 57.3 | 64.4 |
| horror | 5 | 43.1 | 56.6 | 63.4 |
| tycoon | 16 | 36.1 | 40.7 | 46.4 |
| rhythm | 5 | 33.7 | 37.1 | 45.4 |
| shooter | 7 | 33.4 | 35.9 | 39.1 |
| Family | Base | +P2C | + RSIGame | |
|---|---|---|---|---|
| sports | 4 | 49.6 | 61.5 | 72.6 |
| idle | 4 | 39.0 | 59.2 | 68.6 |
| tycoon | 16 | 38.8 | 43.1 | 56.1 |
| racing | 4 | 34.9 | 36.7 | 53.0 |
| horror | 5 | 34.1 | 43.6 | 49.9 |
| visualnovel | 11 | 32.7 | 38.8 | 59.0 |
| Family | Base | +P2C | + RSIGame | |
|---|---|---|---|---|
| horror | 5 | 58.7 | 56.6 | 57.0 |
| idle | 4 | 48.8 | 53.2 | 73.9 |
| cardgame | 5 | 48.3 | 53.1 | 59.3 |
| sports | 4 | 44.8 | 51.2 | 55.6 |
| shooter | 7 | 42.3 | 44.7 | 61.5 |
| visualnovel | 11 | 42.0 | 46.0 | 44.9 |
| Family | Base | Play2Code | RSIGame | |
|---|---|---|---|---|
| idle | 4 | 80.4 | 78.6 | 84.1 |
| sports | 4 | 73.3 | 68.7 | 76.1 |
| horror | 5 | 60.6 | 62.4 | 71.5 |
| tycoon | 16 | 54.7 | 56.4 | 66.3 |
| strategy | 17 | 52.1 | 51.0 | 62.0 |
| visualnovel | 11 | 50.2 | 48.9 | 58.4 |
| Family | Base | +P2C | + RSIGame | |
|---|---|---|---|---|
| idle | 4 | 72.7 | 73.8 | 78.2 |
| cardgame | 5 | 68.9 | 68.9 | 61.8 |
| visualnovel | 11 | 68.0 | 68.6 | 73.5 |
| tycoon | 16 | 59.9 | 62.9 | 74.3 |
| sports | 4 | 57.2 | 59.8 | 64.8 |
| roguelike | 14 | 56.5 | 58.2 | 66.7 |
| Family | Base | +P2C | + RSIGame | |
|---|---|---|---|---|
| sports | 4 | 56.1 | 68.8 | 65.4 |
| horror | 5 | 55.6 | 52.9 | 57.6 |
| simulation | 6 | 55.3 | 53.3 | 67.8 |
| visualnovel | 11 | 53.3 | 54.2 | 52.4 |
| rhythm | 5 | 50.9 | 58.4 | 61.4 |
| tycoon | 16 | 47.5 | 58.9 | 56.7 |
| Family | Base | Play2Code | RSIGame | |
|---|---|---|---|---|
| idle | 4 | 73.0 | 75.2 | 78.5 |
| horror | 5 | 56.0 | 60.7 | 67.5 |
| tycoon | 16 | 53.3 | 62.0 | 66.5 |
| sports | 4 | 51.6 | 49.9 | 56.8 |
| shooter | 7 | 41.8 | 51.4 | 55.4 |
| rhythm | 5 | 46.8 | 61.3 | 66.2 |