A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?
Organizations: KRAFTON · KAIST · Korea National University of Arts
Abstract
Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at https://a2z-gamespec-bench.github.io.
Figures & tables
| Benchmark Setup | Requirement Evaluation | Evaluation Channels | ||||||||
| Benchmark | Full-Game Generation | Input Type | # of Spec. Tokens | Evaluation Target | Relations | Predefined | Cross-Axis | Source Code | Rendered Behavior | Adaptive Play |
| GameDevBench | ✗ | Task + project | 185 | Task-specific tests | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ |
| GameEngineBench | ✗ | Spec. + project | – | Tests + LLM judge | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ |
| OpenGame-Bench | ✓ | Game spec. | 1,830 | Build / visual / intent | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ |
| WebGameBench | ✓ | Structured spec. | – | Runtime quality | ✗ | ✓ | ✗ | ✗ | ✓ | ✓ |
| GameCraft-Bench | ✓ | Game spec. | 1,547 | Predefined rubric | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ |
| Model | Overall GDD Fidelity | Evaluation Axis | Verifiable Rate | ||||||||
| Source | Replay | Playtest | |||||||||
| All | Small | Big | Small | Big | Small | Big | Small | Big | Small | Big | |
| Claude-Fable-5.1 | 77.0 | 82.8 | 71.1 | 89.7 | 83.5 | 73.5 | 55.1 | 85.2 | 74.7 | 98.7% | 98.7% |
| Claude-Opus-5 | 73.9 | 80.9 | 66.8 | 87.9 | 78.4 | 71.2 | 54.4 | 83.6 | 67.7 | 100.0% | 99.3 % |
| Claude-Opus-4.8 | 56.8 | 67.9 | 45.6 | 72.5 | 49.6 | 62.5 | 41.4 | 68.7 | 46.0 | 99.3% | 100.0% |
| GPT-6-Astra | 71.6 | 78.5 | 64.6 | 84.2 | 74.4 | 69.9 | 52.1 | 81.6 | 67.4 | 99.3% | 97.3% |
Appendix figures & tables46 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
| CVs used in the search ( ) | 13 |
| Candidate harnesses per iteration ( ) | 3 |
| Repeats per quality judgment / pairwise comparison | 3 / 3 |
| Proposer and generator | Claude-Opus-4.8 |
| Evaluator (scoring and pairwise judging) | GPT-5.5 |
| Maximum iterations | 20 |
| Criterion | Requirement | Failure if violated | Example violation |
| G1 Structural Conformance | • Standalone, all sections present • No empty or [TBD] section • No dangling identifier | The requirement cannot be checked. | A section cited elsewhere had been removed. |
| G2 Referential Consistency | • One canonical definition ( DBT- ) • Every citation resolves • Repeated values agree | Contradictions; no implementation can satisfy all. | bool[56] declared for 55 species; 1 wave in App. A vs. 3 in App. B. |
| G3 Behavioral Soundness | • Total, deterministic rule tables • No unreachable or dead mechanic • Difficulty never decreases | Faithful implementations are penalized. | A 90 s boss timer never fires; HP decay kills the patient at 40 s. |
| G4 Executable Formulas | • Variables bound to constants • Fixed evaluation order, bounds • Worked example that recomputes | Agent and evaluator can compute different answers. | Kill time ignores the boss’s DEF; ceil(50/4) written as 12. |
| Statistic | Small | Big |
| Documents | 50 | 50 |
| Lines (mean) | 629 | 1,153 |
| Tokens (total) | 704,270 | 1,314,862 |
| Tokens (mean, range) | 14,085 (5,240–24,022) | 26,297 (12,489–51,433) |
| Outcome requirements (total) | 2,710 | 4,202 |
| Outcome requirements (mean, range) | 54.2 (9–151) | 84.0 (10–246) |
| # | Task Family | Games | |
| 1 | Reaction, Inhibition & Speed Classification | 9 | reaction_test , no_response_test , bloom_or_weed , orbit , bin_bit , odd_even_flash , odd_patch , same_or_shift , three_letter_delivery |
| 2 | Temporal Precision & Rhythm | 7 | beat_bistro , centerline , chord_snap , half_beat_bell , ooga_ooga_tower , ozone_dash , rooftop_rumble |
| 3 | Continuous & Precision Motor Control | 9 | cursor_pursuit , still_cursor , chalk_escape , stamp_register , pocket_curling , magnet_dock , safe_dial , warm_ascent , wheel_steps |
| 4 | Short-Term Memory & Attentional Tracking | 3 | echo_three , hidden_spot , key_under_cups |
| 5 | Perceptual Estimation & Psychophysics | 4 | blink_count , equal_slice , mobile_balance , whisper_compass |
| 6 | Spatial Reasoning & Mental Transformation | 5 | one_fold_letter , shadow_pair , lighthouse_mirrors , tiny_moving_van , untangled_cords |
| # | Genre | Games | |
| 1 | Platformer & Metroidvania | 7 | afterglow , chromashade , rainwright , ssitgim , turning_keep , neon_spray , the_librarians_hook |
| 2 | Action Roguelite & Survival | 4 | abyssal_chain , seamline , todays_best_spot , afterglow_network |
| 3 | Shooter & Bullet-hell | 3 | golem_workshop , magma_arc , snowfall_draw |
| 4 | Rhythm | 4 | beat_reroute , beatstorm , cloud_cast , rhythm_of_response |
| 5 | Puzzle | 4 | chameleon_slide , fable_knot , fogfall_delivery , u001_toyfit |
| 6 | Strategy & Tactics | 5 | ashen_citadel , gateline_2d , siege_deck_2d , chroma_bastion , xl001_edens_debt |
| Document | Sentences | Causal (%) |
| The Sky Above, the Sky Below | 2,069 | 38.3 |
| Frontier Pharmacist | 1,996 | 42.5 |
| Leisure Suit Larry 5 | 1,701 | 38.3 |
| Love for Sail | 1,551 | 37.5 |
| Shape Up or Slip Out | 1,444 | 44.6 |
| Claw | 1,428 | 29.6 |
| Corpus | Documents | Sentences | Causal (%) |
| Studio GDDs | 10 | 11,741 | 37.9 |
| Game benchmark specifications | 246 | 11,441 | 32.0 |
| Board game rulebooks | 97 | 50,518 | 45.3 |
| Pooled | 353 | 73,700 | 42.0 |
| Macro average over corpora | – | – | 38.4 |
| Agreement Measure | Contract |
| Matched Items | 100% |
| Dependency Graph | |
| Trigger Name | 0.94 |
| Predicate Structure | 0.70 |
| Game & Rule | GDD statement (abridged) | Generation | Generation | Ratio of ( ) |
| untangled_cords R8 ( Small ) | All other frames: keep state, keep updating t in Play_Active | timer INCREMENT dt | timer SET advanced by dt | 1.00 / 1.00 |
| magma_arc STAR2 ( Big ) | Not STAR3, and hit events this attempt: 2 stars | event <= 3 STAR SET 2 | event IN [1, 2, 3] STAR SET 2 | 1.00 / 1.00 |
| Split | games | rules | Predicate structure | pairs | |
| Small | 50 | 869 | identical | 1,773 | 12.7% |
| different | 834 | 17.7% | |||
| Big | 50 | 1,860 | identical | 4,045 | 8.8% |
| different | 1,535 | 11.1% |
| Judge, effort | Item | Mean reward | vs. ref. | s / call | $ / call |
| sol, xhigh (reference) | – | 0.639 | – | 78 | 0.048 |
| sol, high | 0.88 | 0.667 | 71 | 0.034 | |
| sol, medium | 0.88 | 0.669 | 63 | 0.034 | |
| sol, low | 0.89 | 0.680 | 52 | 0.045 | |
| luna, xhigh | 0.85 | 0.634 | 63 | 0.0074 | |
| luna, high | 0.85 | 0.645 | 41 | 0.0072 |
| Coding agent | Normal (%) | Adversarial gain (pp) | Combined (%) | ||||
| Sat. | Viol. | Cov. | Sat. | Viol. | Score | Cov. | |
| All | |||||||
| Claude-Fable-5.1 | 28.6 | 0.4 | 29.0 | 51.4 | 7.6 | 80.0 | 88.0 |
| Claude-Opus-5 | 26.7 | 0.4 | 27.1 | 49.0 | 9.0 | 75.7 | 85.1 |
| Claude-Opus-4.8 | 19.2 | 0.6 | 19.7 | 38.2 | 13.6 | 57.4 | 71.5 |
| GPT-6-Astra | 25.8 | 0.5 | 26.3 | 48.7 | 8.5 | 74.5 | 83.4 |
| First pass | Second pass | Coverage | |
| Normal | – | 40.5 | – |
| Normal | Normal | 41.0 | |
| Normal | Adversarial | 48.5 | |
| Adversarial | – | 45.0 | – |
| Agent | Equal | Source emphasis | Replay emphasis | Adaptive emphasis |
| Claude-Fable-5.1 | 77.0 (1) | 79.4 (1) | 73.8 (1) | 77.7 (1) |
| Claude-Opus-5 | 73.9 (2) | 76.2 (2) | 71.1 (2) | 74.3 (2) |
| GPT-6-Astra | 71.6 (3) | 73.5 (3) | 68.9 (3) | 72.3 (3) |
| GPT-5.5 | 63.2 (4) | 64.5 (4) | 60.3 (5) | 64.8 (4) |
| GPT-5.6-Sol | 61.8 (5) | 62.0 (5) | 60.4 (4) | 63.1 (5) |
| Measure | Small | Big | All |
| Games | 50 | 50 | 100 |
| Full-credit rules | 738 | 896 | 1,634 |
| Eligible full-credit targets | 377 | 531 | 908 |
| Additional inspection targets | 97 | 311 | 408 |
| Games with additional targets | 17 | 36 | 53 |
| Additional targets / eligible targets | 25.7% | 58.6% | 44.9% |
| Judge | Effort | Rule-Level Correlation | Invariant Agreement | Mean Rule Score | Tokens/Call | Relative Cost |
| GPT-5.6-Sol | Extra high | 0.74 | 0.83 | 0.76 | 213 | 1.00 |
| High | 0.70 | 0.83 | 0.76 | 214 | 1.00 | |
| Medium | 0.75 | 0.88 | 0.78 | 176 | 0.83 | |
| Low | 0.80 | 0.86 | 0.79 | 160 | 0.75 | |
| GPT-5.6-Luna | Extra high | 0.72 | 0.77 | 0.79 | 198 | 0.15 |
| High | 0.72 | 0.76 | 0.79 | 187 | 0.15 |
| Native | 640 px | 480 px | |
| 1920 1080 | |||
| 1280 720 |
| Judge, effort | Small | Big | Tokens / vote |
| Sol, xhigh | 0.959 / 2 | 0.930 / 4 | 485–753K |
| Sol, high (re-run) | 0.959 / 3 | 0.959 / 1 | 516–688K |
| Sol, low | 0.892 / 7 | 0.897 / 3 | 217–315K |
| Luna, xhigh | 0.919 / 3 | 0.893 / 1 | 57–86K |
| Luna, high | 0.910 / 6 | 0.888 / 3 | 51–71K |
| Luna, low | 0.833 / 9 | 0.810 / 10 | 46–64K |
| Audit Outcome | Assessments | Share (%) |
| Fixed-rate preferred | 458 | 11.2 |
| Adaptive preferred | 701 | 17.1 |
| Equivalent | 2,848 | 69.4 |
| Insufficient evidence | 95 | 2.3 |
| Total | 4,102 | 100.0 |
| Gameplay Agent | Adaptive Playtest | Judgment Coverage (%) | ||
| Game | Progress (%) | Requirements | Normal | Normal + Adv. |
| Abyssal Chain | 110 | 6.4 | 89.1 | |
| Alias Alchemy Shop | 133 | 12.0 | 91.0 | |
| Beat Reroute | 84 | 19.1 | 86.9 | |
| Chameleon Slide | 87 | 0.0 | 67.8 | |
| Fogfall Delivery | 119 | 5.0 | 88.2 | |
| Game | Source | Replay | Playtest |
| Sonic: Cascade Coast | 99.71 | 53.50 | 80.00 |
| Diablo Cathedral | 86.74 | 55.00 | 66.55 |
| Rocket League | 83.64 | 57.40 | 77.54 |