Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.
Figures & tables
Figure 1: AAArena overview. Left: the 12 games in ascending rule-description size (abstract syntax tree node count), with player counts, role symmetry, and information structure (Table 1 ). Right: retained-champion Elo for seven models and human SOTA. Each spoke reports native Elo, with game-specific tick intervals; higher Elo lies farther outward. Human SOTA forms the dashed regular dodecagon, with its Elo labeled in red. Model colours match Figure 3 .
Figure 2: Agent for agent iteration loop. An AI agent uses game resources, replays, and retained notes to revise an executable game-agent policy. Validated snapshots enter small matches against selected opponents or full-pool evaluations against the frozen pool. Small matches return dense replays; full-pool evaluations return outcomes, Elo, and rank. The agent retains development history across revisions.
Game
Strategic setting
Pool programs
AST nodes
Rules (RA)
Pacman
Competitive maze collection
44
1,127
219
SnakeGo
Snake movement and territory
141
1,538
308
Rollman
Asymmetric maze pursuit
64
1,539
312
MoneCraft
Mining and resource control
112
1,823
347
AntWar
Tower defence and economy
114
2,160
368
LostSpace
Multiplayer survival and escape
111
2,297
426
Table 1: The 12 games, ordered by increasing abstract syntax tree (AST) node count for their rule representations. RA counts rule atoms, or independently changeable rule propositions. Pool sizes count executable programs, not unique human participants.
Table 2: Main results across 12 adversarial games, ordered by rule complexity. Rule complexity uses AST node counts (Section 3.1 ). A green check marks at least one configuration reaching rank 1; a red cross marks none. Each model entry gives Elo ( ↑ ) and rank ( ↓ ) under a 128-small/16-full budget; a gold medal marks rank 1. Bold marks the best value per row. All models use max reasoning effort. Opus5.5 uses Claude Code; others use Codex.
Figure 3: Per-game Elo of the retained champion. Each game orders model configurations by Elo.
Figure 4: GLM-5.3 on Dorado . Full-pool Elo and rank, small–full interaction order, and token use under the 128/16 budget. Appendix B.5 specifies budget-slot placement, batch markers, rank annotations, and token accounting.
Full-pool eval.
Elo
Rank
Policy adjustment
Code excerpt
2
1782.0
36
Keep mining ; attack only a vulnerable enemy base .
Table 3: Dorado champion milestones. Mechanism descriptions with added (+) and deleted (-) C++ diff excerpts. Ellipses omit context.
128 / 16
384 / 48
Change
Game
Elo
Rank
Elo
Rank
Δ Elo
Δ Rank
Generals
1441.2
17
1441.2
17
+0.0
+0
Miracle
1436.8
26
1486.8
19
+50.0
+7
LOTA
2147.4
22
2147.4
22
+0.0
+0
AntWar2
1977.1
9
2117.0
6
+139.8
+3
Table 4: Threefold-budget continuation for GLM-5.3. Elo and rank of the retained champion at each budget. The extension adds 256 small units and 32 full-pool evaluations to the same run. A positive Δ Rank indicates an improved position.
Figure 5: Policy performance under a threefold budget. Full-pool Elo, interaction order, and token use for four continued runs. Ranks appear at every third full-pool evaluation and local Elo peaks. Plot conventions appear in Appendix B.5 .
Figure 6: Opponent selection. GLM-5.3 retained-champion Elo and rank with model, ladder, random, and top-four selection; all use dense feedback and batch size 4. Plot conventions appear in Appendix B.5 .
Figure 7: Replay feedback. Dense replays versus binary win/non-win feedback under the same 128/16 budget. Labels give Elo and rank. Dense feedback improves performance in all three representative games.
Figure 8: Small-match batch size. Number of opponents per submission, with a fixed total allowance of 128 opponent tests and 16 full-pool evaluations. Labels give Elo and rank. The best size is 4, 2, and 4 for Pacman , AntWar , and Miracle , respectively.
Figure 9: On-policy and off-policy learning trajectories. Full-pool Elo and ranks, cumulative tokens, and token increments for GLM-5.3 with self-generated and other players’ replays. Both conditions receive full-pool feedback on their own policies. Appendix B.5 specifies rank rows, reference levels, and token accounting.
Figure 10: On-policy versus off-policy replay learning. Retained-champion Elo and rank for GLM-5.3 under the same 128/16 budget ceilings. On-policy scores are the main-table results. Labels give Elo and rank; each panel uses the same zero-based Elo scale. Both Pacman runs terminate early upon reaching rank 1.
Work
AI output
Opponents
Feedback
Protocol
Scores
Traces
Direct play
GameBench ( Costarelli et al., 2024 )
Actions
Human + AI players
✓
✓
Direct play
TextArena (competitive) ( Guertler et al., 2025 )
Actions
Human + AI players
✓
✓
Direct play
Program development
GACL ( Game Agent Coding League, n.d. )
Programs
AI-developed programs
–
–
One-shot
Table 5: Representative game-agent evaluation protocols. Feedback denotes information available to the evaluated AI agent: ✓ indicates availability and a dash indicates absence in the named setting. Traces include in-game observations or post-match logs. CodeClash refers to its standard, closed-code tournament; GACL to code generation.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Success. Full-pool Elo and ranks, interaction order, and token use for runs that reach rank 1 and stop early. Plot conventions appear in Appendix B.5 .
Figure 12: Progress. Full-pool Elo and ranks, interaction order, and token use for games with fluctuating improvement. Plot conventions appear in Appendix B.5 .
Figure 13: Stagnation. Full-pool Elo and ranks, interaction order, and token use for games that plateau after initial adaptation. Plot conventions appear in Appendix B.5 .
Game
Model
Ladder
Random
Top-4
Pacman
2096.8 / 4
2127.5 / 3
2014.4 / 6
1897.0 / 10
AntWar
1180.5 / 8
1311.7 / 6
1154.4 / 9
1056.2 / 11
Miracle
1482.4 / 20
1448.8 / 24
1457.0 / 24
1406.0 / 31
Appendix
Table 6: Opponent selection ablation with GLM-5.3. Each cell is Elo / rank under the 128/16 budget.
Game
Dense
Binary
Pacman
2236.1 / 1
1710.1 / 11
AntWar
1118.7 / 9
880.6 / 22
Miracle
1528.2 / 15
1461.1 / 23
Appendix
Table 7: Replay feedback ablation with GLM-5.3. Each cell is Elo / rank under the 128/16 budget.
Game
1
2
4
8
Pacman
2040.6 / 5
2196.3 / 2
2703.9 / 1
2160.4 / 2
AntWar
1130.2 / 9
1433.4 / 5
1208.7 / 8
1118.7 / 9
Miracle
1338.7 / 45
1491.2 / 19
1518.7 / 16
1465.3 / 23
Appendix
Table 8: Small-match batch size ablation with GLM-5.3. Each cell is Elo / rank under the 128/16 budget.
Game
Opus5.5
GPT6-sol
GLM-5.3
Kimi K3
DeepSeek V4 Pro
Qwen 3.8
LongCat 2.0
Rollman
48.75
25.28
23.82
5.50
6.11
14.68
18.35
Pacman
22.16
16.14
33.14
40.17
91.95
163.12
17.52
AntWar
74.19
16.08
142.94
54.59
102.80
284.29
37.56
AquaWar
76.95
26.37
68.68
1.49
60.11
13.92
20.65
Generals
75.44
19.66
137.23
88.88
172.70
60.95
22.30
LostSpace
65.37
21.58
149.86
72.03
141.51
190.16
100.95
Appendix
Table 9: Token use of the main-table runs (millions, input plus output; we count cached input once within input). All runs share the 128/16 interaction budget.
Table 10: Algorithm families in the champion policies. Each bar splits the decision code of one retained champion by algorithm family, weighted by non-comment lines of code. We exclude infrastructure (I/O, parsing, protocol). Row and column means weight each policy equally.
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire and use new knowledge. Existing benchmarks have sought to evaluate this ability, but they primarily evaluate tasks whose rules are provided in the instructions or already familiar to pretrained models, making it difficult to distinguish learning from interactions from reasoning with existing knowledge. To address this, we introduce Learn2Play Bench, a benchmark of newly designed text-based games, whose rules are novel or counterintuitive, requiring agents to acquire knowledge through interaction rather than rely solely on pretrained knowledge. These games provide reproducible feedback and automatic scoring, enabling controlled evaluation of learning across repeated attempts. We also vary game instances to test whether agents can apply what they have learned to new situations. Therefore, we evaluate how backbone models, self-evolving methods, and agent harnesses affect agents' learning ability, revealing three findings: (1) Experience retention: Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies. (2) Human agent gap: Top-performing human players achieve higher peak scores than the evaluated agents. Human explore more varied strategies, and repeat actions less. (3) Harness matters: With the backbone fixed, changing the harness can improve performance while reducing estimated inference cost. Together, these findings provide insights into how LLM agents learn from experience and suggest directions for future work to improve their learning ability. Project website: https://liushiliushi.github.io/learn2play-bench-website/
This paper introduces a new paradigm for AI game programming, leveraging large language models (LLMs) to extend and operationalize Claude Shannon's taxonomy of game-playing machines. Central to this paradigm is Nemobot, an interactive agentic engineering environment that enables users to create, customize, and deploy LLM-powered game agents while actively engaging with AI-driven strategies. The LLM-based chatbot, integrated within Nemobot, demonstrates its capabilities across four distinct classes of games. For dictionary-based games, it compresses state-action mappings into efficient, generalized models for rapid adaptability. In rigorously solvable games, it employs mathematical reasoning to compute optimal strategies and generates human-readable explanations for its decisions. For heuristic-based games, it synthesizes strategies by combining insights from classical minimax algorithms (see, e.g., shannon1950chess) with crowd-sourced data. Finally, in learning-based games, it utilizes reinforcement learning with human feedback and self-critique to iteratively refine strategies through trial-and-error and imitation learning. Nemobot amplifies this framework by offering a programmable environment where users can experiment with tool-augmented generation and fine-tuning of strategic game agents. From strategic games to role-playing games, Nemobot demonstrates how AI agents can achieve a form of self-programming by integrating crowdsourced learning and human creativity to iteratively refine their own logic. This represents a step toward the long-term goal of self-programming AI.
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short string and has unique transformation rules that must be discovered through interaction and experimentation. The levels of the game present a series of challenges to test whether the rules have been discovered, where the win conditions for each level are also unknown. We provide games at seven tiers of difficulty for AI agents. The lowest tier is routinely solvable by multiple models, while the highest tier challenges the best models in agentic harnesses. All 70 games were solved by at least one human on first attempt. A subset of 21 games is released publicly, and the remainder is held private for secure evaluation.
Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13
1Thinking About Thinking · 2Independent · University of Oxford +5