Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.
Figures & tables
Figure 1: AAArena overview. Left: the 12 games in ascending rule-description size (abstract syntax tree node count), with player counts, role symmetry, and information structure (Table 1 ). Right: retained-champion Elo for seven models and human SOTA. Each spoke reports native Elo, with game-specific tick intervals; higher Elo lies farther outward. Human SOTA forms the dashed regular dodecagon, with its Elo labeled in red. Model colours match Figure 3 .
Figure 2: Agent for agent iteration loop. An AI agent uses game resources, replays, and retained notes to revise an executable game-agent policy. Validated snapshots enter small matches against selected opponents or full-pool evaluations against the frozen pool. Small matches return dense replays; full-pool evaluations return outcomes, Elo, and rank. The agent retains development history across revisions.
Game
Strategic setting
Pool programs
AST nodes
Rules (RA)
Pacman
Competitive maze collection
44
1,127
219
SnakeGo
Snake movement and territory
141
1,538
308
Rollman
Asymmetric maze pursuit
64
1,539
312
MoneCraft
Mining and resource control
112
1,823
347
AntWar
Tower defence and economy
114
2,160
368
LostSpace
Multiplayer survival and escape
111
2,297
426
Table 1: The 12 games, ordered by increasing abstract syntax tree (AST) node count for their rule representations. RA counts rule atoms, or independently changeable rule propositions. Pool sizes count executable programs, not unique human participants.
Table 2: Main results across 12 adversarial games, ordered by rule complexity. Rule complexity uses AST node counts (Section 3.1 ). A green check marks at least one configuration reaching rank 1; a red cross marks none. Each model entry gives Elo ( ↑ ) and rank ( ↓ ) under a 128-small/16-full budget; a gold medal marks rank 1. Bold marks the best value per row. All models use max reasoning effort. Opus5.5 uses Claude Code; others use Codex.
Figure 3: Per-game Elo of the retained champion. Each game orders model configurations by Elo.
Figure 4: GLM-5.3 on Dorado . Full-pool Elo and rank, small–full interaction order, and token use under the 128/16 budget. Appendix B.5 specifies budget-slot placement, batch markers, rank annotations, and token accounting.
Full-pool eval.
Elo
Rank
Policy adjustment
Code excerpt
2
1782.0
36
Keep mining ; attack only a vulnerable enemy base .
Table 3: Dorado champion milestones. Mechanism descriptions with added (+) and deleted (-) C++ diff excerpts. Ellipses omit context.
128 / 16
384 / 48
Change
Game
Elo
Rank
Elo
Rank
Δ Elo
Δ Rank
Generals
1441.2
17
1441.2
17
+0.0
+0
Miracle
1436.8
26
1486.8
19
+50.0
+7
LOTA
2147.4
22
2147.4
22
+0.0
+0
AntWar2
1977.1
9
2117.0
6
+139.8
+3
Table 4: Threefold-budget continuation for GLM-5.3. Elo and rank of the retained champion at each budget. The extension adds 256 small units and 32 full-pool evaluations to the same run. A positive Δ Rank indicates an improved position.
Figure 5: Policy performance under a threefold budget. Full-pool Elo, interaction order, and token use for four continued runs. Ranks appear at every third full-pool evaluation and local Elo peaks. Plot conventions appear in Appendix B.5 .
Figure 6: Opponent selection. GLM-5.3 retained-champion Elo and rank with model, ladder, random, and top-four selection; all use dense feedback and batch size 4. Plot conventions appear in Appendix B.5 .
Figure 7: Replay feedback. Dense replays versus binary win/non-win feedback under the same 128/16 budget. Labels give Elo and rank. Dense feedback improves performance in all three representative games.
Figure 8: Small-match batch size. Number of opponents per submission, with a fixed total allowance of 128 opponent tests and 16 full-pool evaluations. Labels give Elo and rank. The best size is 4, 2, and 4 for Pacman , AntWar , and Miracle , respectively.
Figure 9: On-policy and off-policy learning trajectories. Full-pool Elo and ranks, cumulative tokens, and token increments for GLM-5.3 with self-generated and other players’ replays. Both conditions receive full-pool feedback on their own policies. Appendix B.5 specifies rank rows, reference levels, and token accounting.
Figure 10: On-policy versus off-policy replay learning. Retained-champion Elo and rank for GLM-5.3 under the same 128/16 budget ceilings. On-policy scores are the main-table results. Labels give Elo and rank; each panel uses the same zero-based Elo scale. Both Pacman runs terminate early upon reaching rank 1.
Work
AI output
Opponents
Feedback
Protocol
Scores
Traces
Direct play
GameBench ( Costarelli et al., 2024 )
Actions
Human + AI players
✓
✓
Direct play
TextArena (competitive) ( Guertler et al., 2025 )
Actions
Human + AI players
✓
✓
Direct play
Program development
GACL ( Game Agent Coding League, n.d. )
Programs
AI-developed programs
–
–
One-shot
Table 5: Representative game-agent evaluation protocols. Feedback denotes information available to the evaluated AI agent: ✓ indicates availability and a dash indicates absence in the named setting. Traces include in-game observations or post-match logs. CodeClash refers to its standard, closed-code tournament; GACL to code generation.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Success. Full-pool Elo and ranks, interaction order, and token use for runs that reach rank 1 and stop early. Plot conventions appear in Appendix B.5 .
Figure 12: Progress. Full-pool Elo and ranks, interaction order, and token use for games with fluctuating improvement. Plot conventions appear in Appendix B.5 .
Figure 13: Stagnation. Full-pool Elo and ranks, interaction order, and token use for games that plateau after initial adaptation. Plot conventions appear in Appendix B.5 .
Game
Model
Ladder
Random
Top-4
Pacman
2096.8 / 4
2127.5 / 3
2014.4 / 6
1897.0 / 10
AntWar
1180.5 / 8
1311.7 / 6
1154.4 / 9
1056.2 / 11
Miracle
1482.4 / 20
1448.8 / 24
1457.0 / 24
1406.0 / 31
Appendix
Table 6: Opponent selection ablation with GLM-5.3. Each cell is Elo / rank under the 128/16 budget.
Game
Dense
Binary
Pacman
2236.1 / 1
1710.1 / 11
AntWar
1118.7 / 9
880.6 / 22
Miracle
1528.2 / 15
1461.1 / 23
Appendix
Table 7: Replay feedback ablation with GLM-5.3. Each cell is Elo / rank under the 128/16 budget.
Game
1
2
4
8
Pacman
2040.6 / 5
2196.3 / 2
2703.9 / 1
2160.4 / 2
AntWar
1130.2 / 9
1433.4 / 5
1208.7 / 8
1118.7 / 9
Miracle
1338.7 / 45
1491.2 / 19
1518.7 / 16
1465.3 / 23
Appendix
Table 8: Small-match batch size ablation with GLM-5.3. Each cell is Elo / rank under the 128/16 budget.
Game
Opus5.5
GPT6-sol
GLM-5.3
Kimi K3
DeepSeek V4 Pro
Qwen 3.8
LongCat 2.0
Rollman
48.75
25.28
23.82
5.50
6.11
14.68
18.35
Pacman
22.16
16.14
33.14
40.17
91.95
163.12
17.52
AntWar
74.19
16.08
142.94
54.59
102.80
284.29
37.56
AquaWar
76.95
26.37
68.68
1.49
60.11
13.92
20.65
Generals
75.44
19.66
137.23
88.88
172.70
60.95
22.30
LostSpace
65.37
21.58
149.86
72.03
141.51
190.16
100.95
Appendix
Table 9: Token use of the main-table runs (millions, input plus output; we count cached input once within input). All runs share the 128/16 interaction budget.
Table 10: Algorithm families in the champion policies. Each bar splits the decision code of one retained champion by algorithm family, weighted by non-comment lines of code. We exclude infrastructure (I/O, parsing, protocol). Row and column means weight each policy equally.