Our research investigates how two adaptive AI methods, evolutionary transfer learning and TD(lambda), perform in the three-dimensional chess environment Dragonchess. The game challenges players with its unique board structure and computational load, making it an ideal setting to study how adaptive methods can update evaluation heuristics in novel environments. In this work we re-implement the Dragonchess engine, changing it from a PyGame engine to C++. This enables faster gameplay, allowing us to run 10,000 games with confidence intervals and significance tests, rather than a single small tournament. Both adaptive methods outperform all other agents in the round-robin tournament. Our results showed that there is no significant difference in the performance between the evolved and learned evaluations. This research establishes the efficacy of adaptive methods in structurally complex, novel game domains.
Figures & tables
Figure 1: Each layer of the board (Sky, Ground, and Underworld) consists of 8 rows and 12 columns, represented by integer indices from 0 to 287. At any given index of the array, an integer constant is stored, representing piece type and ownership. These constants are positive for Gold, and negative for Scarlet. This is the same structure as the original Dragonchess implementation [ 10 ] .
Figure 2: Screenshot from our Dragonchess engine’s graphical interface built using C++. Dragonchess’ characteristic three-layered board can be seen with the Sky (top), Land (middle), and Underworld (bottom)—with distinct piece types and clear depiction of multi-layer interactions and movements unique to Dragonchess. This is the same GUI as in the original PyGame implementation [ 10 ] .
Agent
W
L
D
Score
Elo
CMA-ES (evolved, original)
2131
615
1254
0.690
1678.4
TD( λ ) self-play (learned)
2314
1088
598
0.653
1653.1
AlphaBeta- d2 (material search)
1932
1380
688
0.569
1595.2
Jackman (handcrafted)
1889
1374
737
0.564
1592.0
Random
47
3856
97
0.024
981.3
Table 1: Depth-matched (AlphaBeta d=2 ) round-robin standings, 1000 games per pair, color-balanced. Elo by Bradley-Terry MLE (draws as 0.5), mean anchored to 1500.
Figure 3: Bar chart visualization of the Bradley-Terry Elo round-robin standings detailed in Table 1 .
A
B
A rate
95% CI
TD( λ )
CMA-ES
0.503
[0.463, 0.544]
TD( λ )
Jackman
0.601 ∗
[0.569, 0.632]
TD( λ )
AlphaBeta
0.560 ∗
[0.528, 0.591]
TD( λ )
Random
0.982 ∗
[0.972, 0.989]
CMA-ES
Jackman
0.743 ∗
[0.706, 0.777]
CMA-ES
AlphaBeta
0.758 ∗
[0.723, 0.790]
Table 2: Head-to-head decisive win rates (draws excluded) with 95% Wilson confidence intervals. Search depth fixed at 2.
Chess is a two player strategic game that is embedded in classical AI culture as it was once the frontier for intelligent behaviour. There was the silent assumption that the advent of computer engines that play better than the best humans will extinguish interest in the game. However, the opposite has come to pass, with a growing following for the game. A lot of the computational resources are now centered around training of players, where the engine output is just one aspect. Access to past games is also an essential part, both in knowing what games a specific player has played previously, and also which continuations at a certain position have led to victory more often for each of the two colour players. We present Chess_db a suite of logic programming tools that can effectively manipulate games both in memory and via creating back end databases. In particular, we provide versatile code that creates databases from PGN (portable game notation) game files and explore the suitability of open source key-value databases for storing position tables that provide near-instant access to information pertaining to substantially large number of games.
Nicos Angelopoulos, Jan Wielemaker
University College & Imperial College, London UK · University College & Imperial College London, UK · SWI-Prolog solutions +1
Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess. However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the granular quality of gameplay. Nevertheless, incorporating move-by-move information into rating adjustments presents a significant challenge given the substantial noise and the vastness of the game-state space. To address this, we propose the Drift-Diffusion-Enhanced Elo Rating System (DD-Elo), a novel skill assessment framework inspired by the drift diffusion model (DDM) from cognitive neuroscience. By modeling skill expression as a decision-making process, our model integrates move-level data to capture rapid skill fluctuations. We provide a rigorous mathematical derivation proving that DD-Elo maintains a bounded deviation from the traditional Elo system, ensuring theoretical alignment. Extensive experiments demonstrate that DD-Elo adapts to skill changes faster than Elo. Our findings suggest that DD-Elo offers an explainable, highly responsive, and backward-compatible solution for chess rating ecosystems. The implementation code is publicly available at https://github.com/Aquila-zhou1/DD-Elo .
Tianyuan Zhou, Zhizheng Fu, Tianming Yang
School of Intelligence Science and Technology Nanjing University Suzhou, China · Center for Excellence in Brain Science and Intelligence Technology Institute of Neuroscience Chinese Academy of Sciences Shanghai, China
Recent advances in LLM-driven code evolution have enabled automated discovery by iteratively generating and improving programs. However, applying these methods to adversarial multi-agent games introduces a fundamental challenge: the evaluation landscape shifts as strategies improve, causing fixed evaluators to become unreliable and evolution to stagnate. We propose three mechanisms to address this challenge: evaluator co-evolution, which incorporates discovered champions into the opponent pool; hierarchical deep evaluation, which replaces noisy few-game scores with statistically reliable assessments; and weakness pressure, which dynamically up-weights the most difficult opponents to break through plateaus. We implement these mechanisms within FAMOU, a framework built upon the same foundation-model code-evolution paradigm as OpenEvolve and ShinkaEvolve. On the MCTF 2026 3v3 maritime capture-the-flag task, FAMOU consistently outperforms both baselines under two backbone LLMs, achieving the highest combined score (0.526) and the best generalization to unseen opponents (61.7% win rate), while ablations confirm that each mechanism contributes to performance. Notably, the LLM mutation process generates tactical structures entirely absent from the seed strategies -- including lookahead search and adaptive interception -- demonstrating that code-level evolution can produce nontrivial algorithmic innovations in adversarial settings. The FAMOU-evolved strategy further achieved 1st place in the hardware round-robin and 3rd in simulation at the AAMAS 2026 MCTF Competition, validating its real-world transferability. The optimized implementation and corresponding evaluation codes developed through our evolutionary process are available at: https://github.com/1xiangliu1/FAMOU-CoEvo
Haoran Li, Zengle Ge, Ziyang Zhang +10
University of Chinese Academy of Sciences · Famou Agent Team, Baidu AI Cloud · The University of Sydney, Australia +3