Amadeus: When Models of People Meet
Organizations: Queen’s University Belfast United Kingdom
Abstract
With the constant advancements in AI, one possibility is to model agents after humans and, in turn, use these agents to carry out synthetic interactions. Such models could be used to predict interactions between their real counterparts, or potentially interactions at larger scales. In this paper, we test a more controlled version of this question through chess. We use 8 elite chess players, seal their direct pairwise games, learn each player independently using different methods, and then compose the resulting models on the withheld dyads. To evaluate the generated interactions, we use two measurements: opening-family total variation distance and win-draw-loss (WDL) total variation distance. M1 primarily improves WDL fidelity while producing smaller opening-family improvements, whereas M2 produces much larger opening-family improvements while having little effect on WDL-TV. For opening-family behaviour under M2, the correct assignment of the eight learned player identities also gives the closest match among all possible assignments. These results show that at least some properties of previously unseen interactions can be recovered from independently learned individuals. The partial recovery observed here may reflect limitations of the current individual modelling methods rather than a fundamental limit on compositional interaction recovery. An additional post-hoc method that combines the two mechanisms improves both measurements, suggesting that recovery across these behavioural properties is not necessarily mutually exclusive.
Figures & tables
| Method | Rank 1 | Mean NLL | Wrong correct | Positions |
|---|---|---|---|---|
| M1 | 8/8 | 0.006802 | 0.004365 | 47,136 |
| M2 | 8/8 | 0.009278 | 0.020449 | 46,950 |
| Method | Metric | GG | AG | GB | AB | |
|---|---|---|---|---|---|---|
| M1 | WDL-TV | 0.164171 | 0.157753 | 0.156141 | 0.150148 | -0.014024 |
| M1 | Opening-family TV | 0.554433 | 0.547406 | 0.544765 | 0.536973 | -0.017461 |
| M2 | WDL-TV | 0.163166 | 0.162582 | 0.162589 | 0.163408 | 0.000241 |
| M2 | Opening-family TV | 0.554711 | 0.501370 | 0.488343 | 0.437202 | -0.117509 |
| Metric | M1 AB | M2 AB | M3 AB |
|---|---|---|---|
| WDL-TV | 0.150148 | 0.163408 | 0.152036 |
| Opening-family TV | 0.536973 | 0.437202 | 0.444428 |
| Method | Metric | Improved | Mean | Median | Min | Max |
|---|---|---|---|---|---|---|
| M1 | WDL-TV | 22/28 | -0.014024 | -0.010900 | -0.088700 | 0.046300 |
| M1 | Opening-TV | 22/28 | -0.017461 | -0.018883 | -0.058385 | 0.027417 |
| M2 | WDL-TV | 13/28 | 0.000241 | 0.001650 | -0.018300 | 0.015700 |
| M2 | Opening-TV | 27/28 | -0.117509 | -0.105987 | -0.364176 | 0.004373 |
| M3 | WDL-TV | 19/28 | -0.011037 | -0.010500 | -0.064400 | 0.032800 |
| M3 | Opening-TV | 27/28 | -0.109967 | -0.090846 | -0.347976 | 0.006572 |
| Method | Metric | Matched TV | Permutation mean | Rank / 40,320 | Substitutions beaten |
|---|---|---|---|---|---|
| M1 | WDL | 0.1501 | 0.1563 | 143 | 60% |
| M1 | Opening | 0.5370 | 0.5460 | 1,789 | 63% |
| M2 | WDL | 0.1634 | 0.1643 | 8,096 | 52% |
| M2 | Opening | 0.4372 | 0.5800 | 1 | 90% |
| M3 | WDL | 0.1520 | 0.1569 | 1,625 | 53% |
| M3 | Opening | 0.4444 | 0.5886 | 1 | 90% |
| Method | Metric | Cond. | Observed TV | Reference TV |
|---|---|---|---|---|
| M1 | WDL | GG | 0.1642 | 0.1048 |
| M1 | WDL | AB | 0.1501 | 0.1070 |
| M1 | Opening | GG | 0.5544 | 0.3813 |
| M1 | Opening | AB | 0.5370 | 0.3805 |
| M2 | WDL | GG | 0.1632 | 0.1049 |
| M2 | WDL | AB | 0.1634 | 0.1050 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage 1 (population) | Stage 2 (Broadcast) | |
|---|---|---|
| Initialization | Random; PyTorch defaults, with Xavier initialization for Elo/GAB embeddings only | Stage-1 final checkpoint |
| Optimizer steps | 1,000,000 | 200,000 |
| Effective batch size | 512 | 512 |
| Peak learning rate | ||
| Learning-rate floor | ||
| Weight decay |
| M1 | M2 | |
|---|---|---|
| Trained parameters | One 128-d vector per player | Candidate CNN, 32-d style vector per player, scale |
| Action space | Full 4,352, legal-masked | Base top-5; human move appended during train/validation if absent |
| Loss | Cross-entropy | Cross-entropy over candidates |
| Training scope | 8 independent runs | 1 joint run over all players |
| Train/val split | 80/20 by game, independently per player | 80/20 by game with shared seeded split state |
| Phase balancing | Per player, per split | Per player, per split |
| Player | Representative Elo |
|---|---|
| Magnus Carlsen | 2855 |
| Wesley So | 2769 |
| Levon Aronian | 2756 |
| Fabiano Caruana | 2786 |
| Maxime Vachier-Lagrave | 2751 |
| Hikaru Nakamura | 2829 |
| Software | |
| Python / PyTorch | 3.12.3 / 2.6.0 (CUDA 12.4) |
| python-chess / pyarrow | 1.11.2 / 25.0.1 |
| Seeds | |
| Training | Seed 0 where specified by the training pipeline; M1/M2 data splitting and phase balancing also use seed 0 |
| Generation root seed | 20260911; per-game seeds derived deterministically |
| Bootstrap | 0 is the script default; the runtime value used for the final artifact is not separately recorded |