With the constant advancements in AI, one possibility is to model agents after humans and, in turn, use these agents to carry out synthetic interactions. Such models could be used to predict interactions between their real counterparts, or potentially interactions at larger scales. In this paper, we test a more controlled version of this question through chess. We use 8 elite chess players, seal their direct pairwise games, learn each player independently using different methods, and then compose the resulting models on the withheld dyads. To evaluate the generated interactions, we use two measurements: opening-family total variation distance and win-draw-loss (WDL) total variation distance. M1 primarily improves WDL fidelity while producing smaller opening-family improvements, whereas M2 produces much larger opening-family improvements while having little effect on WDL-TV. For opening-family behaviour under M2, the correct assignment of the eight learned player identities also gives the closest match among all 8!=40,320 possible assignments. These results show that at least some properties of previously unseen interactions can be recovered from independently learned individuals. The partial recovery observed here may reflect limitations of the current individual modelling methods rather than a fundamental limit on compositional interaction recovery. An additional post-hoc method that combines the two mechanisms improves both measurements, suggesting that recovery across these behavioural properties is not necessarily mutually exclusive.
Figures & tables
Method
Rank 1
Mean Δ NLL
Wrong − correct
Positions
M1
8/8
0.006802
0.004365
47,136
M2
8/8
0.009278
0.020449
46,950
Table 1. Individual behavioural validation on held-out nonsealed data. Δ NLL denotes generic-minus-correct NLL. M2 uses candidate-matched NLL and is therefore not directly comparable in scale with M1.
Method
Metric
GG
AG
GB
AB
AB−GG
M1
WDL-TV
0.164171
0.157753
0.156141
0.150148
-0.014024
M1
Opening-family TV
0.554433
0.547406
0.544765
0.536973
-0.017461
M2
WDL-TV
0.163166
0.162582
0.162589
0.163408
0.000241
M2
Opening-family TV
0.554711
0.501370
0.488343
0.437202
-0.117509
Table 2. Sealed target–target interaction fidelity for M1 and M2. Lower TV is better. The final column reports AB−GG , so negative values indicate that composing both personalized players reduces the distance from the corresponding real interaction.
Metric
M1 AB
M2 AB
M3 AB
WDL-TV
0.150148
0.163408
0.152036
Opening-family TV
0.536973
0.437202
0.444428
Table 3. Exploratory M3 results. The M1 and M2 columns show the corresponding fully personalized AB conditions for reference. Lower TV is better.
Method
Metric
Improved
Mean
Median
Min
Max
M1
WDL-TV
22/28
-0.014024
-0.010900
-0.088700
0.046300
M1
Opening-TV
22/28
-0.017461
-0.018883
-0.058385
0.027417
M2
WDL-TV
13/28
0.000241
0.001650
-0.018300
0.015700
M2
Opening-TV
27/28
-0.117509
-0.105987
-0.364176
0.004373
M3
WDL-TV
19/28
-0.011037
-0.010500
-0.064400
0.032800
M3
Opening-TV
27/28
-0.109967
-0.090846
-0.347976
0.006572
Table 4. Variation in AB−GG across the 28 target-player dyads. “Improved” reports the number of dyads for which AB has lower TV than GG.
Figure 1. Dyad-level change from GG to AB for each method. Negative values indicate that the fully personalized interaction is closer to the withheld real interaction. M1 shows a broader reduction in WDL-TV, whereas the opening-family reductions for M2 and M3 are substantially larger and occur for nearly all dyads.
Method
Metric
Matched TV
Permutation mean
Rank / 40,320
Substitutions beaten
M1
WDL
0.1501
0.1563
143
60%
M1
Opening
0.5370
0.5460
1,789
63%
M2
WDL
0.1634
0.1643
8,096
52%
M2
Opening
0.4372
0.5800
1
90%
M3
WDL
0.1520
0.1569
1,625
53%
M3
Opening
0.4444
0.5886
1
90%
Table 5. Identity-specificity results for the matched AB compositions. Lower TV and rank are better; higher "Substitutions beaten" is better.
Method
Metric
Cond.
Observed TV
Reference TV
M1
WDL
GG
0.1642
0.1048
M1
WDL
AB
0.1501
0.1070
M1
Opening
GG
0.5544
0.3813
M1
Opening
AB
0.5370
0.3805
M2
WDL
GG
0.1632
0.1049
M2
WDL
AB
0.1634
0.1050
Table 6. Observed TV and finite-sample simulation results for GG and AB, averaged equally over the 28 dyads.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Stage 1 (population)
Stage 2 (Broadcast)
Initialization
Random; PyTorch defaults, with Xavier initialization for Elo/GAB embeddings only
Stage-1 final checkpoint
Optimizer steps
1,000,000
200,000
Effective batch size
512
512
Peak learning rate
5×10−5
1×10−5
Learning-rate floor
1×10−5
2×10−6
Weight decay
1×10−6
1×10−6
Appendix
Table 7. Base-model training configuration. Both stages use AdamW, linear warmup with cosine-annealed restarts, mixed precision, and gradient-norm clipping at 3.5.
M1
M2
Trained parameters
One 128-d vector per player
Candidate CNN, 32-d style vector per player, scale σ
Action space
Full 4,352, legal-masked
Base top-5; human move appended during train/validation if absent
Loss
Cross-entropy
Cross-entropy over candidates
Training scope
8 independent runs
1 joint run over all players
Train/val split
80/20 by game, independently per player
80/20 by game with shared seeded split state
Phase balancing
Per player, per split
Per player, per split
Appendix
Table 8. Personalization training configuration. The base model is frozen in both methods.
Player
Representative Elo
Magnus Carlsen
2855
Wesley So
2769
Levon Aronian
2756
Fabiano Caruana
2786
Maxime Vachier-Lagrave
2751
Hikaru Nakamura
2829
Appendix
Table 9. Representative Elo values used for conditioning.
Software
Python / PyTorch
3.12.3 / 2.6.0 (CUDA 12.4)
python-chess / pyarrow
1.11.2 / 25.0.1
Seeds
Training
Seed 0 where specified by the training pipeline; M1/M2 data splitting and phase balancing also use seed 0