Conditional Generation of Creative Chess Puzzles with Diffusion Models
Organizations: Aalto University · Google DeepMind
Abstract
While modern language models demonstrate impressive generative capabilities, they often struggle with constrained, counter-intuitive creative tasks. To address this limitation, we explore chess puzzle generation as a rigorous testbed for computational creativity and reasoning, a domain where altering a single piece can invalidate an entire solution. We propose a novel approach for conditional generation of creative chess puzzles using masked diffusion models. Unlike previous methods, our non-directional diffusion approach allows for conditioning on specific tactical themes and partial board positions. We introduce a novel auxiliary task of simultaneous best-move prediction, which improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%. To further optimize solution uniqueness and theme conditioning, we establish a reinforcement learning framework adapted from Denoising Diffusion Policy Optimization (DDPO). This RL training increases the yield of unique and theme-matching positions by 89.1%. Finally, we release the first open-weights models (Appendix B) for chess puzzle generation, offering a new pathway for controllable, creative generation.
Figures & tables
| Training | Predict best move | Legal (%) | Unique (%) | Counter-Intuitive (%) | Puzzle (%) | Unique ∗ & Theme (%) | Theme | Unique ∗ (%) |
|---|---|---|---|---|---|---|---|
| Supervised | No | 98.97 0.05 | 39.16 0.23 | 1.00 0.08 | 0.39 0.03 | 19.10 0.19 | 82.04 0.39 |
| Supervised | Block | 98.99 0.05 | 41.25 0.24 | 0.99 0.07 | 0.41 0.03 | 20.77 0.20 | 83.64 0.36 |
| Supervised | Yes | 98.96 0.05 | 43.69 0.24 | 0.93 0.07 | 0.41 0.03 | 22.37 0.20 | 84.05 0.34 |
| RL v1 | Yes | 99.76 0.04 | 38.02 0.43 | 2.23 0.20 | 0.85 0.08 | 16.40 0.32 | 73.80 0.82 |
| RL v2 | Yes | 99.94 0.01 | 66.15 0.20 | 0.36 0.03 | 0.23 0.02 | 42.32 0.22 | 87.21 0.21 |
| Supervised ( Feng et al., 2026 ) | 99.72 | 30.89 | 1.11 | 0.34 | — | — | |
| Training | Predict best move | Board (Self) | Board (Lichess) | PV (Self) | PV (Lichess) |
|---|---|---|---|---|---|
| Supervised | No | 11.63 0.01 | 11.16 0.01 | 0.78 0.002 | 0.93 0.001 |
| Supervised | Block | 11.67 0.01 | 11.18 0.01 | 0.80 0.002 | 0.93 0.001 |
| Supervised | Yes | 11.70 0.01 | 11.19 0.01 | 0.81 0.002 | 0.93 0.001 |
| RL v1 | Yes | 15.15 0.01 | 14.34 0.01 | 0.76 0.002 | 0.97 0.001 |
| RL v2 | Yes | 12.12 0.01 | 12.10 0.01 | 0.88 0.002 | 0.97 0.001 |
| Supervised ( Feng et al., 2026 ) | 10.859 | 8.528 | (0.837) | (0.637) | |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | value |
| Model | |
| Parameter count | 268M |
| Transformer block count | 16 |
| Attention head count | 8 |
| Embedding dimension | 1024 |
| Activation | SwiGLU (widening factor 2.66) |
| Category | Themes |
|---|---|
| State-of-game | Opening, Middlegame, Endgame |
| Type-of-endgame | Pawn endgame, Bishop endgame, Knight endgame, Rook endgame, Queen endgame, Queen rook endgame |
| Type-of-checkmate | Mate, Back rank mate, Boden mate, Smothered mate, Hook mate, Double bishop mate, Arabian mate, Dovetail mate, Anastasia mate |
| Length-of-checkmate | Mate in 1, Mate in 2, Mate in 3, Mate in 4, Mate in 5 |
| Length-of-puzzle | One move, Short, Long, Very long |
| Winning | Crushing, Advantage |