cs.AISep 29, 2026

Conditional Generation of Creative Chess Puzzles with Diffusion Models

Authors: Aatu Selkee, Severi Rissanen, Xidong Feng, Tom Zahavy, Eric Malmi

Organizations: Aalto University · Google DeepMind

Abstract

While modern language models demonstrate impressive generative capabilities, they often struggle with constrained, counter-intuitive creative tasks. To address this limitation, we explore chess puzzle generation as a rigorous testbed for computational creativity and reasoning, a domain where altering a single piece can invalidate an entire solution. We propose a novel approach for conditional generation of creative chess puzzles using masked diffusion models. Unlike previous methods, our non-directional diffusion approach allows for conditioning on specific tactical themes and partial board positions. We introduce a novel auxiliary task of simultaneous best-move prediction, which improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%. To further optimize solution uniqueness and theme conditioning, we establish a reinforcement learning framework adapted from Denoising Diffusion Policy Optimization (DDPO). This RL training increases the yield of unique and theme-matching positions by 89.1%. Finally, we release the first open-weights models (Appendix B) for chess puzzle generation, offering a new pathway for controllable, creative generation.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models

    May 17, 2026Ethan TangChessModel Checkpoints

  2. RePAIR: Predictive Self-Supervised Representation Learning in Chess

    Jun 10, 2026Christoph Koller, Johannes Fürnkranz, Timo BertramChessSelf-Supervised Representations

  3. Benchmarking Prompt Optimization of Large Language Models With Chess

    Sep 30, 2026Timothée Lesort, Alejandra López de Aberasturi Gómez, Tristan Karch +4ChessLarge Language Model Benchmarks