Diffusion models excel at image synthesis, but they remain limited in their ability to reliably satisfy structured spatial reasoning constraints. In conditional data distribution modeling tasks with implicit logical structure, such as puzzles defined by visible clues paired with consistent solutions, state-of-the-art generative models tend to approximate pixel-space distributions without learning the underlying logical rules required for inference. To address this limitation, we present a novel framework for spatial reasoning with diffusion models that leverages unsupervised object discovery and abstractions of object relations. We show that the relational knowledge derived from object-centric representations enriches diffusion models with structural primitives, allowing them to effectively guide the generative representation space during both training and inference, and enabling conditional image generation that satisfies reasoning constraints. Additionally, we introduce a large-scale generative spatial reasoning benchmark with four datasets inspired by human-solvable puzzles. Our results show that relational abstractions significantly improve reasoning capabilities of diffusion models on a variety of complex reasoning tasks, while enabling robust generalization in out-of-distribution settings.
Figures & tables
Figure 1: Illustration of RDM with Discoverer and the Reasoner blocks. During training, the clues ( xc ) and solutions (x0) are individually fed to the pretrained frozen ( ) slot encoder to obtain the corresponding feature ( f1:m′ ) and positional ( p1:m′ ) embeddings. Slot Attention yields the slots ( s1:k′ ) for reconstruction by the slot decoder, and object-specific features and positions, z′=(f~1:k′,p~1:k′) for relational abstractions. The relational abstraction module consists of L abstractor layers that optimize the relational embeddings ( ), which are introduced into the U-Net bottleneck through a cross-attention layer. Since the solution is unavailable at inference time, its disentangled representations ( f~1:k0,p~1:k0 ) are suppressed by a default learnable embedding ( ∅ ) with probability pdrop .
Figure 2: Akari as a grid-based image completion puzzle.
Figure 3: Coldoku puzzle, which combines Sudoku with a color arrangement that also satisfies Sudoku rules. Tableau 10 palette is used, which is designed to have a large perceptual difference between colors.
Figure 4: Tangram example. A valid solution uses all pieces matching the colors and shapes, given the (a) and (b) as clues .
Figure 5: LogicFace example with 4 features: smiling ( OR ), male ( AND ), glasses ( XOR ), young ( IMPLIES ). Features from A and B are combined in S via the corresponding logic operators (e.g., glasses : A=0 , B=1 , S=1 , satisfying XOR ).
Akari
Coldoku
in-dist.
low
high
Easy (50-60)
Medium (40-50)
Hard (30-40)
blb
acc
blb
acc
blb
acc
dgt
clr
acc
dgt
clr
acc
dgt
clr
acc
Base
86.02
80.04
43.62
39.48
92.22
88.92
93.53
94.73
89.00
61.81
78.56
49.06
11.38
25.31
3.88
SRM ( Wewer et al., 2025 )
parallel
35.82
18.80
1.60
0.94
54.18
33.34
94.67
96.74
91.06
61.02
78.53
51.01
12.17
25.91
4.00
sequential
00.00
00.00
–
–
–
–
93.05
96.70
92.37
82.21
83.61
68.35
40.95
36.66
20.78
Table 1: Experimental results on Akari and Coldoku datasets.
Tangram
LogicFace
□
△S
△M
△L
shape
acc
smiling ( OR )
male ( AND )
glasses ( XOR )
young ( IMP. )
avg
acc
Base
23.06
24.22
42.82
41.50
55.60
57.72
15.46
54.60
76.52
67.34
38.90
57.84
11.24
SRM ( Wewer et al., 2025 ) 3 3 3 Diffusion forcing led to highly unstable training in Tangram , where both SRM parallel and sequential were unable to outperform Base (similar to LogicFace ). These findings also align with Wewer et al. (2025) .
parallel
–
–
–
–
–
–
–
37.00
38.16
49.26
39.44
40.97
2.60
sequential
–
–
–
–
–
–
–
25.04
58.38
50.58
73.86
51.94
5.70
RDM [S]
Table 2: Experimental results on Tangram and LogicFace datasets.
Figure 6: Performance of RDM [ L ] in Tildoku .
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Akari
Coldoku
Tangram
LogicFace
H×W
198×198
180×180
128×128
128×128
h×w
11×11
9×9
8×8
8×8
k
121 / 16
81 / 16
7
6
# samples
1M
1M
1,013
30,000
neural evaluator
✗
✓
✗
✓
Appendix
Table 3: General statistics: image ( H×W ) and bottleneck ( h×w ) resolution; number of slots ( k ) and unique samples (# samples); and neural evaluation.
Figure 7: Akari samples from the training set.
Figure 8: Visualization of Coldoku samples from the training set (top: solution , bottom: clue ).
Figure 9: Visualization of Tildoku solution samples with different tilting ratios.
Figure 10: Tangram layouts ordered by reasoning difficulty. From easy configurations (left) containing salient geometric cues that simplify piece identification to complex silhouettes (right) characterized by high structural ambiguity, where multiple valid piece arrangements may exist for the same layout.
Figure 11: Visualization of Tangram samples from the training set.
OR
AND
XOR
IMPLIES
A
B
S
A
B
S
A
B
S
A
B
S
0
0
0
0
0
0
0
0
0
0
0
1
0
1
1
0
1
0
0
1
1
0
1
1
1
0
1
1
0
0
1
0
1
1
0
0
1
1
1
1
1
1
1
1
0
1
1
1
Appendix
Table 4: Logical operators between A and B , yielding S . Since there are four operator and each operator has four valid statements, LogicFace has 256 possible configurations.
Figure 13: LogicFace examples from the training set with the four binary features: smiling ( OR ), male ( AND ), glasses ( XOR ), young ( IMPLIES ).
Figure 17
k
d
channels
#conv/layer
#params.
#steps
Accuracy
Slot-Attention
Akari
16
128
[16, 32, 64, 128]
1
1.4M
100K
99.58
Coldoku
16
128
[32, 64, 128, 256]
1
5.4M
400K
88.19
Tangram
7
128
[16, 32, 64, 128]
1
1.4M
300K
16.34
LogicFace
6
256
[32, 64, 128, 256]
2
7.6M
200K
72.54
Tildoku
20
128
[16, 32, 64, 128]
2
7.6M
400K
61.15
Appendix
Table 5: Architecture, training and performance details of the Discoverer in configurations S and M .
d
channels
#conv/layer
#params of D
#params of R
#params of U-Net
# steps
RDM [S]
128
[64, 128, 256, 512]
[2, 2, 1, 1]
(Table 5 )
4M
25M
300K
RDM [M]
256
[64, 128, 256, 512]
[2, 2, 2, 2]
(Table 5 )
4M
35M
300K
RDM [L]
Akari
512
[128, 256, 512, 512]
2
40M
15M
120M
40K
Coldoku
512
[128, 256, 512, 512]
2
40M
15M
120M
45K
Tangram
512
[128, 256, 512, 512]
2
40M
15M
120M
80K
Appendix
Table 6: Architecture and training details of the Reasoner .
Figure 14: Disentanglement of the tans on Tangram inputs.
Figure 15: Disentanglement of the facial attributes on LogicFace inputs.
ID
OoD
5-10% (low)
35-40% (high)
unl
dst
ovp
blb
acc
unl
dst
ovp
blb
acc
unl
dst
ovp
blb
acc
τ=10
Base
0.46
0.10
0.03
74.08
67.35
3.22
0.22
0.48
17.02
13.80
0.17
0.04
0.01
89.78
86.80
SRM
parallel
2.38
0.81
0.15
26.52
12.26
15.43
0.63
0.77
1.04
0.54
1.27
0.63
0.06
45.42
27.22
sequential
7.01
9.81
17.79
0.00
0.00
0.26
4.64
87.15
0.00
0.00
13.87
13.47
6.28
0.00
0.00
Appendix
Table 7: Detailed results on Akari .
acc
Easy (50-60)
Medium (40-50)
Hard (20-30)
dgt
clr
acc
dgt
clr
acc
dgt
clr
acc
Base
59.20
66.84
40.28
22.62
45.92
12.48
1.50
9.42
0.24
SRM ( seq. )
44.87
25.87
13.80
24.31
21.31
7.08
3.62
5.62
0.62
RDM [L]
clue
62.52
79.74
50.28
33.02
52.84
19.18
3.78
10.12
0.64
solution
65.72
80.92
53.26
43.70
60.12
28.40
8.34
14.48
1.92
Appendix
Table 8: Detailed results on the Tildoku dataset.
Configuration [S]
Configuration [L]
(∗↑)
□
△S
△M
△L
shape
acc
□
△S
△M
△L
shape
acc
τ=10
Base
17.96
17.82
29.78
34.98
39.33
40.09
12.37
16.08
15.24
21.78
24.64
28.08
28.64
11.46
RDM
clue
24.84
25.37
37.55
40.71
41.90
42.41
22.20
34.72
34.26
41.80
43.34
44.88
44.98
31.96
solution
31.43
31.80
36.73
38.69
40.06
40.51
28.45
33.34
32.66
39.90
41.64
43.16
43.52
29.94
full
34.49
34.69
38.65
42.67
43.82
4.31
30.18
22.20
21.38
33.26
36.28
40.48
41.44
16.74
Appendix
Table 9: Detailed results on Tangram .
(∗↑)
Configuration [S]
Configuration [L]
smiling ( OR )
male ( AND )
glasses ( XOR )
young ( IMP. )
avg
acc
fid
smiling ( OR )
male ( AND )
glasses ( XOR )
young ( IMP. )
avg
acc
fid
τ=10
Base
71.28
79.39
60.77
73.50
71.24
26.03
44.95
55.08
55.42
39.24
70.28
54.51
8.62
35.92
RDM
clue
87.74
96.66
93.61
81.67
89.92
66.08
37.66
83.34
94.52
49.60
70.02
74.37
27.98
37.18
solution
68.53
78.34
49.81
72.10
67.19
18.98
50.21
65.41
66.44
49.41
66.05
61.83
14.95
39.71
full
84.97
94.74
90.81
78.55
87.27
59.57
36.37
78.38
91.74
50.68
65.66
71.61
25.08
38.04
Appendix
Table 10: Detailed results on LogicFace .
Figure 18: Examples of generated Tangram solutions.
Figure 20: Examples of generated LogicFace solutions shown in S , when conditioned on A and B .
Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzzles. We ask how such reasoning can be carried over to pixels, where no symbolic representation is available. We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters. The Thinker is trained with the standard reconstruction loss alone, without symbolic targets, a solver, or a verifier. PaTh solves 92.5% of hard MNIST Sudoku puzzles (prior best 75%) and 71.2% of extreme ones (prior best 4.1%), with 10M parameters against 82M for a standard diffusion model. It also improves on mazes, Queens, and CLEVR scenes with specified spatial relations, and its advantage grows with problem size. Diagnostic experiments show that PaTh recovers from injected mistakes that the diffusion model cannot repair, especially when many cells are wrong. Together, these results show that reasoning mechanisms developed for symbolic data can be integrated into pixel-space diffusion without symbolic supervision, opening a path toward generating data under increasingly complex constraints.
Paweł Skierś, Małgorzata Grzanka, Wojciech Masarczyk +2
Warsaw University of Technology · IDEAS Research Institute · University of Amsterdam
Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled spatial data from external vision sources or synthetic engines. In contrast, we argue that for many tasks, spatial reasoning capabilities are already present in pre-trained LRMs but require alignment through logical coherence under geometric 2D and 3D constraints. In this work, we propose a self-supervised reinforcement learning (RL) framework that targets the internal reasoning process without requiring ground-truth annotations. By formalizing the notion of consistency verifiers -- reward functions that check for geometric and semantic consistency under transformations -- we demonstrate that models can improve their spatial reasoning abilities. We use both image transformations, like flipping, and textual transformations, like swapping the order of objects in the question, and propose a new optimal transport-based RL strategy, OT-GRPO, which is a minimal-matching variant of group relative policy optimization tailored to pairwise verifiers. We show that this label-free consistency training approaches the accuracy of models trained with ground-truth supervision and achieves similar generalization across diverse tasks and data domains.
Theo Uscidda, Marta Tintore Gazulla, Maks Ovsjanikov +2
CREST, ENSAE, Institut Polytechnique de Paris · Google Zurich · Google DeepMind
Diffusion models have achieved success in high-fidelity data synthesis, yet their capacity for more complex, structured reasoning like text following tasks remains constrained. While advances in language models have leveraged strategies such as latent reasoning and recursion to enhance text understanding capabilities, extending these to multimodal text-to-image generation tasks is challenging due to the continuous and non-discrete nature of visual tokens. To tackle this problem, we draw inspiration from modular human cognition and propose a recursive, sparse mixture-of-experts framework integrated into conventional diffusion models. Our approach introduces a recursive component within joint attention layers that iteratively refines visual tokens over multiple latent steps while efficiently sharing parameters via sparse selection of neural modules. At each step, a gating network is devised to dynamically select specialized neural modules, conditioned on the current visual tokens, the diffusion timestep, and the conditioning information. Comprehensive evaluation on class-conditioned ImageNet image generation tasks and additional studies on the GenEval and DPG benchmark demonstrate the superiority of the proposed method in enhancing model image generation performance.
Yuwei Sun, Yuxuan Yao, Hui Li +1
Shanghai Academy of AI for Science · Fudan University