I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein's curved universe.'' So wrote M.C. Escher about his 1956 lithograph Print Gallery. Nearly half a century later, a mathematical analysis related its geometry to an untwisted source image through a conformal power map z↦zα, α∈C. Building on this construction, we use a frozen text-to-image diffusion model to generate new self-referential scenes. Prompting alone does not enforce the recursion, while a post-hoc transformation can leave structures poorly connected. Applying the transformation during sampling is also insufficient: the denoiser may "repair" the intended distortion or drift out of the prescribed geometry. We construct a generalized inverse T† of the non-invertible image transformation T, adapted to its recursive constraint. In the idealized formulation, the Penrose identity TT†T=T makes TT† an idempotent projection onto geometrically admissible images. Yet denoising only the transformed image remains an out-of-distribution task, even with projection. We therefore braid denoising steps with T and T†: source-space steps develop the untwisted scene, while transformed-space steps refine its appearance and connections in the final geometry. We generate Print Gallery-like compositions and explore further transformations. Rather than distorting a finished image, we let the scene and its distortion develop together.
Figures & tables
Figure 1
Figure 2: The Droste effect in three representations. Left: a tile repeated in logarithmic coordinates, with horizontal period L=logλ and vertical period 2π . Middle: returning to image coordinates gives scale repetition, s(λz)=s(z) . Right: the conformal power map adds rotation, producing the Escher geometry with y(eαLw)=y(w) . All three panels use the same pattern, with λ=16 .
Figure 3: Braided sampling. Warm-up interleaves denoising with untwisted repetition T0 . The sampler switches between Droste and Escher representations using T and T† , with denoising blocks D(n) between switches. It ends with refinement in the Escher representation. At each switch, 4× super-resolution precedes the map, then resizing to working resolution. Dashed arrows indicate repetition; the noise-level axis is schematic. The early Droste preview is an illustrative blur of the later panel, not a recorded checkpoint.
Figure 4: Overview of generated results. Diverse scenes with conformal twists, square spirals, and multi-center maps. See the animated supplementary material .
Figure 5: Conformal results with input prompts. Longer prompts are excerpted. Animated views appear in the supplementary material .
Figure 6: Twist and transformation choices. Each row shares a prompt and initial seed. The conformal columns use ∣p∣=1,2 (negative p for the bookshop), followed by Poles, Möbius, Square, and Rimrings. Longer prompts are excerpted. Rimrings omits inverse steps. See the supplementary material for animated views.
Figure 7: Rimrings: repetition with visible boundaries. Bands accumulate near an outer rim; their divisions form material layers, ledges, coils, or plate edges. These examples omit the unstable inverse. Animated views appear in the supplementary material .
Figure 8: Baselines and cumulative ablation. Rows fix the base prompt and seed; the blue callout supplies the prompt extension. The first three columns show prompt-only, extended-prompt, and post-hoc T baselines. Starting from T at σ=0.87 , subsequent columns add warm-up, T† followed by a final T at σ=0.5 , in-loop SR (ours), and optional time travel. Identical central crops are enlarged 8× below.
Figure 9: Failure cases. Left: recursion anchors on the window rather than the canvas. Right: time travel amplifies the seam (without, then with).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
Base model
FLUX.1-dev ( 12 B rectified-flow DiT), FluxPipeline , fp16
Guidance
distilled guidance embedding, scale 3.5 (no CFG pass)
Table 3: Operator schedule over the 128 steps ( σ:1→0 ).
Name
Symbol
Default
Swept
Inset scale
sin
1/4
1/8,1/16
Periods per turn
p
1
−1,2
Warm-up start
σw
0.95
off, 0.99
First op
σhi
0.87
0.80,0.92
Op gap
steps
9
5,13
Time travel
σtt
off
0.60,0.75
Appendix
Table 4: Swept hyper-parameters and their defaults.
Figure 10: Checking geometric round trips. From left to right: x , Tx , P(Tx) , P2(Tx) , and the residual ∣P2(Tx)−P(Tx)∣ , where P=TT† . Rows use deck ratios λ=2,4,16 . The three transformed images coincide in the continuous construction; the rasterized comparison shows blur and residual differences near checkerboard edges. The measured residual ∣P2(Tx)−P(Tx)∣ is 0.011 , 0.012 , 0.014 (top to bottom), each at the bilinear-resampling floor.
Figure 11: Transfer to a pixel-space backbone. The conformal T -cycle recipe applied to PixelDiT-1300M, with the same prompts, numerical seeds, transform, and scale as the FLUX configuration, and adapted intervention noise levels. The warm-up start changes from 0.95 to 0.96 , σhi from 0.87 to 0.92 , and σlo from 0.5 to 0.69 . Where time travel is used, its noise level changes from 0.75 to 0.82 . Per-cell tags retain the FLUX naming convention (e.g. tt75 ) and do not denote the adapted noise levels.
Figure 12: Additional Escher-recursion results (FLUX.1-dev). Curated gallery across subjects, transforms (conformal, square, poles, Möbius), inset scales and spiral orders; (Page 1 of 2.)
Planar tiled diffusion denoises overlapping windows of one rectangular canvas. The hyperbolic plane has no such canvas, and its area grows exponentially with radius. We introduce HyperbolicDiffusion, a training-free method for generating finite visual fields directly on the hyperbolic plane H2. Our Hyperbolic Blooming Cover reduces window placement to a compact dynamic program that runs in seconds while providing strong theoretical guarantees. Permanent surface IDs form a shared latent canvas: a standard diffusion model denoises local windows, whose predictions are fused back onto H2. Because curvature causes residual disagreement and blur at multi-window junctions, a geometry-derived second stage re-noises and repairs precisely those regions. The resulting fields are sharp, reprojectable, and consistent across viewpoints, providing a prompt-driven generative counterpart to Escher's Circle Limit series.
Photomosaics are large images whose local regions are seen as independent tiles while their overall arrangement forms a coherent scene. Generating them at high resolution, with every tile convincing in its own right, is computationally expensive, since the canvas must hold many detailed tiles at once. We present PhotoQuilt, a training-free framework that generates photomosaics at arbitrary resolution. Diffusion models struggle to satisfy both scales at once, as direct high-resolution generation is costly and tends toward one smooth image rather than a mosaic, while patch-based tiling keeps local detail but loses global structure. PhotoQuilt resolves this with a bootstrapped tiled denoising procedure. We first produce a global composition at low resolution to fix the layout, then upscale it in latent space and re-inject noise to restore generative capacity. Denoising proceeds within fixed tiles, so each forms its own image while the shared global structure holds them in one layout. Because tile generation is handled separately, PhotoQuilt scales to large canvases without quadratic attention cost. Experiments show that PhotoQuilt outperforms current baselines on both global structure and local realism.
Koorosh Roohi, Javad Rajabi, Andrew Fleet +1
University of Toronto · Vector Institute · KITE Research Institute +2
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary-condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify-then-Diffuse (RTD), a training-free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft-Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout-agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state-of-the-art compositional fidelity and robust gains. On the AE-Bench object pair subset, RTD improves BLIP-VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3× faster. Code will be released at https://github.com/Z-yiwei/rectify-then-diffuse
Ning Zhu, An Chen, Mengfei Zhao +4
Glasgow College, University of Electronic Science and Technology of China · School of Mathematical Sciences, University of Electronic Science and Technology of China