Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale σ, we treat the denoiser as a self-map and study its dynamics. For an exact denoiser, fixed points correspond to critical points of the smoothed data density, while attractors correspond to its modes; as σ increases, sample-level modes merge into progressively coarser ones. This suggests a geometric view of memorization: examples that receive excess probability mass due to duplication or overfitting, as well as outliers, should remain distinguishable under stronger smoothing than ordinary examples. We quantify this persistence by the critical scale σc, the largest noise scale at which an example is retained by the fixed-scale dynamics. In conditional models, the same construction extends naturally to image--caption pairs. Experiments in controlled settings and on large-scale models show that σc tracks memorization arising from duplication, overfitting, and outliers, and identifies both memorized and partially memorized examples in Stable Diffusion. Moreover, σc yields interpretable measures of the image spatial distribution and caption dependence of memorization.
Figures & tables
Figure 1: A diffusion model as a family of dynamical systems. Fixing the noise level σ and feeding the denoiser its own output defines a self map Mσ ( left ). A memorized training image is a fixed point of this map: its orbit stays at the image, whereas the orbit of a control image drifts away and crosses the escape radius r ( bottom left ). Repeating this retention test at every σ ( right ) yields the critical scale σc , the largest noise level at which the image is still returned; it is 6.7× larger for the memorized image, effectively detecting it.
Figure 2: One model, a family of dynamical systems. Orbits of the exact map Mσ (Eq. 1 ) for samples from an 8-mode mixture of gaussians. Left: at small σ each datum carries its own attractor. Middle: at intermediate σ the attractors are the 8 population modes. Right: at large σ a single attractor remains, at the data mean.
Figure 3: Scale space of attractors. Exact map Mσ on a 2-D training set: three groups of three sub-clusters (black dots), plus one datum repeated 15 times (red star). Top: basins of attraction at four scales, coloured by the reached attractor ( × ); arrows are the drift Mσ(x)−x . Bottom left: merge tree of the attractors reached from the data (width = number of data collected). Single data lose their attractor at small σ ; the duplicated datum (red) keeps it until merging with the neighbouring sub-clusters. Bottom right: number of attractors. A denoiser trained on the same data reproduces this scale space, basins, merge tree and counts alike (Appendix C , Fig. 8 ).
Figure 4: Duplication on full CIFAR-10. (a) σc of each image, grouped by multiplicity, in the model trained with duplicates (circles) and in the released model trained on the same images without duplicates (triangles). (b) σc as a function of the iteration budget K (median per group): constant for images with m≥8 , which are fixed points, and decreasing as 1/K for held-out and m=1 images, which are transients. (c) Orbits Mσk(x) of a memorized training image ( m=400 ) and of a training image seen once ( m=1 ), with σ fixed along each row and k along the columns; the dashed line marks the σc of each image ( K=50 , r=0.5 ; 1.62 and 0.14 ). The memorized image is still reproduced after 50 iterations at σ=1 , where the other one has already vanished.
Figure 5: Rarity: detection thresholds of σc and of sampling. Fraction of added images detected as a function of their multiplicity m , for sources of increasing atypicality. Orange : the image appears at least once among 105 samples; the threshold is m≈8 for every source. Blue : σc exceeds the 95 th percentile of the images of the same source not added to the model; the threshold decreases to m=1 as the images become more atypical. The shaded region is memorization detected by σc but not by sampling.
Score
MV vs. control
MV vs. swap
TV vs. control
TV vs. swap
MV vs. TV
∥ϵc−ϵ∅∥ Wen et al. (2024)
0.999
0.606
0.999
0.184
0.902
∥ϵc(x0)−ϵ∅(x0)∥
0.999
0.887
0.996
0.696
0.868
σc , K=16
0.917
0.904
0.695
0.803
0.909
Δlogσc , K=16
0.8569
0.908
0.824
0.939
0.820
∣Δlogσc∣ , K=16
0.984
0.860
0.879
0.354
0.905
local ∣Δlogσc∣ , K=16
0.997
0.865
0.993
0.458
0.935
Table 1: Memorization-detection on Stable Diffusion v1.4. AUC for distinguishing MV and TV examples from controls, caption-swapped pairs, and MV from TV. Our global scores use r=0.25 , and local ∣Δlogσc∣ uses r=0.1 . For both variants of Wen et al. (2024) , we report its best-performing configuration (Appendix F ).
Figure 6: Critical-scale gaps reveal spatial and image-driven memorization. Left: A TV example: the training image, generations from its training caption, and the local gap ∣Δlogσcu∣ . The gap is large where all generations copy the training image (the room) and near zero where they vary (the curtain pattern). Right: A wrong caption exposes images already retained by the unconditional model. Center: wrong-caption gap Δlogσc against unconditional inpainting RMSE (top) and unconditional σc (bottom); bars include std deviation over 5 captions. Side: a negative-gap image is retained by unconditional dynamics but lost when conditioning on a wrong caption; unconditional inpainting restores the masked region.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Definition 1 on Stable Diffusion. Top: the map Mσ . The image is encoded once into the latent x0 ; the U-Net is evaluated with the caption at noise level σ , and its output xk−σϵ^ is fed back as the next input. Bottom: each row is an orbit (Mσk(x))k≥0 , decoded for display, of a memorized image under its own caption ( w=1 ). First row: started at the image, the orbit does not move; the image is a fixed point x∗ . Second row: started with part of the image replaced by a grey square, the orbit restores it and converges to the same x∗ . The occluded image lies in the basin Bσ(x∗) , so x∗ is an attractor. Third row: the same start at a larger σ , which defines a different map in the family {Mσ}σ>0 . There the start is no longer in the basin, and the orbit leaves.
Figure 8: A trained denoiser reproduces the scale space of Fig. 3 . Same training set and same panels as Fig. 3 , with the exact map Mσ replaced by Mσ(x)=Dσ(x) for an EDM-preconditioned MLP trained on the 141 points. Top: basins at the four scales of Fig. 3 , coloured by the attractor ( × ) reached from the data; grey marks the grid points that reach none of them, in the corners of the domain where the network saw no training data. Bottom left: merge tree of the trained map. Bottom right: its attractor count against that of the exact map. The two maps agree: the counts are equal at 80% of the 90 scales, the partitions of the training set they induce have adjusted Rand index 0.99 on average ( 0.76 at worst, at the smallest scales where single data split), and the duplicated datum keeps its attractor over the same range of scales in both.
Figure 9: The exact denoiser follows the law. σcemp of the added images of the rarity experiment against their kernel distance dkern to the training set (Eq. 28 ), for images added once (grey) and 32 times (red), under the escape test with K=1 , r=0.2 . The lines are the law σc=dkern/2log(N/m) of Theorem 5 , drawn without fitting. Doubling the distance doubles the critical scale, whereas 32 copies increase it by a factor 1.26 .
intervention
group
σcemp
σc
κ
duplication
m=1
4.99
0.142
0.028
m=8
5.60
1.036
0.194
m=800
9.91
1.431
0.145
m=800 , control model
9.91
0.269
0.024
overfitting
N=5⋅104
4.93
0.056
0.012
N=2000
6.11
0.243
0.043
Appendix
Table 2: The three interventions on a common scale. σcemp is the critical scale of the exact empirical denoiser of the training set of each model under the same escape test, equation 40 ; κ=σc/σcemp is the fraction of this solution that the network implements. Group medians, training images only: the ceiling varies little, and the effect is carried by κ .
Figure 10: Orbits across scales, one image per group. Orbits of the model trained with duplicates, with σ fixed along each row and the iteration index k along each column, for a memorized image ( m=800 ), a training image with m=1 and a held-out image, each chosen at the median σc of its group. The dashed line marks the σc of the image ( K=64 , r=0.5 ). The memorized image is unchanged after 64 iterations at σ=1.43 , whereas the other two converge to a uniform colour at σ=0.24 , with nearly identical critical scales ( 0.122 and 0.118 ).
Figure 11: Overfitting: the same images across training-set sizes. (a) σc of the 100 training images shared by all models (blue) and of 100 held-out images (orange): medians with interquartile range, individual images in the background. Images that escape at the lower end of the bisection interval are drawn at that end, so the held-out medians for N≤500 are upper bounds. (b) Fraction of generated samples that are copies (orange, left axis; when no copy is found the point is drawn at the 95% upper bound) and AUC of σc between training and held-out images (blue, right axis). At N=2000 sampling no longer detects memorization while σc still does.
Figure 12: Paired comparison on three images that are never sampled. One added image per source, chosen at the median of its group. Each row shows Mσ(x) , one denoising step applied to the clean image at noise level σ , in the model trained on the image (top, blue) and in the model not trained on it (bottom, grey), together with the relative displacement ∥Mσ(x)−x∥/∥x∥ whose crossing of r defines σc . The model not trained on the image moves it towards what it has learned (blur, generic texture) at a much smaller σ .
Figure 13: Examples removed when filtering the memorization dataset. (a) Each row shows a group of duplicate examples, of which we retain only one. (b) The first image in each row is the stored training image, followed by three generations from its training caption. We remove cases in which the caption does not reproduce the stored image.
Figure 14: The full (σ,k) grid behind Figure 1 : the critical scale on Stable Diffusion v1.4, for the same memorized caption–image pair (left) and control pair (middle), each run with its own caption at guidance scale 1 . Each row iterates Mσ from the image x at one noise scale σ ; a framed iterate Mσk(x) is still within r∥x∥ of x . At small σ the model keeps returning x , which therefore sits in a basin of its own; at large σ the orbit drifts away towards content shared with other data. The critical scale σc (dashed line) separates the two regimes: the memorized image withstands far more noise than the control before it escapes. Right: the drift ∥Mσk(x)−x∥/∥x∥ along the orbit, one curve per scale coloured by σ ; the thick curve is a scale above σc and the dot marks where it crosses r . Here K=50 and r=0.5 .
Figure 15: Ablation on the number of iterations K . Detection performance as a function of K for escape radii r=0.1 and r=0.25 . The first four columns report AUC for the different critical-scale scores; the final column reports the mean critical scale (solid line) ± one standard deviation (shaded region). Performance is stable around the primary choice K=16,r=0.25 .
Figure 16: Ablation on the escape radius r . Detection performance (AUC) as a function of r using K=16 . For this ablation, σc is approximated over a logarithmic grid of 24 noise levels rather than estimated by log-space bisection. Performance is stable around the primary choice K=16,r=0.25 (global measures), r=0.1 (local ∣Δlogσc∣ ).
Figure 17: Critical-scale score distributions across evaluation groups. Histograms of the scores used to separate the different groups. Left: Local ∣Δlogσc∣ for MV, TV, and control examples ( K=16 , r=0.1 ). Center-Right: Signed Δlogσc for MV and TV pairs with their original and swapped captions ( K=16 , r=0.25 ). Swapped pairs frequently have negative gaps.
Figure 18: Additional examples of localized memorization. Each panel shows a TV training image, generations from its caption, and the corresponding local ∣Δlogσcu∣ map. Large gaps align with regions reproduced consistently across generations, whereas variable regions generally have gaps near zero.
Figure 19: Near-zero gaps are ambiguous. The two examples have similar mismatched-caption gaps but different absolute retention. Top: Both branches retain the image, and unconditional reconstruction succeeds. Bottom: Neither branch retains the image, and unconditional reconstruction fails.
Figure 20: Additional examples with negative mismatched-caption gaps. Both examples exhibit unconditional retention and successful inpainting. The more negative gap in the top example corresponds to more faithful reconstructions than the less negative gap in the bottom example.
Figure 21: Trajectory and critical scale MV vs control. Top: critical scale on for a totally memorized caption–image pair ( left ) and a control pair ( right ), each run with its own caption at guidance scale 1 . Each row iterates Mσ from the image x at one noise scale σ ; a framed iterate Mσk(x) is still within r∥x∥ of x . At small σ the model keeps returning x , which therefore sits in a basin of its own; at large σ the orbit drifts away towards content shared with other data. The critical scale σc (dashed colored line) separates the two regimes: the memorized image withstands far more noise than the control before it escapes. Bottom: the drift ∥Mσk(x)−x∥/∥x∥ along the orbit, one curve per scale coloured by σ ; the thick curve is a scale above σc and the dot marks where it crosses r .
Figure 22: Trajectory and critical scale TV vs control. Top: critical scale on for a partially memorized caption–image pair ( left ) and a control pair ( right ), each run with its own caption at guidance scale 1 . Each row iterates Mσ from the image x at one noise scale σ ; a framed iterate Mσk(x) is still within r∥x∥ of x . At small σ the model keeps returning x , which therefore sits in a basin of its own; at large σ the orbit drifts away towards content shared with other data. The critical scale σc (dashed colored line) separates the two regimes: the memorized image withstands far more noise than the control before it escapes. Bottom: the drift ∥Mσk(x)−x∥/∥x∥ along the orbit, one curve per scale coloured by σ ; the thick curve is a scale above σc and the dot marks where it crosses r .
Figure 23: Trajectory and critical scale MV vs TV. Top: critical scale on for a totally memorized caption–image pair ( left ) and a partially memorized pair ( right ), each run with its own caption at guidance scale 1 . Each row iterates Mσ from the image x at one noise scale σ ; a framed iterate Mσk(x) is still within r∥x∥ of x . At small σ the model keeps returning x , which therefore sits in a basin of its own; at large σ the orbit drifts away towards content shared with other data. The critical scale σc (dashed colored line) separates the two regimes: the memorized image withstands far more noise than the control before it escapes. Bottom: the drift ∥Mσk(x)−x∥/∥x∥ along the orbit, one curve per scale coloured by σ ; the thick curve is a scale above σc and the dot marks where it crosses r .
Figure 24: Δlogσc for MV and control images. The top row shows an MV image–caption pair and the bottom row a control. The left and center panels show fixed-scale trajectories across noise levels under training-caption and unconditional conditioning, respectively. The right panels plot maxk≤K∣∣Mσk(x)−x∣∣/∣∣x∣∣ for both branches. Their crossings with the dotted threshold r define the conditional and unconditional critical scales, whose log difference is Δlogσc . The MV pair has a large gap because its caption substantially increases retention, whereas the control has nearly equal critical scales and a gap close to zero.
Figure 25: Fixed-scale trajectories across noise levels. Each comparison shows the iterate MσK(x) across noise scales (top) and the maximum normalized displacement maxk≤K∣Mσk(x)−x∣/∣x∣ as a function of σ (bottom). The dotted line marks the escape radius r ; its crossing determines σc . The upper panel compares an MV, a TV, and a control image, ordered from longest to shortest retention. The lower panel compares two TV images with different amounts of memorized content. The image with more memorized content is retained longer.
May 28, 2026·Marta Aparicio Rodriguez, Anastasia Borovykh, Grigorios A. Pavliotis +1Diffusion ModelsMemorization
Department of Mathematics, Imperial College London, UK · ML Lab, Capital Fund Management, France · Department of Physics, École Polytechnique Fédérale de Lausanne (EPFL), Switzerland