Diffusion samplers can reduce computation by generating selected spectral coordinates and filling the remaining directions with noise. How many directions must they retain? We study this question for data with power-law covariance spectra. For Gaussian data compared to a smoothed target, we prove matching bounds on the required number of retained directions, provided that the ambient dimension is sufficiently large. The truncation error depends on the combined Wiener gains of the omitted directions, regardless of the accuracy of the sampler on the retained coordinates. Keeping only directions whose signal exceeds the output noise level can therefore leave a non-vanishing error: many individually weak directions remain significant in aggregate. Combining this characterization with a diffusion convergence bound yields sufficient sampling-step complexity under exact scores. The upper bounds also extend to estimated principal components and, componentwise, to Gaussian mixtures. The practical prescription is to select the retained subspace using an aggregate spectral-tail error budget, then to choose the diffusion noise level accordingly.
Figures & tables
Figure 1: Roadmap of the results. The error decomposition of Section 2 splits the problem into a truncation error, computed exactly for Gaussian data in Section 3 , and a sampling error on the retained coordinates, bounded for general data in Section 4 ; the two combine in Theorem 4 . Grey boxes are supporting results in the appendix; bold boxes are the main theorems.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Patch
d
Vectors
Sources
Window [ℓ,h]
b^/d
CIFAR-10
8 × 8
192
6000/2000/2000
6000/2000/2000
[8,48]
0.09
CIFAR-10
16 × 16
768
6000/2000/2000
6000/2000/2000
[8,192]
0.10
CIFAR-100
8 × 8
192
6000/2000/2000
6000/2000/2000
[8,48]
0.17
CIFAR-100
16 × 16
768
6000/2000/2000
6000/2000/2000
[8,192]
0.10
DIV2K
8 × 8
192
1600/1280/1600
100/80/100
[8,48]
0.17
DIV2K
16 × 16
768
1600/1280/1600
100/80/100
[8,192]
0.10
Appendix
Table 1: Data used for the check. Vectors and independent sources (images; artists for FMA( Defferrard et al. (2017) )) per train/val/test split; fitting window [ℓ,h] ; fitted breakpoint relative to dimension, b^/d . † Audio, outside the scope of the paper.
β^
Test log-MSE
Broken power law
Dataset
Patch
train
val
test
power
exp.
broken
β^1
β^2
b^
CIFAR-10
8 × 8
1.82
1.85
1.83
0.009
0.022
0.004
1.28
2.04
16
CIFAR-10
16 × 16
1.57
1.60
1.60
0.018
0.161
0.003
1.42
2.05
74
CIFAR-100
8 × 8
1.84
1.83
1.83
0.009
0.031
0.005
1.74
2.43
34
CIFAR-100
16 × 16
1.56
1.57
1.58
0.016
0.167
0.004
1.43
1.96
74
DIV2K
8 × 8
1.65
1.68
1.70
0.011
0.029
0.010
1.51
2.43
34
Appendix
Table 2: Model comparison on the common window. β^ : power-law exponent fitted separately on each split. Test log-MSE of the power, exponential and broken power laws (smallest in bold). β^1,β^2,b^ : broken-law head and tail exponents and breakpoint (training split).
Window [ℓ,h]
Bulk [8,m]
Dataset
Patch
train
val
test
broken law
m
m/d
train
val
test
U
CIFAR-10
8 × 8
1.66
1.68
1.55
1.47
89
0.46
3.21
3.20
3.28
1.0
CIFAR-10
16 × 16
1.63
1.70
1.79
1.57
295
0.38
3.25
3.18
3.30
1.0
CIFAR-100
8 × 8
1.62
1.64
1.53
1.23
96
0.50
3.11
3.29
3.05
1.0
CIFAR-100
16 × 16
1.55
1.61
1.68
1.46
306
0.40
3.12
3.05
3.21
1.1
DIV2K
8 × 8
1.51
1.37
1.39
1.32
76
0.40
2.93
2.21
2.57
5.1
Appendix
Table 3: Envelope ratios C/c of Assumption 1 with the training exponent β^ . Window: observed ratio per split, and the ratio implied by the fitted broken power law alone (Proposition 5 ). Bulk: end m on the test split, and the ratio on [8,m] per split. U : factor by which including every resolved rank (in particular i<8 ) raises the upper constant, maximum over splits.
Figure 2: Sample covariance spectra of image patches (train and test splits; FMA audio in the last two panels). Shading: fitting window [ℓ,h] . Dashed: power-law fit; dotted: broken power-law fit; thin vertical line: fitted breakpoint b^ . Ranks beyond the sample rank are not shown.
Figure 3: Compensated spectra σ^i2iβ^ , normalised by the training-window minimum c , for the three disjoint splits; β^ is fitted on the training split only. An exact power law is a horizontal line. Solid horizontal lines: training envelope [c,C] on the window (shaded). Dashed line: c/2 , which defines the bulk end m (dotted vertical line, test split). Values below 0.03 are clipped.
Figure 4: Dependence on patch size (image datasets). (a) Single power-law exponent; bars span the estimates on the three disjoint splits. (b) Broken-law head and tail exponents. (c) Envelope ratio on the test split, over the fitting window and over the bulk [8,m] . Dashed horizontal lines in (a,b): β=1 .
β
points
used
TV/∥(gi)i>K∥2
TV/Pinsker
1.25
4779
3242
0.276-0.290
0.78-0.82
1.5
4757
3763
0.276-0.290
0.77-0.81
2
4725
3838
0.276-0.302
0.77-0.81
Appendix
Table 4: Theorem 1 on all truncation curves. Points: evaluated (η,K) pairs; ratio range over the points with 10−5<∥(gi)i>K∥2<21 and a Monte Carlo signal-to-noise ratio above 5 . No point violates the bounds 1001GK≤TV≤23GK or the Pinsker bound.
Figure 5: Theorem 1. (a) Truncation error TV(q~Kη,qη) against K for β=1.5 and five resolutions (solid), with the explicit Pinsker bound of Theorem 1(a) (dashed). (b) All points for all β and η against the Wiener-gain tail norm; dashed lines are the constants 3/2 and 1/100 of Theorem 1(a). (c) Their ratio, which is constant until the error saturates at one.
Asymptotic points
K⋆≥16kη
β
2β−12
1/ε
Csp/η2
n
1/ε
Csp/η2
Kε/K⋆
K⋆/Knec
xK⋆+1K⋆/ε
1.25
1.333
1.347±0.005
1.340±0.007
32
1.337
1.333
1.32-1.51
521-528
4.24-4.57
1.5
1.000
1.005±0.003
1.001±0.003
34
0.997
0.999
1.24-1.33
110-114
4.77-5.21
2
0.667
0.670±0.002
0.669±0.001
31
0.668
0.668
1.15-1.21
26-26
5.60-6.18
Appendix
Table 5: Theorem 2. Fitted exponents of 1/ε and of Csp/η2 in K⋆ (joint least squares, 95% intervals), over asymptotic points ( 4kη≤K⋆≤d/8 ) and over K⋆≥16kη ; sharpness of the sufficient choice Kε of Theorem 2(a) and of the necessary bound Knec of Theorem 2(b) ( ε<1/200 ); and the signal-to-noise ratio of the first discarded direction at the required cut, xK⋆+1K⋆/ε (Theorem 2(c)).
Figure 6: Theorem 2. (a) Required dimension against accuracy for β=1.5 and five resolutions (markers), with the sufficient choice of Theorem 2(a) (dashed) and the necessary bound of Theorem 2(b) where it applies, ε<1/200 (dotted). (b) Required dimension for all settings against the predicted variable; open markers are outside the asymptotic range. (c) Fitted exponents against the prediction 2/(2β−1) .
Figure 7: (a) Signal-to-noise ratio of the first discarded direction at the required cut, scaled by K⋆/ε (Theorem 2(c)), for all asymptotic points. (b) Error of truncating exactly at the noise level, K=kη , against kη (Corollary 1); dotted lines, the guaranteed lower bound. (c) Required dimension for four ambient dimensions; crosses mark 2K⋆>d .
Rule
Meets
K/K⋆ : median, range
Noise level, K=kη
1%
0.034
0.000-1.33
95% variance
36%
0.35
0.002-2,047
99% variance
64%
4.10
0.008-21,479
99.9% variance
90%
30
0.078-47,499
Theorem 2(a), K2ε(η)
100%
1.26
1.04-2.00
Tail budget, Theorem 1(a)
100%
1.25
1.00-1.60
Appendix
Table 6: Selection rules over all (β,kη,ε) : share of cases meeting TV≤ε (upper end of the Monte Carlo interval), and retained dimension relative to the minimum K⋆ . Variance rules keep the smallest K explaining the stated share of ∑iσi2 ; the Theorem 2(a) rule is the smallest K whose explicit bound (11) is at most ε ; the tail-budget rule is the smallest K whose Pinsker bound in Theorem 1(a) is at most ε .
Figure 8: Selection rules over all (β,kη,ε) . (a) Achieved truncation error relative to the target (values below 10−3 are drawn at 10−3 ). (b) Retained dimension relative to the minimum.
ε
K
σ2
T
trunc.
N
total
TV(p,qK)
0.1
320
1.7×10−4
17.0
0.039
32
0.074
0.78
0.05
640
6.2×10−5
19.4
0.020
64
0.027
0.63
0.02
1600
1.6×10−5
22.6
0.008
128
0.010
0.57
Appendix
Table 7: Exact-score sampler in the settings of Theorem 4 ( β=1.5 , kη=16 , Wiener-gain rates). Trunc.: truncation error TV(q~Kη,qη) ; N : fewest grid steps with total TV≤ε ; total: TV(p~η,qη) at that N ; TV(p,qK) : unsmoothed sampling error at that N .
Figure 9: (a) Samplers with perturbed retained variances: total error (markers), truncation error (dashed), and the upper bound of (10) (dotted). (b) Exact-score sampler in the settings of Theorem 4: total error (solid), truncation error (dashed), unsmoothed sampling error (dotted). (c) Steps needed for unsmoothed sampling error ≤ε/2 against K , for Wiener-gain and equal rates; the dashed line has the slope of the bound of Theorem 3.
CIFAR-10
CIFAR-100
M
val.
test
EM iter.
val.
test
EM iter.
1
0.000
0.000
-
0.000
0.000
-
2
0.294
0.293
16
0.333
0.333
17
4
0.416
0.420
58
0.453
0.452
50
8
0.488
0.493
100 ‡
0.550
0.549
72
16
0.549
0.558
100 ‡
0.611
0.611
100 ‡
Appendix
Table 8: Held-out log-likelihood gain over a single Gaussian (bits per dimension) and number of EM iterations; ‡ iteration cap reached, ( ⋅ ) number of reseeded components.
CIFAR-10
CIFAR-100
pooled β^
1.82
1.84
pooled C^sp/c^sp
1.60
1.65
components: β^m , weighted median
2.04
2.00
weighted 10%-90% quant.
1.70-2.27
1.72-2.25
range over components
1.57-2.57
1.56-2.57
weight with β^m>1
1.00
1.00
Appendix
Table 9: Power-law fits on ranks 8 - 48 for the M=32 mixtures. Weighted quantiles use the mixture weights πm .
Figure 10: (a) Held-out log-likelihood gain of the patch mixtures over a single Gaussian. (b,c) Eigenvalues of the 32 component covariances (blue) and of the pooled covariance Σ (black). The shaded band is the rank window [8,48] used for the power-law fits; the dashed line has the fitted pooled slope, shifted vertically for visibility.
Figure 11: Truncation error on CIFAR-10 patches ( M=32 mixture, shared basis); lighter colours are coarser resolutions kη . (a) Mixture error (solid) and the error predicted by Theorem 1 for the moment-matched Gaussian (dashed). (b) Ratio of the two. (c) Error achieved by six truncation rules (Appendix I.3.5 ), relative to the target ε∈{0.3,0.2,0.1,0.05,0.02,0.01} , with the percentage of the 30 (resolution, ε ) pairs that meet the target: K=kη ; the smallest K keeping 95 , 99 or 99.9% of the variance; the smallest K at which the explicit bound of Theorem 1(a) for N(0,Σ) is at most ε (T1(a)); and the smallest K at which bound (11) with fitted (β^,C^sp) is at most ε (T2(a)). Triangles mark T2(a) choices with K=d , i.e. no truncation. CIFAR-100 is in Figure 13 .
ε=0.1
ε=0.01
kη
η2/σ12
TVNG
TV(kη)
K⋆
KG
KT1
KT2
Kˉcomp
K⋆
KG
KT1
KT2
Kˉcomp
2
0.1922
0.039
0.22
6
6
7
9
7.3
35
34
39
46
32.3
4
0.1081
0.064
0.21
10
10
12
13
11.3
47
45
51
71
41.4
8
0.0234
0.199
0.41
31
30
35
40
28.8
84
79
84
192 ∗
67.4
16
0.0102
0.330
0.43
49
46
52
75
41.8
108
97
102 †
192 ∗
82.1
32
0.0025
0.602
0.56
84
78
83 †
192 ∗
66.1
152
128
134 †
192 ∗
107.5
Appendix
Table 10: Retained directions needed on CIFAR-10 patches. η2/σ12 : evaluation level relative to the top pooled eigenvalue. TVNG : distance TV(qη,N(0,Σ+η2I)) between the smoothed mixture and its moment-matched Gaussian. TV(kη) : error when exactly the directions above the evaluation level are retained. K⋆ and KG : smallest K with error at most ε for the mixture and for N(0,Σ) . KT1 , KT2 : rules T1(a) and T2(a) of Appendix I.3.5 ; † misses the target, ∗ no truncation ( K=d ). Kˉcomp : expected number of retained directions, ∑mπmKm , when each component uses its own basis with per-component budget ε , which guarantees error at most ε by Proposition 4 (Appendix I.3.6 ). CIFAR-100 is in Table 11 .
Figure 12: Component-basis truncation. (a) Mixture truncation error against the bound of Proposition 4 for all budgets, resolutions and both datasets. (b) CIFAR-10: expected number of retained directions ∑mπmKm for component bases (solid), and the number K⋆ needed with the shared basis (dashed), as functions of the achieved mixture error.