Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by high sampling costs. While prior work focuses on training objectives or individual solvers, the broader sampling design problem, specifically solver selection and scheduling, remains largely governed by static heuristics. We propose SDM, a principled, training-free sampling framework that adapts both the numerical solver and the timestep schedule to the intrinsic properties of the diffusion trajectory. By analyzing the PF-ODE dynamics, we show that velocity variation is small in high-noise stages and increases near the data manifold, identifying intervals where solver order is most consequential. In parallel, we introduce an offline-calibrated adaptive scheduling method that explicitly controls the local Wasserstein discretization error and projects the calibrated trajectory to a prescribed NFE budget. We further extend the formulation to a mixed-transition Wasserstein error bound, providing a unified error-propagation view of adaptive scheduling and solver selection within the overall SDM framework. Across standard benchmarks, with extensions to modern ODE samplers, high-resolution synthesis, and text-to-image generation, SDM achieves improved sample quality compared to baseline methods, attaining an FID of 1.93 on CIFAR-10, 2.41 on FFHQ, and 1.98 on AFHQv2, with a reduced number of function evaluations compared to existing samplers. Our code is available at https://github.com/aiimaginglab/sdm.
Figures & tables
Figure 1: Overview of SDM. (Inset) We introduce an adaptive timestep optimization derived from an analytical upper bound on the Wasserstein error. Here, ηi serves as a controllable schedule that explicitly governs the allowable error budget at each step, tightening step sizes where the flow is most sensitive. (Main) Complementing this, SDM adapts the solver order to the trajectory’s geometry. In the early high-noise regime (black), the flow is nearly linear, allowing efficient low-order integration. As the trajectory approaches the data manifold (blue) , the curvature increases, necessitating high-order solvers. These two strategies act as independent and complementary components, jointly improving both sample quality and computational efficiency.
Unconditional
Unconditional
Unconditional
CIFAR-10 32×32
FFHQ 64×64
AFHQv2 64×64
VP
VE
VP
VE
VP
VE
Solver
Schedule
Euler
EDM (ρ=7)
7.61
7.75
4.59
4.76
2.93
3.21
COS ( Williams et al., 2024 )
7.43
7.25
4.29
4.46
2.89
3.10
SDM (Adaptive Scheduling)
6.18
6.48
4.16
4.51
2.71
2.97
Table 1: Quantitative results on unconditional generation for CIFAR-10, FFHQ, and AFHQv2 dataset, measured by FID and NFE. “SDM (Adaptive Scheduling)” denotes the Wasserstein-bounded timestep schedule for the solver indicated in the leftmost column. Rows under “SDM (Adaptive Solver)” use the curvature-based solver selection with the optimized threshold parameter τκ . The best FID within each solver block for a given dataset and parameterization are highlighted in bold , while the overall best FID in each column are highlighted in bold . Overall, the proposed SDM sampler achieves superior performance in both sample quality and computational efficiency over all baselines.
Schedule
FID
CLIP
Uniform
23.41
0.3162
SDM
22.06
0.3151
Table 2: Quantitative results on text-to-image generation with Stable Diffusion 3 ( Left ) and FLUX ( Right ). The baseline timestep schedule, denoted as “Uniform”, includes the shifting term used by each corresponding model. SDM remains achieving the best model performance in FID while remaining competitive in CLIP for recent rectified flow-based text-to-image models.
Figure 2: Empirical validation of trajectory geometry and Wasserstein error allocation. ( Left ) Relative curvature κ^rel as a function of noise level σ for standard benchmarks. The curvature exhibits an approximately linear correlation with noise levels σ in log scale, consistent with the theoretical derivation of second order probability flow ODE. ( Right ) Distribution of local Wasserstein error bound η^i over diffusion timesteps for ImageNet 64×64 . EDM schedules exhibit an initially increasing trend with a subsequent decay, reaching maximum during the intermediate sampling stages. In contrast, SDM schedules allocate a larger portion of the error budget to the early high-noise stages, resulting in improved sample quality.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Timestep Schedule
CIFAR-10
FFHQ
AFHQv2
ImageNet
Grid search
EDM
2×10−4
1×10−4
1×10−3
1×10−4
{2,5,10,20,50,100}×10−5
SDM
2×10−4
1×10−4
2×10−4(VP)
1×10−4
{2,5,10,20,50,100}×10−5
1×10−3(VE)
Appendix
Table 3: Parameter search grid for threshold τκ for step schedule based adaptive solver selection.
Parameter
CIFAR-10
Grid search
Uncond. VP
Uncond. VE
Cond. VP
Cond. VE
ηmin
0.01
0.01
0.01
0.02
0.01, 0.02, 0.03, 0.04, 0.05
ηmax
0.40
0.40
0.40
0.10
0.10, 0.20, 0.30, 0.40, 0.50
p
1.0
1.0
1.0
1.0
0.8, 1.0, 1.2
q
0.1
0.25
0.1
0.25
0.1, 0.25
Appendix
Table 4: Parameter search grid for Wasserstein error tolerance and N -step resampling parameters for CIFAR-10.
Conditional
Conditional
CIFAR-10 32×32
ImageNet 64×64
VP
VE
ADM
Solver
Schedule
Euler
EDM (ρ=7)
7.09
6.75
3.48
COS ( Williams et al., 2024 )
7.05
6.66
3.74
SDM (Adaptive Scheduling)
6.10
5.12
2.32
Appendix
Table 5: Quantitative results on conditional generation for CIFAR-10 32×32 and ImageNet 64×64 , measured by FID and NFE. “SDM (Adaptive Scheduling)” denotes the Wasserstein-bounded timestep schedule for the solver indicated in the leftmost column. Rows under “SDM (Adaptive Solver)” use the curvature-based solver selection with the optimized threshold parameter τκ . The best FID within each solver block for a given dataset and parameterization are highlighted in bold , while the overall best FID in each column are highlighted in bold .
Unconditional
Unconditional
Unconditional
CIFAR-10 32×32
FFHQ 64×64
AFHQv2 64×64
VP
VE
VP
VE
VP
VE
Solver
Schedule
DPM-Solver++(2M)
EDM (ρ=7)
2.57
2.52
2.60
2.71
2.12
2.26
SDM (Adaptive Scheduling)
2.33
2.33
2.59
2.73
2.07
2.19
UniPC
EDM (ρ=7)
2.32
2.28
2.51
2.62
2.07
2.20
Appendix
Table 6: Quantitative results on modern ODE samplers, including DPM-Solver++ and UniPC, for CIFAR-10 32×32 , FFHQ 64×64 , and AFHQv2 64×64 . Overall, SDM schedules achieves superior performance in FID compared to EDM baselines.
Unconditional
Unconditional
FFHQ 64×64
AFHQv2 64×64
VP
VE
VP
VE
Solver
Schedule
Euler
EDM (ρ=7)
20.92
21.06
10.14
10.73
SDM (Adaptive Scheduling)
17.94
19.03
9.73
10.07
Heun
EDM (ρ=7)
5.27
5.44
2.99
3.16
Appendix
Table 7: Quantitative results on 10-step generation for FFHQ 64×64 and AFHQv2 64×64 . Results for the adaptive solver are reported with the optimal threshold parameter τκ . Adaptive scheduling is applied to Euler, Heun, and SDM-based adaptive solvers. The best-performing configurations compared to the corresponding baselines are highlighted in bold , while the overall best results across all methods are highlighted in bold . Overall, the proposed SDM sampler consistently shows superior performance in both sample quality and computational efficiency for 10-step generation.
Unconditional
Unconditional
FFHQ 64×64
AFHQv2 64×64
VP
VE
VP
VE
Solver
Schedule
Euler
EDM (ρ=7)
8.64
8.75
4.61
5.04
SDM (Adaptive Scheduling)
7.37
8.13
4.21
4.73
Heun
EDM (ρ=7)
2.64
2.76
2.08
2.22
Appendix
Table 8: Quantitative results on 20-step generation for FFHQ 64×64 and AFHQv2 64×64 . Results for the adaptive solver are reported with the optimal threshold parameter τκ . Adaptive scheduling is applied to Euler, Heun, and SDM-based adaptive solvers. The best-performing configurations compared to the corresponding baselines are highlighted in bold , while the overall best results across all methods are highlighted in bold . Overall, the proposed SDM sampler consistently shows superior performance in both sample quality and computational efficiency for 20-step generation.
Solver
Schedule
FID
DPM-Solver++(2M)
EDM
2.48
DPM-Solver++(2M)
SDM
2.36
UniPC
EDM
2.45
UniPC
SDM
2.37
Appendix
Table 9: Quantitative results on high-resolution generation for ImageNet 512×512 using the EDM2-M pretrained model. ( Left ) FID comparison under EDM and SDM schedules with fixed solvers, DPM-Solver++(2M) and UniPC, respectively. ( Right ) FID comparison with changes in adaptive solver (DPM-Solver++(1M) → DPM-Solver++(2M)) and adaptive scheduling.
Schedule
FID
CLIP
Uniform
18.29
0.3209
EDM
18.10
0.3238
AYS
17.90
0.3205
OTS ( Xue et al., 2024 )
17.49
0.3219
AutoDiffusion
17.76
0.3182
SDM
17.72
0.3201
Appendix
Table 10: Quantitative results on text-to-image generation with SDXL. Comparison of SDM against standard schedules, including time-uniform, EDM, AYS, OTS ( Xue et al., 2024 ) , and AutoDiffusion. SDM demonstrates competitive performance against recent schedule optimization baselines. Bold and underlined values indicate the best and second-best results, respectively.
Figure 3: Ablation on threshold τκ . FID as a function of the curvature threshold τκ for CIFAR-10 ( 32×32 , solid) and AFHQv2 ( 64×64 , dashed), evaluated under unconditional and conditional settings using the curvature-thresholded adaptive solver. Markers denote the selected τκ values that yield the best model performance for each dataset and training configuration.
Figure 4: Qualitative comparison on AFHQv2 ( 64×64 ) across VP (left) and VE (right) parameterizations. The top row shows EDM (Heun), while the bottom row shows SDM with adaptive solver and scheduling. Each panel shows a 3×5 grid of generated samples; corresponding FID and NFE are reported below.
Figure 5: Qualitative comparison on ImageNet 512×512 using the pretrained EDM2-M model. For each pair of rows, the top and bottom row represents results from EDM and SDM schedules, respectively. Across the examples, the EDM schedule exhibits more noticeable structural inconsistencies in generated objects, whereas SDM produces more coherent object structures. All results are generated using DPM-Solver++(2M) as the numerical solver.
Figure 6: Qualitative comparison of different sampling schedules using SDXL for text-to-image generation. From left to right: Time uniform, EDM, AYS, and SDM. All results are generated using DPM-Solver++(2M) as the numerical solver.
Figure 7: Qualitative comparison of different sampling schedules using Stable Diffusion 3 for text-to-image generation. From left to right: Time uniform and SDM. All results are generated using Euler as the numerical solver.
ECE & CSL University of Illinois Urbana-Champaign · Computer Science Department Carnegie Mellon University · ECE, CSL & NCSA University of Illinois Urbana-Champaign