Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference speed due to the numerous iterative passes of Diffusion Transformers. To reduce the exhaustive compute, recent works resort to the feature caching and reusing scheme that skips network evaluations at selected diffusion steps by using cached features in previous steps. However, their preliminary design solely relies on local approximation, causing errors to grow rapidly with large skips and leading to degraded sample quality at high speedups. In this work, we propose spectral diffusion feature forecaster (Spectrum), a training-free approach that enables global, long-range feature reuse with tightly controlled error. In particular, we view the latent features of the denoiser as functions over time and approximate them with Chebyshev polynomials. Specifically, we fit the coefficient for each basis via ridge regression, which is then leveraged to forecast features at multiple future diffusion steps. We theoretically reveal that our approach admits more favorable long-horizon behavior and yields an error bound that does not compound with the step size. Extensive experiments on various state-of-the-art image and video diffusion models consistently verify the superiority of our approach. Notably, we achieve up to 4.79× speedup on FLUX.1 and 4.67× speedup on Wan2.1-14B, while maintaining much higher sample quality compared with the baselines.
Figures & tables
FLUX.1 [ 17 ]
Stable Diffusion 3.5-Large [ 4 ]
Acceleration
Quality
Image Reward ↑
CLIP ↑
Acceleration
Quality
Image Rewad ↑
CLIP ↑
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
50 steps †
26.03
1.00
-
-
-
1.00
27.49
25.05
1.00
-
-
-
1.05
28.66
25 steps
13.23
1.97
18.11
0.752
0.322
1.01
27.58
12.67
1.98
12.41
0.593
0.464
1.02
28.78
15 steps
8.04
3.24
15.77
0.673
0.432
1.00
27.56
7.71
3.25
10.51
0.453
0.595
0.85
28.75
FORA (N=4) [ 39 ]
8.40
3.19
15.15
0.651
0.454
0.97
27.55
8.63
2.90
9.54
0.437
0.606
0.27
27.03
Table 1 : Benchmark results of text-to-image generation task on DrawBench with Flux and Stable Diffusion 3.5-Large. We use 50 steps as the reference ( † ). Our Spectrum achieves higher speedup while maintaining better sample quality across two speedup scenarios consistently.
Figure 2 : Qualitative comparison on text-to-image generation using FLUX.1. Spectrum aligns consistently with the 50-step reference while accelerating it by a factor of 4.79 × . Other baselines show noticeable degradation in color and prompt consistency.
Wan2.1-14B [ 50 ]
HunyuanVideo [ 16 ]
Acceleration
Quality
VBench Quality ↑
Acceleration
Quality
VBench Quality ↑
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
50 steps †
486.12
1.00
-
-
-
83.15
378.30
1.00
-
-
-
84.61
25 steps
246.30
1.97
14.94
0.459
0.481
81.74
193.19
1.96
19.64
0.673
0.376
84.78
15 steps
150.50
3.23
13.92
0.397
0.546
80.77
119.58
3.16
17.40
0.620
0.459
83.21
FORA (N=4) [ 39 ]
155.80
3.12
14.47
0.426
0.532
80.53
122.42
3.09
18.24
0.663
0.427
83.34
Table 2 : Benchmark results of text-to-video generation task on VBench with Wan2.1-14B and HunyuanVideo. We use 50 steps as the reference ( † ). Our Spectrum achieves higher speedup while maintaining better sample quality across two speedup scenarios consistently.
Figure 3 : Qualitative comparisons on HunyuanVideo. Spectrum achieves higher sample fidelity while delivering more speedup.
Figure 4 : Qualitative comparison on text-to-video generation using Wan2.1-14B. Spectrum aligns consistently with the high-quality 50-step reference using only 14 network evaluations, while TaylorSeer is slower and exhibits noticeable artifacts on character and background.
Figure 5 : Ablation study on the regularization weight λ .
Stable Diffusion 3.5-Large [ 4 ]
Adaptive
PSNR ↑
SSIM ↑
LPIPS ↓
Image Reward ↑
Taylor (N=8)
×
10.40
0.460
0.594
-0.22
Taylor (α=3.0)
✓
13.25
0.585
0.454
0.65
Spectrum (N=8)
×
11.34
0.501
0.552
0.23
Spectrum (α=3.0)
✓
15.68
0.620
0.430
0.82
Wan2.1-14B [ 50 ]
Table 3 : Ablation study of adaptive scheduling on SD3.5 and Wan.
Figure 6 : Ablation on the degree of Chebyshev polynomials M .
Last only
Latency(s) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Image Reward ↑
Taylor
×
7.36
17.93
0.718
0.375
0.99
Taylor
✓
5.45
18.61
0.738
0.356
0.99
Spectrum
×
15.44
19.34
0.725
0.362
1.02
Spectrum
✓
5.59
19.66
0.741
0.341
1.03
Table 4 : Ablation study of the last-block-only caching strategy on FLUX.1 with N=8 for both Taylor and Spectrum .
Diffusion step
10
20
30
40
50
Taylor ( N=8 )
0.0121
0.0303
0.0629
0.1226
0.2510
Spectrum ( α=3.0 )
0.0040
0.0164
0.0358
0.0742
0.1674
Table 5: RMSE between predicted latents and oracle on Wan2.1.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Reference
N
W
α
NFE
FORA (N=4)
Table 1 , 2
4
1
0.0
13
ToCa (N=4)
Table 1 , 2
4
3
0.0
14
TaylorSeer (N=4)
Table 1 , 2
4
5
0.0
16
Spectrum (α=0.75)
Table 1 , 2
2
5
0.75
14
FORA (N=6)
Table 1 , 2
6
1
0.0
9
ToCa (N=6)
Table 1 , 2
6
3
0.0
10
Appendix
Table 6: Detailed specification on the scheduler and the total number of network evaluations (NFE) for all methods. “Reference” refers to the tables in the main paper where the corresponding method was mentioned.
Figure 7 : Additional qualitative comparison on text-to-video generation using Wan2.1-14B and HunyuanVideo.
FLUX.1 [ 17 ]
Acceleration
Quality
Image Reward ↑
CLIP ↑
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
50 steps †
26.46
1.00
-
-
-
1.13
25.92
FORA (N=6) [ 39 ]
6.37
4.16
13.94
0.604
0.523
1.04
26.10
ToCa (N=6,R=0.9) [ 58 ]
13.51
1.96
16.14
0.652
0.462
1.10
26.10
TeaCache (δ=0.8) [ 21 ]
6.73
3.94
16.07
0.664
0.436
1.03
25.88
Appendix
Table 7 : Benchmark results of text-to-image generation on COCO2017 [ 19 ] using FLUX.1. We use 50 steps as the reference ( † ). Our Spectrum achieves higher speedup while maintaining better sample quality.
Figure 8 : Qualitative comparison on text-to-image generation using Stable Diffusion 3.5-Large
Figure 9 : Additional qualitative comparison on text-to-image generation using FLUX.1.
Figure 16
Figure 11 : More text-to-video generation samples on HunyuanVideo with Spectrum . Samples were generated using only 14 network evaluations , leading to a significant speedup of 3.5 × without quality degradation.
We propose Hybrid-Basis Feature Forecasting (HybridFF), a training-free, plug-and-play framework for accelerating diffusion sampling. To capture local smoothness, long-range trends, and complex non-monotonic variations when modeling feature evolution, HybridFF first estimates coefficients using moving least squares (MLS) for each of multiple complementary basis families and then combines the corresponding predictors using fusion weights. In addition to the choice of basis functions, the fusion weights also play a critical role. We introduce two strategies to balance quality and speedup. HybridFF (Fixed) prioritizes efficiency with model-specific fusion weights calibrated on a small set and held constant during inference. HybridFF (Adaptive) updates the fusion weights online using branch reliability scores computed from an exponential moving average of full-step prediction errors, improving prediction fidelity and generation quality under aggressive caching while retaining substantial acceleration. Experiments across DiT-XL/2, FLUX.1-dev, SD3.5-Large, and HunyuanVideo demonstrate a favorable speedup--quality trade-off over representative single-basis forecasters and caching baselines.
Kai-Liang Cheng, Yuan-Yuan Cheng, Yu-fan Jin +1
University of Science and Technology of China, China
Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer--timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to 6.70× over Vanilla while maintaining competitive output quality.
Hanshuai Cui, Zhiqing Tang, Zhi Yao +3
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China
Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoising process while high-frequency details emerge only in later timesteps. This structure offers a natural opportunity for efficient generation, as high-resolution computation on noise-dominated frequencies is largely redundant. We propose Spectral Progressive Diffusion, a general framework that progressively grows resolution along the denoising trajectory of pretrained diffusion models. To this end, we develop a spectral noise expansion mechanism and derive an optimal resolution schedule from the model's power spectrum. Our framework supports training-free acceleration and a novel fine-tuning recipe that further improves efficiency and quality. We demonstrate significant speedups on state-of-the-art pretrained image and video generation models while preserving visual quality.