Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference speed due to the numerous iterative passes of Diffusion Transformers. To reduce the exhaustive compute, recent works resort to the feature caching and reusing scheme that skips network evaluations at selected diffusion steps by using cached features in previous steps. However, their preliminary design solely relies on local approximation, causing errors to grow rapidly with large skips and leading to degraded sample quality at high speedups. In this work, we propose spectral diffusion feature forecaster (Spectrum), a training-free approach that enables global, long-range feature reuse with tightly controlled error. In particular, we view the latent features of the denoiser as functions over time and approximate them with Chebyshev polynomials. Specifically, we fit the coefficient for each basis via ridge regression, which is then leveraged to forecast features at multiple future diffusion steps. We theoretically reveal that our approach admits more favorable long-horizon behavior and yields an error bound that does not compound with the step size. Extensive experiments on various state-of-the-art image and video diffusion models consistently verify the superiority of our approach. Notably, we achieve up to 4.79× speedup on FLUX.1 and 4.67× speedup on Wan2.1-14B, while maintaining much higher sample quality compared with the baselines.
Figures & tables
FLUX.1 [ 17 ]
Stable Diffusion 3.5-Large [ 4 ]
Acceleration
Quality
Image Reward ↑
CLIP ↑
Acceleration
Quality
Image Rewad ↑
CLIP ↑
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
50 steps †
26.03
1.00
-
-
-
1.00
27.49
25.05
1.00
-
-
-
1.05
28.66
25 steps
13.23
1.97
18.11
0.752
0.322
1.01
27.58
12.67
1.98
12.41
0.593
0.464
1.02
28.78
15 steps
8.04
3.24
15.77
0.673
0.432
1.00
27.56
7.71
3.25
10.51
0.453
0.595
0.85
28.75
FORA (N=4) [ 39 ]
8.40
3.19
15.15
0.651
0.454
0.97
27.55
8.63
2.90
9.54
0.437
0.606
0.27
27.03
Table 1 : Benchmark results of text-to-image generation task on DrawBench with Flux and Stable Diffusion 3.5-Large. We use 50 steps as the reference ( † ). Our Spectrum achieves higher speedup while maintaining better sample quality across two speedup scenarios consistently.
Figure 2 : Qualitative comparison on text-to-image generation using FLUX.1. Spectrum aligns consistently with the 50-step reference while accelerating it by a factor of 4.79 × . Other baselines show noticeable degradation in color and prompt consistency.
Wan2.1-14B [ 50 ]
HunyuanVideo [ 16 ]
Acceleration
Quality
VBench Quality ↑
Acceleration
Quality
VBench Quality ↑
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
50 steps †
486.12
1.00
-
-
-
83.15
378.30
1.00
-
-
-
84.61
25 steps
246.30
1.97
14.94
0.459
0.481
81.74
193.19
1.96
19.64
0.673
0.376
84.78
15 steps
150.50
3.23
13.92
0.397
0.546
80.77
119.58
3.16
17.40
0.620
0.459
83.21
FORA (N=4) [ 39 ]
155.80
3.12
14.47
0.426
0.532
80.53
122.42
3.09
18.24
0.663
0.427
83.34
Table 2 : Benchmark results of text-to-video generation task on VBench with Wan2.1-14B and HunyuanVideo. We use 50 steps as the reference ( † ). Our Spectrum achieves higher speedup while maintaining better sample quality across two speedup scenarios consistently.
Figure 3 : Qualitative comparisons on HunyuanVideo. Spectrum achieves higher sample fidelity while delivering more speedup.
Figure 4 : Qualitative comparison on text-to-video generation using Wan2.1-14B. Spectrum aligns consistently with the high-quality 50-step reference using only 14 network evaluations, while TaylorSeer is slower and exhibits noticeable artifacts on character and background.
Figure 5 : Ablation study on the regularization weight λ .
Stable Diffusion 3.5-Large [ 4 ]
Adaptive
PSNR ↑
SSIM ↑
LPIPS ↓
Image Reward ↑
Taylor (N=8)
×
10.40
0.460
0.594
-0.22
Taylor (α=3.0)
✓
13.25
0.585
0.454
0.65
Spectrum (N=8)
×
11.34
0.501
0.552
0.23
Spectrum (α=3.0)
✓
15.68
0.620
0.430
0.82
Wan2.1-14B [ 50 ]
Table 3 : Ablation study of adaptive scheduling on SD3.5 and Wan.
Figure 6 : Ablation on the degree of Chebyshev polynomials M .
Last only
Latency(s) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Image Reward ↑
Taylor
×
7.36
17.93
0.718
0.375
0.99
Taylor
✓
5.45
18.61
0.738
0.356
0.99
Spectrum
×
15.44
19.34
0.725
0.362
1.02
Spectrum
✓
5.59
19.66
0.741
0.341
1.03
Table 4 : Ablation study of the last-block-only caching strategy on FLUX.1 with N=8 for both Taylor and Spectrum .
Diffusion step
10
20
30
40
50
Taylor ( N=8 )
0.0121
0.0303
0.0629
0.1226
0.2510
Spectrum ( α=3.0 )
0.0040
0.0164
0.0358
0.0742
0.1674
Table 5: RMSE between predicted latents and oracle on Wan2.1.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Reference
N
W
α
NFE
FORA (N=4)
Table 1 , 2
4
1
0.0
13
ToCa (N=4)
Table 1 , 2
4
3
0.0
14
TaylorSeer (N=4)
Table 1 , 2
4
5
0.0
16
Spectrum (α=0.75)
Table 1 , 2
2
5
0.75
14
FORA (N=6)
Table 1 , 2
6
1
0.0
9
ToCa (N=6)
Table 1 , 2
6
3
0.0
10
Appendix
Table 6: Detailed specification on the scheduler and the total number of network evaluations (NFE) for all methods. “Reference” refers to the tables in the main paper where the corresponding method was mentioned.
Figure 7 : Additional qualitative comparison on text-to-video generation using Wan2.1-14B and HunyuanVideo.
FLUX.1 [ 17 ]
Acceleration
Quality
Image Reward ↑
CLIP ↑
Latency(s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
50 steps †
26.46
1.00
-
-
-
1.13
25.92
FORA (N=6) [ 39 ]
6.37
4.16
13.94
0.604
0.523
1.04
26.10
ToCa (N=6,R=0.9) [ 58 ]
13.51
1.96
16.14
0.652
0.462
1.10
26.10
TeaCache (δ=0.8) [ 21 ]
6.73
3.94
16.07
0.664
0.436
1.03
25.88
Appendix
Table 7 : Benchmark results of text-to-image generation on COCO2017 [ 19 ] using FLUX.1. We use 50 steps as the reference ( † ). Our Spectrum achieves higher speedup while maintaining better sample quality.
Figure 8 : Qualitative comparison on text-to-image generation using Stable Diffusion 3.5-Large
Figure 9 : Additional qualitative comparison on text-to-image generation using FLUX.1.
Figure 16
Figure 11 : More text-to-video generation samples on HunyuanVideo with Spectrum . Samples were generated using only 14 network evaluations , leading to a significant speedup of 3.5 × without quality degradation.
Jul 30, 2026·Hanshuai Cui, Zhiqing Tang, Zhi Yao +3DrafterCorrection
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China