We propose Hybrid-Basis Feature Forecasting (HybridFF), a training-free, plug-and-play framework for accelerating diffusion sampling. To capture local smoothness, long-range trends, and complex non-monotonic variations when modeling feature evolution, HybridFF first estimates coefficients using moving least squares (MLS) for each of multiple complementary basis families and then combines the corresponding predictors using fusion weights. In addition to the choice of basis functions, the fusion weights also play a critical role. We introduce two strategies to balance quality and speedup. HybridFF (Fixed) prioritizes efficiency with model-specific fusion weights calibrated on a small set and held constant during inference. HybridFF (Adaptive) updates the fusion weights online using branch reliability scores computed from an exponential moving average of full-step prediction errors, improving prediction fidelity and generation quality under aggressive caching while retaining substantial acceleration. Experiments across DiT-XL/2, FLUX.1-dev, SD3.5-Large, and HunyuanVideo demonstrate a favorable speedup--quality trade-off over representative single-basis forecasters and caching baselines.
Figures & tables
Figure 1: HybridFF results: HunyuanVideo (top row, 3.85× ) and FLUX.1-dev (bottom two rows, 4.37× ) at NFE = 10.
Figure 2: DiT-XL/2 feature-trajectory analysis. (a) Intermediate features follow structured trajectories along denoising steps, and full-step observations can be fitted by polynomial bases to forecast skipped-step features. (b) Hybrid-basis polynomial forecasting gives lower feature-prediction error than single-basis predictors, motivating HybridFF’s coefficient-level fusion.
Figure 3: Overview of the proposed hybrid-basis forecasting pipeline. The pipeline proceeds in three stages: (1) full steps compute and cache exact intermediate features; (2) predict steps fit Taylor, Hermite, and Chebyshev predictors from the same normalized history; (3) the branch coefficient vectors are fused by simplex weights and used to reconstruct the de-normalized target block feature.
NFE
Method
SSIM ↑
PSNR ↑
LPIPS ↓
FID ↓
IS ↑
sFID ↓
Recall ↑
Precision ↑
Latency (s) ↓
Speedup ↑
50
Baseline (50-step)
–
–
–
2.276
240.855
4.258
0.592
0.802
0.971
1.00 ×
14
TaylorSeer
0.927
30.710
0.042
2.528
235.309
5.121
0.589
0.803
0.568
1.71 ×
14
HiCache
0.927
30.811
0.041
2.518
235.894
5.165
0.592
0.801
0.531
1.83 ×
14
Spectrum
0.914
29.233
0.052
2.604
230.968
4.577
0.596
0.794
0.542
1.79 ×
14
HybridFF (Fixed)
0.942
31.982
0.032
2.467
236.325
4.838
0.593
0.800
0.520
1.87 ×
14
HybridFF (Adaptive)
0.945
32.053
0.031
2.416
236.801
4.762
0.591
0.801
0.527
1.84 ×
Table 1: Comparison on class-conditional image generation with DiT-XL/2. Bold and underline mark the best and second-best entries in each block, respectively.
Figure 4: Comparison on class-conditional generation with DiT-XL/2 at NFE = 10. F/A denote fixed/adaptive fusion weights.
Figure 5: Comparison on text-to-image generation with SD3.5-Large at NFE = 14. F/A denote fixed/adaptive fusion weights.
NFE
Method
SSIM ↑
PSNR ↑
LPIPS ↓
VBench Quality ↑
Latency (s) ↓
Speedup ↑
50
Baseline (50-step)
–
–
–
80.958
351.227
1.00 ×
14
TaylorSeer
0.794
24.568
0.182
79.576
140.757
2.50 ×
14
HiCache
0.826
24.984
0.165
80.292
138.094
2.54 ×
14
Spectrum
0.834
26.237
0.140
80.259
128.548
2.73 ×
14
HybridFF (Fixed)
0.853
26.400
0.124
80.343
123.702
2.84 ×
14
HybridFF (Adaptive)
0.867
26.809
0.113
80.614
126.076
2.79 ×
Table 3: Comparison on text-to-video generation with HunyuanVideo. Bold and underline mark the best and second-best entries in each block, respectively.
Figure 6: Comparison on text-to-video generation with HunyuanVideo at NFE = 10. F/A denote fixed/adaptive fusion weights.
Figure 7: Ablations on the ridge regularizer and locality bandwidth with DiT-XL/2 at NFE = 10.
Config
SSIM ↑
PSNR ↑
LPIPS ↓
FID ↓
IS ↑
sFID ↓
Latency (s) ↓
Speedup ↑
D1W4
0.636
21.117
0.283
9.449
165.282
16.268
0.410
2.37 ×
D1W5
0.579
20.179
0.348
16.084
125.705
23.489
0.419
2.32 ×
D1W6
0.538
19.496
0.400
24.096
95.156
31.739
0.426
2.28 ×
D1W7
0.513
19.019
0.437
31.397
76.817
39.063
0.433
2.24 ×
D1W8
0.499
18.752
0.457
36.362
67.100
43.727
0.442
2.20 ×
D2W4
0.800
24.029
0.129
3.179
224.385
6.757
0.414
2.35 ×
Table 4: Ablation on the MLS polynomial degree D and history window W with DiT-XL/2 at NFE = 10.
Figure 8: SSIM landscape of HybridFF over the basis-weight simplex on a small FLUX.1-dev calibration set. The optimum (yellow star) lies inside the simplex at (π(Chebyshev),π(Hermite),π(Taylor))=(0.6,0.1,0.3) .
Figure 9: Additional qualitative comparisons on FLUX.1-dev at NFE = 10. The left and right panels show two prompt sets. F/A denote fixed/adaptive fusion weights.
Figure 10: Additional qualitative comparison on HunyuanVideo at NFE = 10. F/A denote fixed/adaptive fusion weights.
Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference speed due to the numerous iterative passes of Diffusion Transformers. To reduce the exhaustive compute, recent works resort to the feature caching and reusing scheme that skips network evaluations at selected diffusion steps by using cached features in previous steps. However, their preliminary design solely relies on local approximation, causing errors to grow rapidly with large skips and leading to degraded sample quality at high speedups. In this work, we propose spectral diffusion feature forecaster (Spectrum), a training-free approach that enables global, long-range feature reuse with tightly controlled error. In particular, we view the latent features of the denoiser as functions over time and approximate them with Chebyshev polynomials. Specifically, we fit the coefficient for each basis via ridge regression, which is then leveraged to forecast features at multiple future diffusion steps. We theoretically reveal that our approach admits more favorable long-horizon behavior and yields an error bound that does not compound with the step size. Extensive experiments on various state-of-the-art image and video diffusion models consistently verify the superiority of our approach. Notably, we achieve up to 4.79× speedup on FLUX.1 and 4.67× speedup on Wan2.1-14B, while maintaining much higher sample quality compared with the baselines.
Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on designing stronger forecasters. Yet forecast error varies sharply across steps, and open-loop caches trust the forecast in full at every skipped step. This fixed trust is what breaks as acceleration turns aggressive. The missing question is not only how to forecast better, but when and how much to trust a forecast. We show that reliability can be observed from the cache itself. Two forecasts agree where the feature trajectory is smooth, and they diverge where prediction turns hard. Their disagreement is a cheap runtime signal, and it costs no extra denoiser evaluation. Based on this signal, we introduce RACER, a training-free closed-loop controller with two responses. It continuously shrinks uncertain forecasts toward the last computed feature. At the riskiest steps, RACER refreshes the feature and repays the added evaluation by skipping a later scheduled one. We derive a deterministic error bound for the shrinkage and empirically evaluate its validity and tightness across acceleration regimes. At the same number of denoiser evaluations, RACER improves the strongest open-loop baseline across SD3.5-Large, FLUX.1-dev, Wan2.1-14B, and HunyuanVideo on DrawBench, VBench, and COCO. On SD3.5, we further show that RACER samples faster at equal quality. RACER generalizes across forecasting designs as well. For example, it recovers much of the quality lost on a Taylor base. These results show that reliable diffusion acceleration also depends on how forecasts are used. Code is available at https://github.com/LiZaiyuan0619/RACER
Yanchao Li, Jiaqing Xie, Ben Gao +6
1Nanjing University · 2Shanghai Artificial Intelligence Laboratory · 3North University of China +2
To address the high sampling cost of Diffusion Transformers (DiTs), feature caching offers a training-free acceleration method. However, existing methods rely on hand-crafted forecasting formulas that fail under aggressive skipping. We propose L2P (Learnable Linear Predictor), a simple data-driven caching framework that replaces fixed coefficients with learnable per-timestep weights. Rapidly trained in ~20 seconds on a single GPU, L2P accurately reconstructs current features from past trajectories. L2P significantly outperforms existing baselines: it achieves a 4.55x FLOPs reduction and 4.15x latency speedup on FLUX.1-dev, and maintains high visual fidelity under up to 7.18x acceleration on Qwen-Image models, where prior methods show noticeable quality degradation. Our results show learning linear predictors is highly effective for efficient DiT inference. Code is available at https://github.com/Aredstone/L2P-Cache.
Zhirong Shen, Rui Huang, Jiacheng Liu +6
Shanghai Jiao Tong University · University of Electronic Science and Technology of China · Shandong University +2