We propose Hybrid-Basis Feature Forecasting (HybridFF), a training-free, plug-and-play framework for accelerating diffusion sampling. To capture local smoothness, long-range trends, and complex non-monotonic variations when modeling feature evolution, HybridFF first estimates coefficients using moving least squares (MLS) for each of multiple complementary basis families and then combines the corresponding predictors using fusion weights. In addition to the choice of basis functions, the fusion weights also play a critical role. We introduce two strategies to balance quality and speedup. HybridFF (Fixed) prioritizes efficiency with model-specific fusion weights calibrated on a small set and held constant during inference. HybridFF (Adaptive) updates the fusion weights online using branch reliability scores computed from an exponential moving average of full-step prediction errors, improving prediction fidelity and generation quality under aggressive caching while retaining substantial acceleration. Experiments across DiT-XL/2, FLUX.1-dev, SD3.5-Large, and HunyuanVideo demonstrate a favorable speedup--quality trade-off over representative single-basis forecasters and caching baselines.
Figures & tables
Figure 1: HybridFF results: HunyuanVideo (top row, 3.85× ) and FLUX.1-dev (bottom two rows, 4.37× ) at NFE = 10.
Figure 2: DiT-XL/2 feature-trajectory analysis. (a) Intermediate features follow structured trajectories along denoising steps, and full-step observations can be fitted by polynomial bases to forecast skipped-step features. (b) Hybrid-basis polynomial forecasting gives lower feature-prediction error than single-basis predictors, motivating HybridFF’s coefficient-level fusion.
Figure 3: Overview of the proposed hybrid-basis forecasting pipeline. The pipeline proceeds in three stages: (1) full steps compute and cache exact intermediate features; (2) predict steps fit Taylor, Hermite, and Chebyshev predictors from the same normalized history; (3) the branch coefficient vectors are fused by simplex weights and used to reconstruct the de-normalized target block feature.
NFE
Method
SSIM ↑
PSNR ↑
LPIPS ↓
FID ↓
IS ↑
sFID ↓
Recall ↑
Precision ↑
Latency (s) ↓
Speedup ↑
50
Baseline (50-step)
–
–
–
2.276
240.855
4.258
0.592
0.802
0.971
1.00 ×
14
TaylorSeer
0.927
30.710
0.042
2.528
235.309
5.121
0.589
0.803
0.568
1.71 ×
14
HiCache
0.927
30.811
0.041
2.518
235.894
5.165
0.592
0.801
0.531
1.83 ×
14
Spectrum
0.914
29.233
0.052
2.604
230.968
4.577
0.596
0.794
0.542
1.79 ×
14
HybridFF (Fixed)
0.942
31.982
0.032
2.467
236.325
4.838
0.593
0.800
0.520
1.87 ×
14
HybridFF (Adaptive)
0.945
32.053
0.031
2.416
236.801
4.762
0.591
0.801
0.527
1.84 ×
Table 1: Comparison on class-conditional image generation with DiT-XL/2. Bold and underline mark the best and second-best entries in each block, respectively.
Figure 4: Comparison on class-conditional generation with DiT-XL/2 at NFE = 10. F/A denote fixed/adaptive fusion weights.
Figure 5: Comparison on text-to-image generation with SD3.5-Large at NFE = 14. F/A denote fixed/adaptive fusion weights.
NFE
Method
SSIM ↑
PSNR ↑
LPIPS ↓
VBench Quality ↑
Latency (s) ↓
Speedup ↑
50
Baseline (50-step)
–
–
–
80.958
351.227
1.00 ×
14
TaylorSeer
0.794
24.568
0.182
79.576
140.757
2.50 ×
14
HiCache
0.826
24.984
0.165
80.292
138.094
2.54 ×
14
Spectrum
0.834
26.237
0.140
80.259
128.548
2.73 ×
14
HybridFF (Fixed)
0.853
26.400
0.124
80.343
123.702
2.84 ×
14
HybridFF (Adaptive)
0.867
26.809
0.113
80.614
126.076
2.79 ×
Table 3: Comparison on text-to-video generation with HunyuanVideo. Bold and underline mark the best and second-best entries in each block, respectively.
Figure 6: Comparison on text-to-video generation with HunyuanVideo at NFE = 10. F/A denote fixed/adaptive fusion weights.
Figure 7: Ablations on the ridge regularizer and locality bandwidth with DiT-XL/2 at NFE = 10.
Config
SSIM ↑
PSNR ↑
LPIPS ↓
FID ↓
IS ↑
sFID ↓
Latency (s) ↓
Speedup ↑
D1W4
0.636
21.117
0.283
9.449
165.282
16.268
0.410
2.37 ×
D1W5
0.579
20.179
0.348
16.084
125.705
23.489
0.419
2.32 ×
D1W6
0.538
19.496
0.400
24.096
95.156
31.739
0.426
2.28 ×
D1W7
0.513
19.019
0.437
31.397
76.817
39.063
0.433
2.24 ×
D1W8
0.499
18.752
0.457
36.362
67.100
43.727
0.442
2.20 ×
D2W4
0.800
24.029
0.129
3.179
224.385
6.757
0.414
2.35 ×
Table 4: Ablation on the MLS polynomial degree D and history window W with DiT-XL/2 at NFE = 10.
Figure 8: SSIM landscape of HybridFF over the basis-weight simplex on a small FLUX.1-dev calibration set. The optimum (yellow star) lies inside the simplex at (π(Chebyshev),π(Hermite),π(Taylor))=(0.6,0.1,0.3) .
Figure 9: Additional qualitative comparisons on FLUX.1-dev at NFE = 10. The left and right panels show two prompt sets. F/A denote fixed/adaptive fusion weights.
Figure 10: Additional qualitative comparison on HunyuanVideo at NFE = 10. F/A denote fixed/adaptive fusion weights.