Feature caching accelerates diffusion sampling by replacing expensive network evaluations with predictions from previously computed activations. However, forecasts based only on past features cannot directly incorporate changes in the current denoising state. We investigate whether inexpensive, freshly computed features can serve as observations for correcting these predictions. We introduce an observation-corrected Koopman framework for accelerating frozen diffusion models. Using calibration trajectories, we identify finite-dimensional, time-dependent Koopman approximations that jointly describe the increments of shallow and deep network features. During accelerated sampling, these operators predict the evolution of expensive deep features, while innovations in the observed shallow features correct the predicted state. Periodic full evaluations refresh the observer, and all generative-model parameters remain unchanged. This formulation enables controlled comparisons of temporal prediction and observation correction. Across three 10,000-image runs per dataset, our method reduces paired Inception-feature MSE by 19.9% on CIFAR-10 and 11.9% on a ten-class ImageNet subset relative to channelwise affine prediction under the same four-partial-step schedule. Matched ablations attribute additional reductions of 4.54% and 4.67% to observation correction. The observer achieves 1.89× and 1.85× measured speedups over DDIM-50, supporting improved reference-sampler fidelity without retraining the denoiser.
Figures & tables
Figure 1: Observation correction on selected ImageNet examples. Each row compares full-model generation, open-loop prediction, and observation-corrected prediction using the same initial noise and class label. The accelerated variants share the full-evaluation schedule. Examples are selected for large pixel-MSE reductions among samples that also improve in Inception-feature MSE, with one example per class. They illustrate improvement cases, not average performance. Images are generated at 64×64 on a ten-class ImageNet subset. Full-model outputs are generated references, not ground-truth images. Heatmaps show per-pixel RGB RMSE on [0,1] -scaled images using a shared color scale. Numerical values and percentage reductions report Inception-feature MSE relative to the full-model reference.
Figure 2: Koopman observers for diffusion acceleration. Left: Fresh shallow observations correct the predicted joint increment state. The reconstructed deep feature and current shallow skips feed the retained readout, while expensive deep blocks are skipped. Projection bases, scales, time-dependent operators, and correction gains are identified from offline calibration trajectories. Right: Two full evaluations (F) initialize the observer state. Partial evaluations (P) use the Koopman observer Ok to predict, correct, and reconstruct deep features. Subsequent full evaluations refresh the estimates at the current accelerated samples. The schedule shows an illustrative interior segment. All denoiser weights remain frozen and the sampler is unchanged.
CIFAR-10
ImageNet-10
Method
FID ↓
103Eϕ↓
Speedup ↑
FID ↓
103Eϕ↓
Speedup ↑
Full DDIM-50
13.39±0.05
0.000
1.00×
13.72±0.08
0.000
1.00×
Reduced-step DDIM
15.43±0.10
10.419
1.79×
14.55±0.10
21.492
1.67×
DPM-Solver++ 2M
16.84±0.10
28.282
1.79×
14.76±0.09
38.720
1.67×
Feature reuse
15.90±0.09
15.985
1.71×
15.95±0.04
37.214
1.69×
Linear extrapolation
13.53±0.07
2.982
1.71×
14.08±0.09
14.924
1.69×
Table 1: Generation quality and full-sampler fidelity at the middle latency budget. Each result uses 10,000 images per sampling run and three runs. FID reports mean and sample SD. Eϕ is mean paired Inception-feature MSE relative to full DDIM-50. All seven baseline families are within 3.81% of the observer latency. † denotes an adaptation to the native U-Net.
Figure 3: Distributional quality and reference fidelity are different objectives. FID (top) and paired Inception-feature MSE (bottom) against measured speedup on CIFAR-10 (left) and ImageNet-10 (right). Points show means over three sampling runs, with error bars indicating sample SD. The bottom row uses a logarithmic vertical axis. Baselines are the configurations selected by timing for the three observer budgets, with repeated selections shown once. Every point is plotted at its actual measured speedup, including configurations outside the 5% matching tolerance. Lines connect evaluated configurations as visual guides. Dashed horizontal lines show full DDIM-50 FID.
Dataset
Correction input
FID ↓
103Eϕ↓
103Ex↓
s/batch
CIFAR-10
Disabled
13.638±0.070
2.842±0.050
0.244±0.003
3.757
CIFAR-10
One-step delayed
13.807±0.067
3.110±0.027
0.308±0.003
3.787
CIFAR-10
Current
13.669±0.071
2.713±0.047
0.224±0.005
3.781
ImageNet-10
Disabled
14.266±0.107
18.172±0.230
4.277±0.052
5.310
ImageNet-10
One-step delayed
14.335±0.021
20.984±0.124
5.086±0.009
5.357
ImageNet-10
Current
14.257±0.073
17.323±0.235
3.900±0.045
5.344
Table 2: Effect of current observations with a fixed fitted predictor and the same s=4 schedule. All three variants use the unfused implementation and fresh shallow decoder skips. Ex is pixel MSE on floating-point images in [−1,1] . Errors and FID report mean and sample SD across three sampling runs. Latency is the median of three synchronized repetitions.
Figure 4: Stage-dependent prediction and correction. Columns show CIFAR-10 and ImageNet-10. Top: open-loop deep-feature squared error normalized by feature-reuse error at the same position. Bottom: correction gain, 100(1−Eobs/Eopen) . Positive values indicate improvement. Each cell aggregates 60 CIFAR-10 or 40 ImageNet-10 development trajectories. Anchor steps are zero based and follow sampling order from high to low noise. Gray cells target full evaluations and are excluded.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Calibration
Validation
Diagnostic
Rank r
CIFAR-10
120
60
60
256
ImageNet-10
80
40
40
256
Appendix
Table 3: Complete-trajectory splits for offline identification. The diagnostic split is separate from the formal generation evaluation. Rank is the dimension of each projected increment.
Dataset
s
FID ↓
103KID↓
103Eϕ↓
Speedup ↑
CIFAR-10
2
13.495±0.050
6.319±0.133
0.743
1.527×
CIFAR-10
3
13.515±0.079
6.224±0.098
1.776
1.733×
CIFAR-10
4
13.669±0.071
6.393±0.115
2.713
1.890×
ImageNet-10
2
13.914±0.079
2.665±0.024
4.662
1.504×
ImageNet-10
3
14.133±0.097
2.760±0.032
9.741
1.701×
ImageNet-10
4
14.257±0.074
2.802±0.029
17.323
1.845×
Appendix
Table 4: Observer results across the three refresh budgets. FID and KID report mean and sample SD across three sampling runs. Each run generates 10,000 images. Speedup is relative to full DDIM-50 at the main batch size. Eϕ reports mean paired feature error.
CIFAR-10
ImageNet-10
Predictor
103Eϕ↓
Speedup
103Eϕ↓
Speedup
Feature reuse
18.270±0.144
1.93×
50.543±0.169
1.89×
Linear extrapolation
4.395±0.058
1.92×
21.794±0.277
1.88×
Channelwise affine
3.388±0.067
1.92×
19.653±0.339
1.88×
Static regression, r=256
12.275±0.029
1.89×
34.979±0.361
1.85×
Diagonal observer, r=64
3.478±0.056
1.90×
19.560±0.239
1.85×
Appendix
Table 5: Predictor controls on the same s=4 refresh schedule. Feature errors report mean and sample SD across three sampling runs. The rank-256 observer uses fused inference. Other configurations can differ in rank, fitting, or implementation. The strict correction intervention is reported in Table 2 .
Predictor
CIFAR-10
ImageNet-10
Feature reuse
1.0000
1.0000
Linear extrapolation
0.6022
0.6886
Static shallow-to-deep
0.5964
0.5643
Open-loop dynamics
0.3829
0.3907
Observation-corrected dynamics
0.3664
0.3763
Appendix
Table 6: Local deep-feature prediction on earlier development diagnostic trajectories. Errors are normalized by feature reuse and aggregated over horizons of one, two, and four steps.
Dataset
Comparison
Run 1
Run 2
Run 3
CIFAR-10
Current / disabled
0.9529
0.9591
0.9518
[0.9463,0.9594]
[0.9521,0.9667]
[0.9414,0.9616]
CIFAR-10
Current / delayed
0.8622
0.8774
0.8770
[0.8500,0.8743]
[0.8657,0.8887]
[0.8646,0.8902]
ImageNet-10
Current / disabled
0.9541
0.9541
0.9516
[0.9481,0.9602]
[0.9478,0.9605]
[0.9456,0.9578]
Appendix
Table 7: Paired feature-error ratios for the strict observation interventions at s=4 . Brackets give per-run 95% paired-bootstrap intervals from 1,000 resamples. Ratios below one favor current observations. Intervals are not pooled across runs.
Dataset
Batch
Full
Reuse
Channel
Observer
Speedup
CIFAR-10
10
3.399
1.586
1.589
1.603
2.121×
CIFAR-10
20
4.166
2.119
2.124
2.149
1.939×
CIFAR-10
50
7.043
3.657
3.671
3.728
1.889×
CIFAR-10
100
12.507
6.565
6.594
6.706
1.865×
ImageNet-10
5
4.093
2.098
2.102
2.130
1.922×
ImageNet-10
10
5.959
3.069
3.082
3.131
1.903×
Appendix
Table 8: Separate batch-size timing measurements for s=4 . Latencies are seconds per batch, measured as the median of three synchronized repetitions after warmup. Each speedup uses the full sampler timed again at the corresponding batch size.
Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on designing stronger forecasters. Yet forecast error varies sharply across steps, and open-loop caches trust the forecast in full at every skipped step. This fixed trust is what breaks as acceleration turns aggressive. The missing question is not only how to forecast better, but when and how much to trust a forecast. We show that reliability can be observed from the cache itself. Two forecasts agree where the feature trajectory is smooth, and they diverge where prediction turns hard. Their disagreement is a cheap runtime signal, and it costs no extra denoiser evaluation. Based on this signal, we introduce RACER, a training-free closed-loop controller with two responses. It continuously shrinks uncertain forecasts toward the last computed feature. At the riskiest steps, RACER refreshes the feature and repays the added evaluation by skipping a later scheduled one. We derive a deterministic error bound for the shrinkage and empirically evaluate its validity and tightness across acceleration regimes. At the same number of denoiser evaluations, RACER improves the strongest open-loop baseline across SD3.5-Large, FLUX.1-dev, Wan2.1-14B, and HunyuanVideo on DrawBench, VBench, and COCO. On SD3.5, we further show that RACER samples faster at equal quality. RACER generalizes across forecasting designs as well. For example, it recovers much of the quality lost on a Taylor base. These results show that reliable diffusion acceleration also depends on how forecasts are used. Code is available at https://github.com/LiZaiyuan0619/RACER
Yanchao Li, Jiaqing Xie, Ben Gao +6
1Nanjing University · 2Shanghai Artificial Intelligence Laboratory · 3North University of China +2
Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer--timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to 6.70× over Vanilla while maintaining competitive output quality.
Hanshuai Cui, Zhiqing Tang, Zhi Yao +3
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China
Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference speed due to the numerous iterative passes of Diffusion Transformers. To reduce the exhaustive compute, recent works resort to the feature caching and reusing scheme that skips network evaluations at selected diffusion steps by using cached features in previous steps. However, their preliminary design solely relies on local approximation, causing errors to grow rapidly with large skips and leading to degraded sample quality at high speedups. In this work, we propose spectral diffusion feature forecaster (Spectrum), a training-free approach that enables global, long-range feature reuse with tightly controlled error. In particular, we view the latent features of the denoiser as functions over time and approximate them with Chebyshev polynomials. Specifically, we fit the coefficient for each basis via ridge regression, which is then leveraged to forecast features at multiple future diffusion steps. We theoretically reveal that our approach admits more favorable long-horizon behavior and yields an error bound that does not compound with the step size. Extensive experiments on various state-of-the-art image and video diffusion models consistently verify the superiority of our approach. Notably, we achieve up to 4.79× speedup on FLUX.1 and 4.67× speedup on Wan2.1-14B, while maintaining much higher sample quality compared with the baselines.