Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces this cost by reuse of previously computed features or predictions. However, existing branch-local reuse criteria do not explicitly account for how cache errors combine under CFG or how local perturbations affect the final output. We identify two misalignments in cache control: a branch-guided mismatch, where guided error depends on both the magnitudes and alignment of branch errors, and a local-final mismatch, where the downstream impact of a local error varies across timesteps. We propose RA-CFGCache, a Risk-Aligned Caching framework under CFG that incorporates both factors while keeping the sampling schedule and guidance rule fixed. CFG-aware Guided-Risk Composition combines existing branch-wise proxies using CFG coefficients and offline-calibrated cross-branch alignment. Propagation-Aware Rescaling further weights the resulting guided-risk estimate with a timestep-dependent propagation prior calibrated from isolated reuse perturbations. An online threshold controller then determines when to jointly refresh or reuse both branches. Experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B demonstrate improved efficiency--fidelity trade-offs over evaluated training-free caching baselines. Moreover, RA-CFGCache is compatible with diverse base proxy families, including TeaCache-, DiCache-, and MagCache-style estimators, and consistently improves fidelity at nearly unchanged latency. Code is available at https://github.com/yiming-l21/RA-CFGCache.git.
Figures & tables
Figure 1 : Standard caching vs. CFG caching. Standard caching reuses a single prediction whose cache error directly enters the denoising update, whereas CFG caching reuses conditional and unconditional branches whose errors are composed into the guided prediction.
Figure 2 : Two mismatches under CFG caching. (a) The first mismatch: branch-wise cache errors do not directly reflect the guided error that enters the denoising update. (b) The second mismatch: local guided error does not directly predict the final deviation after error propagation across timesteps.
Figure 3 : Two sources of mismatch under CFG caching. (a) The alignment ρt=cos(δut,δct) varies across timesteps, inducing timestep-dependent cancellation or amplification in Eq. ( 7 ). (b) The propagation gain gt varies across timesteps, showing that similar local guided errors can induce different final deviations. Solid curves show the mean over 100 prompts, and shaded bands indicate ± one standard deviation across prompts.
Figure 4 : Overview of RA-CFGCache. Under classifier-free guidance (CFG), cache control should be aligned with the guided prediction rather than either branch alone. RA-CFGCache first combines branch-wise proxy estimates into a guided-risk estimate, then rescales it using a timestep-dependent propagation prior, and finally uses the resulting risk estimate for online refresh decisions.
Figure 5 : Trade-off curves. RA-CFGCache achieves lower LPIPS at comparable latency across models.
Model
Method
LPIPS ↓
SSIM ↑
PSNR ↑
Speedup ↑
Latency (s) ↓
FLUX.1-dev
Vanilla ( T=50 )
0.0000
1.0000
∞
1.00 ×
21.02
Vanilla ( T=25 )
0.4556
0.6207
13.09
1.96 ×
10.71
TeaCache ( τ =0.4)
0.4471
0.6174
13.80
2.53 ×
8.31
TaylorSeer ( N =3, O =1)
0.4379
0.6398
13.75
2.37 ×
8.86
DiCache ( m=2,τ =0.4)
0.1431
0.8229
22.13
2.94 ×
7.15
MagCache ( τ =0.24)
0.1636
0.8177
21.46
3.18 ×
6.60
Table 1 : Main results on text-to-image and text-to-video generation. Comparison of RA-CFGCache with recent diffusion acceleration baselines.
Figure 6 : Qualitative comparison on FLUX.1-dev. We show three representative text-to-image examples under the same prompts and random seeds. Compared with TeaCache, MagCache, and DiCache, RA-CFGCache better preserves local textures, object structures, and semantic details under comparable acceleration settings.
Table 8Table 9Table 10
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Temporal dynamics of the conditional correction signal under CFG. From left to right: the norm of Δt , the adjacent-step variation ∥Δt−Δt−1∥1 , and the cosine similarity cos(Δt,Δt−1) between consecutive correction vectors.
Figure 8 : Refresh redistribution under matched budget. Left: per-step refresh frequency of the branch-wise and guided-composition controllers. Right: cumulative refresh fraction over timesteps. Guided-composition shifts refresh decisions toward earlier timesteps, showing that preserving CFG error geometry changes the actual scheduling behavior rather than only the local risk value.
Figure 9 : Best reuse action frequency across timesteps under different CFG scales (top-left: 1.5, top-right: 3.5, bottom-left: 5.5, bottom-right: 7.5). As CFG scale increases, reuse-both dominates a larger portion of timesteps, while reuse- Δ becomes favorable mainly in the very late stage. Reuse-cond is rarely optimal throughout.
Method
Speedup ↑
LPIPS ↓
SSIM ↑
PSNR ↑
CLIP ↑
ImageReward ↑
HPSv2 ↑
Vanilla ( T=50 )
1.00 ×
0.0000
1.0000
∞
27.9063
0.8003
0.2776
Vanilla ( T=25 )
1.96 ×
0.4556
0.6207
13.09
27.0429
0.2629
0.2364
FasterCache
2.09 ×
0.2158
0.7593
20.88
27.4779
0.6241
0.2643
MAMBO-G (T=25)
1.98 ×
0.5616
0.5192
11.64
28.4720
0.9800
0.2930
CFG Scheduler
1.00 ×
0.5327
0.5503
12.26
28.1390
1.0160
0.3000
OUSAC
3.44 ×
0.5530
0.5092
11.74
27.8770
0.9300
0.2880
Appendix
Table 9 : Comparison with CFG-aware guidance and acceleration methods on FLUX.1-dev. RA-CFGCache preserves the original sampler and fixed CFG rule. MAMBO-G, CFG Scheduler, and OUSAC modify the guidance policy, while FasterCache exploits CFG branch redundancy. LPIPS, SSIM, and PSNR measure sampler fidelity, with CLIPScore, ImageReward, and HPSv2 reported as complementary quality metrics.
Base method
Variant
Speedup ↑
LPIPS ↓
SSIM ↑
PSNR ↑
TeaCache
Original
3.27 ×
0.3617
0.7112
17.05
TeaCache
+ Propagation Rescaling
3.32 ×
0.3207
0.7305
18.02
MagCache
Original
3.11 ×
0.1964
0.8169
22.05
MagCache
+ Propagation Rescaling
3.12 ×
0.1822
0.8266
22.57
DiCache
Original
3.20 ×
0.2999
0.7744
21.58
DiCache
+ Propagation Rescaling
3.19 ×
0.2393
0.7966
22.38
Appendix
Table 10: Generality of propagation-aware rescaling in the standard single-prediction setting. We apply propagation-aware rescaling to several representative non-CFG caching baselines and observe consistent fidelity improvements across the evaluated proxy families.
Base proxy family
Controller
Latency ↓
Speedup ↑
LPIPS ↓
SSIM ↑
PSNR ↑
Tea-style proxy
Branch-wise accumulation
6.16 s
3.41 ×
0.1852
0.7977
20.60
+ RA-CFGCache wrapper
6.16 s
3.41 ×
0.1492
0.8297
22.39
Mag-style proxy
Branch-wise accumulation
6.18 s
3.40 ×
0.1683
0.8144
21.35
+ RA-CFGCache wrapper
6.17 s
3.41 ×
0.1232
0.8590
23.53
DiCache-style proxy
Branch-wise accumulation
6.13 s
3.43 ×
0.2177
0.7706
19.48
+ RA-CFGCache wrapper
6.13 s
3.43 ×
0.1453
0.8364
22.57
Appendix
Table 11: Compatibility with different base proxy estimators. We apply RA-CFGCache on top of different branch-wise proxy families and compare it against their original branch-wise accumulation controllers. Across all proxy types, RA-CFGCache consistently improves fidelity at nearly unchanged latency, showing that its gain is complementary to the underlying proxy design rather than tied to a specific estimator.
Figure 10 : Additional calibration visualizations. Top-left: pairwise branch-error alignment heatmap. The other three panels show model-specific power-law fits of the propagation gain curve gt .
Figure 11 : Alignment curves across CFG scales for different diffusion models. Across models, ρt generally remains lower or less stable in the early-to-middle stage and tends toward a higher-alignment regime in later timesteps, although detailed local fluctuations are model-dependent.
Figure 12 : Propagation gain curves gt across CFG scales for different diffusion models. The propagation gain is strongly timestep-dependent, with larger values in the early or early-middle stages and much smaller values in later steps. Across CFG scales, the curves show model-dependent magnitude differences and tend to become closer in later steps, indicating that local guided errors have substantially different final impacts depending on their injection timestep.
Surrogate
R2↑
NMAE ↓
Speedup ↑
LPIPS ↓
SSIM ↑
PSNR ↑
No prop. rescaling
–
–
3.74 ×
0.1699
0.8162
21.41
Fixed linear decay
-0.0024
0.3174
3.75 ×
0.1469
0.8297
22.74
Fitted linear decay
0.7164
0.1521
3.76 ×
0.1448
0.8379
22.80
Power-law decay
0.8201
0.1010
3.77 ×
0.1406
0.8398
22.93
Appendix
Table 12: Effect of propagation-gain surrogate. We compare compact surrogate forms for propagation-aware rescaling. The power-law surrogate provides the best fidelity among the tested smooth parameterizations, achieving the lowest LPIPS and highest SSIM with competitive PSNR.
Controller
Reuse both (%) ↑
One-sided reuse (%) ↓
Full-step refresh (%)
Latency (s) ↓
Branch-wise caching
35.1
53.8
11.1
6.81
Ours (reuse both)
62.0
0.0
38.0
4.21
Appendix
Table 13: Effect of caching strategies under CFG parallelism. Controlled FLUX.1-dev study at 1024 × 1024 resolution, CFG scale 3.5, 50 denoising steps, comparing independent branch-wise caching and joint guided caching under a matched branch-evaluation budget.
Figure 13 : Qualitative comparison on FLUX.1-dev.
Figure 14 : Qualitative comparison on CogVideoX-2B (Part I).
Figure 15 : Qualitative comparison on CogVideoX-2B (Part II).
Figure 16 : Qualitative comparison on Wan2.1-T2V-1.3B (Part I).
Figure 17 : Qualitative comparison on Wan2.1-T2V-1.3B (Part II).
Modern diffusion models generate high-quality images and videos, but their iterative denoising process makes inference expensive. Feature caching accelerates sampling by reusing or predicting intermediate activations across neighboring denoising steps, exploiting the redundancy of computations along the reverse trajectory. In this work, we focus on the caching schedule: selecting which denoising steps should be fully recomputed. Existing schedules are either fixed (e.g. uniform) or chosen adaptively from per-step error heuristics; in both cases, the actual compute cost is a side-effect of hand-tuned thresholds rather than a quantity the user can specify. We propose ReCache, which inverts this: given a target budget k, it learns the recomputation schedule that maximizes generation quality, turning compute into a directly controllable input. ReCache trains via policy gradients, sidestepping backpropagation through full diffusion inference, and uses no labelled data. Generations from uncached inference serve as matching targets, paired with a reward for generation quality. ReCache is compatible with any caching mechanism, including feature reuse and feature forecasting; for each mechanism, a single trained policy adapts across computational budgets at inference time. ReCache consistently outperforms scheduling baselines: under a ×5.04 FLOPs reduction on FLUX, it reduces LPIPS by 31% (from 0.456 to 0.316) compared to DiCache; on Wan 2.1 at a ∼×2.6 speedup, it drops LPIPS by 65% (from 0.480 to 0.169) and boosts the VBench score by 7% (5.6 points, from 70.4 to 76.0) over uniform HiCache. Code is available at https://github.com/thecrazymage/ReCache.
Mishan Aliev, Eva Neudachina, Ilya Bykov +4
1HSE University, Russia · 2Yandex Research, Russia
Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising. While cache-based strategies accelerate inference by reusing intermediate features, they largely rely on static, sample-agnostic schedules. We argue that this rigidity overlooks two facts empirically validated in this paper: (i) generation difficulty varies across prompts, requiring adaptive resource allocation--complex inputs demand more computation while simpler ones require less; (ii) error sensitivity fluctuates across timesteps, where static policies may cache high-error steps or waste computation on low-error ones. We therefore propose OnlineCache, a dynamic caching framework that jointly learns when to cache and how to correct approximation errors. We leverage policy gradient to train a lightweight network for adaptive speed-quality trade-offs, and incorporate a learnable corrector to mitigate caching-induced errors. Both modules are jointly optimized under a bilevel optimization framework, with the policy targeting global generation quality and the corrector minimizing local errors. Our method automatically allocates computational resources across both samples and timesteps, improving overall generation quality. Extensive experiments demonstrate clear superiority. On FLUX.1-dev model, OnlineCache achieves nearly 3 speedup while preserving generation fidelity. On DiT and CogVideoX, it similarly delivers competitive acceleration without compromising quality; across all scenarios, it consistently outperforms existing cache-based acceleration baselines.
Zhikang Xie, Xichen Ye, Yifan Wu +5
College of Computer Science and Artificial Intelligence, Fudan University · School of Data Science, Fudan University · 3Shanghai Key Laboratory of Intelligent Information Processing
Diffusion Transformers require repeated denoiser evaluations during iterative sampling, making inference computationally expensive. Cache-based acceleration reduces this cost by reusing intermediate representations across denoising steps, but can introduce representation deviations and degrade generation quality. In this paper, we analyze these deviations and show that effective calibration should consider both the direct mismatch caused by reuse and the subsequent trajectory shift induced by earlier corrections. To address this challenge, we propose Trajectory-Consistent Calibration (TCC), a training-free method that calibrates cached representations toward their full-computation counterparts. Specifically, rather than estimating all calibration priors from a single uncorrected cache trajectory, TCC uses an offline iterative procedure so that each prior accounts for the trajectory shift induced by preceding calibrations. Experiments on PixArt-alpha and DiT-XL/2 show that TCC consistently improves FID across representative cache-based acceleration methods while preserving their underlying reuse policies. Notably, in a representative PixArt-alpha cache-acceleration setting based on FORA, TCC reduces FID from 29.83 to 27.35, slightly surpassing the full-computation baseline.
Mingyu Liang, Dingkun Xu, Jingwei Xu
Laboratory for Novel Software Technology, Nanjing University, China