Diffusion Transformers (DiTs) achieve remarkable performance in video synthesis, but their iterative denoising process suffers from high inference latency. To address this, caching has emerged as an effective acceleration strategy by capitalizing on inter-step redundancy during denoising. Existing dynamic caching methods typically estimate the error that cache reuse would introduce at each denoising step (step error) to guide cache decisions, whereas our concern is how much quality loss cache reuse would cause in the final generated video (terminal error). We show that step error does not directly correspond to terminal error and that latent information helps capture their relationship, thereby informing cache decisions. Moreover, existing threshold-based methods cannot provide precise speedup control, making it difficult to meet practical requirements for user-specified acceleration targets. To address these limitations, we introduce MORCA, a cache scheduling framework trained through offline-to-online reinforcement learning to make latent-aware reuse/recompute decisions under user-specified acceleration targets. Extensive experiments on different video generation models across multiple target acceleration ratios demonstrate that MORCA achieves better generation fidelity than state-of-the-art caching methods under comparable computational budgets. Code is available at https://github.com/x10ngyx/MORCA.
Figures & tables
Figure 1: Step error and terminal error. (a) Following SeaCache, we cache and reuse the residuals of the DiT blocks. For each step, we evaluate two cases: (i) reusing the residual cached at the previous step and (ii) recomputing the DiT blocks, with all other steps computed normally. We measure step error as the relative L1 difference between the cached and recomputed residuals, and terminal error using the PSNR of the resulting video with cache reuse relative to the no-cache reference. (b) Step error does not directly correspond to terminal error. Results are shown for Wan2.1.
Input
RMSE (dB) ↓
Spearman ↑
Step error
6.971
0.5288
Step error + step index
4.421
0.7528
Step error + latent
2.849
0.9357
Step error + latent + step index
2.832
0.9361
Table 1: Latent features improve terminal-error prediction. Terminal error is measured using PSNR. Results on Wan2.1.
Figure 2: Overview of MORCA. Bottom: A learned scheduler makes reuse/recompute decisions at each denoising step. Top left: The scheduler is budget-constrained and latent-aware. The target speedup is mapped to the required total number of reuse steps, which, together with step-error proxies and step statistics, latent features, and the current reuse count extracted from the denoising state, is fed into the scheduler for decision-making (Sec. 3.2 ). Top right: The scheduler is optimized through offline-to-online reinforcement learning using IQL (Sec. 3.3 ).
Figure 3: Calibration of reuse budget K against (a) inference latency and (b) speedup on Wan2.1. The fitted functions take the forms (a) K=α−βTinf and (b) K=α−β/ρ .
Model
Target
Method
Latency (s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Wan2.2
—
Original
990.56
1.00×
—
—
—
1.8×
TeaCache
562.63
1.76×
17.04
0.5984
0.2882
MagCache
544.78
1.82×
16.41
0.5662
0.3041
DiCache
551.29
1.80×
18.51
0.6516
0.2232
SeaCache
546.99
1.81×
19.91
0.7016
0.1888
Ours
549.19
1.80×
22.13
0.7507
0.1681
Table 2: Quantitative comparison with baselines. The best result is highlighted in bold, while the second-best result is underlined.
Figure 4: Qualitative comparison with baselines. Differences are highlighted in the boxed regions.
Variant
P3(x)
P2(μT(x))
P2(σT2(x))
Latency (s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
A0
⋅
⋅
⋅
85.46
2.99×
19.93
0.6420
0.2240
A1
✓
⋅
⋅
86.10
2.97×
19.92
0.6515
0.2273
A2
⋅
✓
⋅
86.89
2.95×
19.76
0.6429
0.2314
A3
⋅
⋅
✓
85.39
3.00×
19.27
0.6151
0.2441
A4
✓
✓
⋅
85.38
3.00×
20.06
0.6598
0.2197
A5
✓
⋅
✓
85.29
3.00×
19.83
0.6484
0.2320
Table 3: Ablation on Latent-Aware State Design. Offline training results on Wan2.1 at approximately 3.0× speedup.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Wan2.2
Wan2.1
Resolution ( W×H )
832×480
832×480
Frames
45
81
Frame rate
16 fps
16 fps
Sampler
DPM++
UniPC
Denoising steps
50
50
Flow shift
12
5
Appendix
Table 5: Inference configurations for Wan2.2 and Wan2.1.
Group
Setting
Offline IQL
Online fine-tuning
Training
Training duration
200 epochs
16 rounds
Optimizer
AdamW
AdamW
π LR
10−4
4×10−5
Q/V LR
10−4
10−4
Batch size
256
256
Adam betas
(0.9,0.999)
(0.9,0.999)
Appendix
Table 6: Training and IQL hyperparameters for the two stages, shared across Wan2.2 and Wan2.1.
Model
Method
1.8×
2.4×
3.0×
Wan2.2
TeaCache
δ=0.29
δ=0.45
δ=0.68
MagCache
δ=0.072 , K=2 , R=0.2
δ=0.199 , K=4 , R=0.2
δ=0.198 , K=5 , R=0.1
DiCache
δ=0.075 , R=0.2
δ=0.172 , R=0.2
δ=0.396 , R=0.2
SeaCache
δ=0.24
δ=0.38
δ=0.55
Model
Method
2.4×
3.0×
4.0×
Wan2.1
TeaCache
δ=0.142
δ=0.195
δ=0.36
Appendix
Table 7: Baseline configurations at three target acceleration ratios for each model. δ denotes the cache threshold, K the maximum number of consecutive reuse steps for MagCache, and R the retention ratio.
Target
Latency (s)
Scheduler time (s)
Proportion (%)
1.8×
549.19
3.95
0.72
2.4×
421.15
3.94
0.94
3.0×
326.93
3.93
1.20
Appendix
Table 8: Scheduler overhead on Wan2.2. The scheduler introduces acceptable overhead.
Figure 5: Visualization of schedules and corresponding generation results under the same reuse budget on Wan2.2.
Figure 6: Speed–quality trade-offs of MORCA and SeaCache on Wan2.2. MORCA consistently achieves better fidelity at comparable speedups.
Figure 7: Additional qualitative comparisons on Wan2.2. From top to bottom: target acceleration ratios of 1.8× , 2.4× , and 3.0× .
Figure 8: Additional qualitative comparisons on Wan2.1. From top to bottom: target acceleration ratios of 2.4× , 3.0× , and 4.0× .
Diffusion models achieve state-of-the-art video generation quality, but their inference remains expensive due to the large number of sequential denoising steps. This has motivated a growing line of research on accelerating diffusion inference. Among training-free acceleration methods, caching reduces computation by reusing previously computed model outputs across timesteps. Existing caching methods rely on heuristic criteria to choose cache/reuse timesteps and require extensive tuning. We address this limitation with a principled sensitivity-aware caching framework. Specifically, we formalize the caching error through an analysis of the model output sensitivity to perturbations in the denoising inputs, i.e., the noisy latent and the timestep, and show that this sensitivity is a key predictor of caching error. Based on this analysis, we propose Sensitivity-Aware Caching (SenCache), a dynamic caching policy that adaptively selects caching timesteps on a per-sample basis. Our framework provides a theoretical basis for adaptive caching, explains why prior empirical heuristics can be partially effective, and extends them to a dynamic, sample-specific approach. Experiments on Wan 2.1, CogVideoX, and LTX-Video show that SenCache achieves better visual quality than existing caching methods under similar computational budgets.
Video diffusion models produce high-quality generations but remain slow at inference due to their sequential denoising procedure. Caching-based acceleration methods address this by reusing intermediate model outputs: leading dynamic approaches such as TeaCache, EasyCache, and DiCache accumulate a drift signal and skip expensive model evaluations when accumulated drift stays below a fixed threshold τ. This threshold controls an apparent tradeoff - raising it yields faster generation at the cost of visual quality, while lowering it preserves quality but sacrifices speed. We show this tradeoff is not fundamental; it is an artifact of holding τ constant throughout denoising. We identify the existence of critical steps - timesteps where the drift signal changes rapidly - and show that applying a low threshold selectively at these steps while caching aggressively elsewhere recovers most of the quality of conservative caching at substantially higher inference speeds. Building on this insight, we propose ACID, a lightweight, training-free wrapper that monitors the rate of change of each method's existing drift signal to dynamically switch between a low and a high threshold. ACID is signal-agnostic and modular: it requires no retraining and plugs directly into existing dynamic caching methods without modifying their core mechanisms. Evaluated across three caching methods (TeaCache, EasyCache, DiCache) and three open-source video diffusion models (HunyuanVideo, Wan 2.1, CogVideoX), ACID consistently expands the Pareto frontier of visual quality versus inference speed beyond what any fixed threshold achieves. In particular, on TeaCache and HunyuanVideo, ACID achieves up to 2.16x speedup over the no-caching baseline, and up to 38% additional speedup over the conservative fixed-threshold baseline with negligible (<0.3 dB PSNR, <0.01 SSIM, <0.01 LPIPS) quality degradation.
Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising. While cache-based strategies accelerate inference by reusing intermediate features, they largely rely on static, sample-agnostic schedules. We argue that this rigidity overlooks two facts empirically validated in this paper: (i) generation difficulty varies across prompts, requiring adaptive resource allocation--complex inputs demand more computation while simpler ones require less; (ii) error sensitivity fluctuates across timesteps, where static policies may cache high-error steps or waste computation on low-error ones. We therefore propose OnlineCache, a dynamic caching framework that jointly learns when to cache and how to correct approximation errors. We leverage policy gradient to train a lightweight network for adaptive speed-quality trade-offs, and incorporate a learnable corrector to mitigate caching-induced errors. Both modules are jointly optimized under a bilevel optimization framework, with the policy targeting global generation quality and the corrector minimizing local errors. Our method automatically allocates computational resources across both samples and timesteps, improving overall generation quality. Extensive experiments demonstrate clear superiority. On FLUX.1-dev model, OnlineCache achieves nearly 3 speedup while preserving generation fidelity. On DiT and CogVideoX, it similarly delivers competitive acceleration without compromising quality; across all scenarios, it consistently outperforms existing cache-based acceleration baselines.
Zhikang Xie, Xichen Ye, Yifan Wu +5
College of Computer Science and Artificial Intelligence, Fudan University · School of Data Science, Fudan University · 3Shanghai Key Laboratory of Intelligent Information Processing