Diffusion Transformers (DiTs) achieve remarkable performance in video synthesis, but their iterative denoising process suffers from high inference latency. To address this, caching has emerged as an effective acceleration strategy by capitalizing on inter-step redundancy during denoising. Existing dynamic caching methods typically estimate the error that cache reuse would introduce at each denoising step (step error) to guide cache decisions, whereas our concern is how much quality loss cache reuse would cause in the final generated video (terminal error). We show that step error does not directly correspond to terminal error and that latent information helps capture their relationship, thereby informing cache decisions. Moreover, existing threshold-based methods cannot provide precise speedup control, making it difficult to meet practical requirements for user-specified acceleration targets. To address these limitations, we introduce MORCA, a cache scheduling framework trained through offline-to-online reinforcement learning to make latent-aware reuse/recompute decisions under user-specified acceleration targets. Extensive experiments on different video generation models across multiple target acceleration ratios demonstrate that MORCA achieves better generation fidelity than state-of-the-art caching methods under comparable computational budgets. Code is available at https://github.com/x10ngyx/MORCA.
Figures & tables
Figure 1: Step error and terminal error. (a) Following SeaCache, we cache and reuse the residuals of the DiT blocks. For each step, we evaluate two cases: (i) reusing the residual cached at the previous step and (ii) recomputing the DiT blocks, with all other steps computed normally. We measure step error as the relative L1 difference between the cached and recomputed residuals, and terminal error using the PSNR of the resulting video with cache reuse relative to the no-cache reference. (b) Step error does not directly correspond to terminal error. Results are shown for Wan2.1.
Input
RMSE (dB) ↓
Spearman ↑
Step error
6.971
0.5288
Step error + step index
4.421
0.7528
Step error + latent
2.849
0.9357
Step error + latent + step index
2.832
0.9361
Table 1: Latent features improve terminal-error prediction. Terminal error is measured using PSNR. Results on Wan2.1.
Figure 2: Overview of MORCA. Bottom: A learned scheduler makes reuse/recompute decisions at each denoising step. Top left: The scheduler is budget-constrained and latent-aware. The target speedup is mapped to the required total number of reuse steps, which, together with step-error proxies and step statistics, latent features, and the current reuse count extracted from the denoising state, is fed into the scheduler for decision-making (Sec. 3.2 ). Top right: The scheduler is optimized through offline-to-online reinforcement learning using IQL (Sec. 3.3 ).
Figure 3: Calibration of reuse budget K against (a) inference latency and (b) speedup on Wan2.1. The fitted functions take the forms (a) K=α−βTinf and (b) K=α−β/ρ .
Model
Target
Method
Latency (s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Wan2.2
—
Original
990.56
1.00×
—
—
—
1.8×
TeaCache
562.63
1.76×
17.04
0.5984
0.2882
MagCache
544.78
1.82×
16.41
0.5662
0.3041
DiCache
551.29
1.80×
18.51
0.6516
0.2232
SeaCache
546.99
1.81×
19.91
0.7016
0.1888
Ours
549.19
1.80×
22.13
0.7507
0.1681
Table 2: Quantitative comparison with baselines. The best result is highlighted in bold, while the second-best result is underlined.
Figure 4: Qualitative comparison with baselines. Differences are highlighted in the boxed regions.
Variant
P3(x)
P2(μT(x))
P2(σT2(x))
Latency (s) ↓
Speedup ↑
PSNR ↑
SSIM ↑
LPIPS ↓
A0
⋅
⋅
⋅
85.46
2.99×
19.93
0.6420
0.2240
A1
✓
⋅
⋅
86.10
2.97×
19.92
0.6515
0.2273
A2
⋅
✓
⋅
86.89
2.95×
19.76
0.6429
0.2314
A3
⋅
⋅
✓
85.39
3.00×
19.27
0.6151
0.2441
A4
✓
✓
⋅
85.38
3.00×
20.06
0.6598
0.2197
A5
✓
⋅
✓
85.29
3.00×
19.83
0.6484
0.2320
Table 3: Ablation on Latent-Aware State Design. Offline training results on Wan2.1 at approximately 3.0× speedup.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Wan2.2
Wan2.1
Resolution ( W×H )
832×480
832×480
Frames
45
81
Frame rate
16 fps
16 fps
Sampler
DPM++
UniPC
Denoising steps
50
50
Flow shift
12
5
Appendix
Table 5: Inference configurations for Wan2.2 and Wan2.1.
Group
Setting
Offline IQL
Online fine-tuning
Training
Training duration
200 epochs
16 rounds
Optimizer
AdamW
AdamW
π LR
10−4
4×10−5
Q/V LR
10−4
10−4
Batch size
256
256
Adam betas
(0.9,0.999)
(0.9,0.999)
Appendix
Table 6: Training and IQL hyperparameters for the two stages, shared across Wan2.2 and Wan2.1.
Model
Method
1.8×
2.4×
3.0×
Wan2.2
TeaCache
δ=0.29
δ=0.45
δ=0.68
MagCache
δ=0.072 , K=2 , R=0.2
δ=0.199 , K=4 , R=0.2
δ=0.198 , K=5 , R=0.1
DiCache
δ=0.075 , R=0.2
δ=0.172 , R=0.2
δ=0.396 , R=0.2
SeaCache
δ=0.24
δ=0.38
δ=0.55
Model
Method
2.4×
3.0×
4.0×
Wan2.1
TeaCache
δ=0.142
δ=0.195
δ=0.36
Appendix
Table 7: Baseline configurations at three target acceleration ratios for each model. δ denotes the cache threshold, K the maximum number of consecutive reuse steps for MagCache, and R the retention ratio.
Target
Latency (s)
Scheduler time (s)
Proportion (%)
1.8×
549.19
3.95
0.72
2.4×
421.15
3.94
0.94
3.0×
326.93
3.93
1.20
Appendix
Table 8: Scheduler overhead on Wan2.2. The scheduler introduces acceptable overhead.
Figure 5: Visualization of schedules and corresponding generation results under the same reuse budget on Wan2.2.
Figure 6: Speed–quality trade-offs of MORCA and SeaCache on Wan2.2. MORCA consistently achieves better fidelity at comparable speedups.
Figure 7: Additional qualitative comparisons on Wan2.2. From top to bottom: target acceleration ratios of 1.8× , 2.4× , and 3.0× .
Figure 8: Additional qualitative comparisons on Wan2.1. From top to bottom: target acceleration ratios of 2.4× , 3.0× , and 4.0× .
College of Computer Science and Artificial Intelligence, Fudan University · School of Data Science, Fudan University · 3Shanghai Key Laboratory of Intelligent Information Processing