Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At 1080×1920 output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72× faster while using 38.0% less peak allocated memory than FlashVSR Tiny. The code is available at https://github.com/kopperx/ReCaVSR.
Figures & tables
Figure 1: Streaming inference with ReCaVSR. Recycled SR latents convey local temporal context from the preceding block, while layer-wise cache routing retains only the historical KV states required by each DiT layer.
Figure 2: Overview of ReCaVSR. The generator conditions on LR observations and recycled SR latents, with layer-wise routing selecting historical KV states. In Stage 1, cache allocation is learned under budget and entropy regularization and exported as a static per-layer schedule. In Stage 2, the generator performs one-step sequential rollout under this schedule, supervised by the MSQ discriminator through global, spatial-window, and temporal-tube queries.
Dataset
Metric
RealViformer
UAV
STAR
DOVE
SeedVR2
SwiftVR
FlashVSR
Ours
REDS30
PSNR ↑
23.32
21.15
22.14
23.28
22.17
21.33
21.41
21.67
SSIM ↑
0.5967
0.5143
0.5432
0.6103
0.5860
0.5234
0.5399
0.5569
LPIPS ↓
0.3043
0.4036
0.4967
0.3773
0.3158
0.3564
0.3311
0.3129
NIQE ↓
3.0804
3.0017
5.2237
4.2282
3.5157
3.4153
2.9504
2.9368
MUSIQ ↑
59.12
60.02
37.95
50.44
57.65
63.77
56.01
60.42
CLIP-IQA ↑
0.3236
0.3698
0.2132
0.2904
0.3041
0.3914
0.3160
0.3396
Table 1: Quantitative comparison on synthetic and real-world VSR benchmarks. The best and second performances are marked in red and blue , respectively.
Figure 3: Qualitative comparisons on challenging synthetic and real-world VSR examples. ReCaVSR restores sharper structures, more natural textures, and cleaner face, text, and logo details.
Method
NIQE ↓
MUSIQ ↑
CLIP-IQA ↑
DOVER ↑
SC ↑
BC ↑
MS ↑
RealViformer
5.0770
46.17
0.3819
0.5053
0.8846
0.9198
0.9848
Stream-DiffVSR
4.0491
54.52
0.4694
0.5704
0.8830
0.9161
0.9813
SwiftVR
4.2822
53.89
0.4640
0.5512
0.8831
0.9177
0.9824
FlashVSR
3.8160
55.15
0.4747
0.5895
0.8829
0.9144
0.9802
Ours
3.9921
57.89
0.4932
0.6075
0.8834
0.9133
0.9828
Table 2: Quantitative comparison on LongVSR60 (Real + AIGC). The best and second-best results are marked in red and blue , respectively.
Metric
DOVE
SeedVR2-3B
SparkVSR
Stream-DiffVSR
FlashVSR
ReCaVSR
First-output latency (s) ↓
356.801
255.199
356.482
0.705
2.830
0.982
FPS ↑
0.563
0.788
0.564
0.849
7.799
21.202
Peak Mem. (GB) ↓
41.079
73.889
41.296
25.923
24.447
15.159
Table 3: GPU inference efficiency at 1080×1920 output resolution on 201 real input frames using a single NVIDIA A100-80GB GPU.
Variant
DOVER ↑
MS ↑
KV (GB) ↓
FLOPs ↓
Uniform ( 3R )
0.5601
0.9801
6.44
1.00×
Uniform ( R )
0.5261
0.9722
2.15
0.27×
Ours †
0.5130
0.9748
2.17
0.28×
Ours
0.5587
0.9785
2.17
0.28×
Table 7
Decoder
PSNR ↑
FPS ↑
Mem. (GB)
Wan2.1 VAE
38.5332
4.51
22.13
Wan2.2 VAE
39.1866
3.77
25.80
TCDecoder
36.8953
33.06
5.56
SwiftVR ReAE
30.9792
141.35
18.11
Ours ‡
34.4523
93.20
1.28
Ours
36.2525
91.71
1.33
Table 6: Decoder quality and efficiency. ‡ : without LR conditioning. Best and second-best values are bold and underlined.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: Blockwise streaming inference with ReCaVSR. Each block uses one DiT forward pass to generate SR latents, which are decoded with LR observations from the same block. Green dotted paths recycle the generated SR latents; orange paths carry the historical KV selected for each layer; gray paths preserve the LR projector and decoder states. The first block starts with zero recycled input and an empty DiT KV cache.
Generator
LR-conditioned decoder
Backbone
Wan2.2-TI2V-5B
Architecture
FlashDecoder
DiT layers
30
Backbone layers
12
Hidden dimension
3072
Refinement layers
2
Latent channels
48
Hidden dimension
512
Patch size (T,H,W)
1×2×2
Latent projection
48→512
LoRA rank
512
LR stem
768→48
Appendix
Table B.1: Architecture configurations of ReCaVSR.
Metric
DOVE
SeedVR2-3B
SwiftVR
FlashVSR Tiny
ReCaVSR
MOS-Q ↑
3.15
3.40
3.32
3.78
3.84
MOS-D ↑
3.17
3.35
3.15
3.72
3.81
MOS-T ↑
3.62
3.29
2.98
3.69
3.87
Appendix
Table C.2: Human evaluation on 35 VideoLQ videos with 15 raters. MOS-Q, MOS-D, and MOS-T measure overall visual quality, fine-detail quality, and temporal stability, respectively. Best results are shown in bold.
720p
1080p
1440p
4K
Method
FPS ↑
GB ↓
FPS ↑
GB ↓
FPS ↑
GB ↓
FPS ↑
GB ↓
RealViformer
31.88
5.44
14.87
12.36
8.69
21.66
3.88
39.03
DOVE
1.53
30.52
0.56
41.08
0.24
57.12
0.15 †
42.11 †
SeedVR2-3B
2.65
75.05
0.79 †
73.89 †
0.44 †
74.88 †
0.15 †
61.46 †
Stream-DiffVSR
1.76
10.24
0.85
25.92
0.47
61.67
0.17 †
44.39 †
FlashVSR Tiny
16.52
12.90
7.80
24.45
4.41
40.61
1.30
67.99
Appendix
Table C.3: Resolution scaling on one A100-80GB. GPU throughput and peak allocated memory for 201 input frames. Memory is reported in decimal GB.
Action
DiT layers
Layers
Slots per layer
∅
1, 2, 4, 6, 9, 22, 25, 27, 29, 30
10
0
W1
3, 28
2
1
W2
5, 7, 13, 17, 21, 24, 26
7
2
W4
8, 11, 20
3
4
A
15
1
2
W2+A
10, 19, 23
3
4
Appendix
Table C.4: Deployed layer-wise historical cache policy. Layer indices are one-based; reserved slots count latent positions per layer.
Figure C.2: Additional visual comparisons. Each example shows the input frame with a marked region and the corresponding crops from VSR methods. The scenes include building facades, masonry, and construction machinery.