Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At 1080×1920 output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72× faster while using 38.0% less peak allocated memory than FlashVSR Tiny. The code is available at https://github.com/kopperx/ReCaVSR.
Figures & tables
Figure 1: Streaming inference with ReCaVSR. Recycled SR latents convey local temporal context from the preceding block, while layer-wise cache routing retains only the historical KV states required by each DiT layer.
Figure 2: Overview of ReCaVSR. The generator conditions on LR observations and recycled SR latents, with layer-wise routing selecting historical KV states. In Stage 1, cache allocation is learned under budget and entropy regularization and exported as a static per-layer schedule. In Stage 2, the generator performs one-step sequential rollout under this schedule, supervised by the MSQ discriminator through global, spatial-window, and temporal-tube queries.
Dataset
Metric
RealViformer
UAV
STAR
DOVE
SeedVR2
SwiftVR
FlashVSR
Ours
REDS30
PSNR ↑
23.32
21.15
22.14
23.28
22.17
21.33
21.41
21.67
SSIM ↑
0.5967
0.5143
0.5432
0.6103
0.5860
0.5234
0.5399
0.5569
LPIPS ↓
0.3043
0.4036
0.4967
0.3773
0.3158
0.3564
0.3311
0.3129
NIQE ↓
3.0804
3.0017
5.2237
4.2282
3.5157
3.4153
2.9504
2.9368
MUSIQ ↑
59.12
60.02
37.95
50.44
57.65
63.77
56.01
60.42
CLIP-IQA ↑
0.3236
0.3698
0.2132
0.2904
0.3041
0.3914
0.3160
0.3396
Table 1: Quantitative comparison on synthetic and real-world VSR benchmarks. The best and second performances are marked in red and blue , respectively.
Figure 3: Qualitative comparisons on challenging synthetic and real-world VSR examples. ReCaVSR restores sharper structures, more natural textures, and cleaner face, text, and logo details.
Method
NIQE ↓
MUSIQ ↑
CLIP-IQA ↑
DOVER ↑
SC ↑
BC ↑
MS ↑
RealViformer
5.0770
46.17
0.3819
0.5053
0.8846
0.9198
0.9848
Stream-DiffVSR
4.0491
54.52
0.4694
0.5704
0.8830
0.9161
0.9813
SwiftVR
4.2822
53.89
0.4640
0.5512
0.8831
0.9177
0.9824
FlashVSR
3.8160
55.15
0.4747
0.5895
0.8829
0.9144
0.9802
Ours
3.9921
57.89
0.4932
0.6075
0.8834
0.9133
0.9828
Table 2: Quantitative comparison on LongVSR60 (Real + AIGC). The best and second-best results are marked in red and blue , respectively.
Metric
DOVE
SeedVR2-3B
SparkVSR
Stream-DiffVSR
FlashVSR
ReCaVSR
First-output latency (s) ↓
356.801
255.199
356.482
0.705
2.830
0.982
FPS ↑
0.563
0.788
0.564
0.849
7.799
21.202
Peak Mem. (GB) ↓
41.079
73.889
41.296
25.923
24.447
15.159
Table 3: GPU inference efficiency at 1080×1920 output resolution on 201 real input frames using a single NVIDIA A100-80GB GPU.
Variant
DOVER ↑
MS ↑
KV (GB) ↓
FLOPs ↓
Uniform ( 3R )
0.5601
0.9801
6.44
1.00×
Uniform ( R )
0.5261
0.9722
2.15
0.27×
Ours †
0.5130
0.9748
2.17
0.28×
Ours
0.5587
0.9785
2.17
0.28×
Table 7
Decoder
PSNR ↑
FPS ↑
Mem. (GB)
Wan2.1 VAE
38.5332
4.51
22.13
Wan2.2 VAE
39.1866
3.77
25.80
TCDecoder
36.8953
33.06
5.56
SwiftVR ReAE
30.9792
141.35
18.11
Ours ‡
34.4523
93.20
1.28
Ours
36.2525
91.71
1.33
Table 6: Decoder quality and efficiency. ‡ : without LR conditioning. Best and second-best values are bold and underlined.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: Blockwise streaming inference with ReCaVSR. Each block uses one DiT forward pass to generate SR latents, which are decoded with LR observations from the same block. Green dotted paths recycle the generated SR latents; orange paths carry the historical KV selected for each layer; gray paths preserve the LR projector and decoder states. The first block starts with zero recycled input and an empty DiT KV cache.
Generator
LR-conditioned decoder
Backbone
Wan2.2-TI2V-5B
Architecture
FlashDecoder
DiT layers
30
Backbone layers
12
Hidden dimension
3072
Refinement layers
2
Latent channels
48
Hidden dimension
512
Patch size (T,H,W)
1×2×2
Latent projection
48→512
LoRA rank
512
LR stem
768→48
Appendix
Table B.1: Architecture configurations of ReCaVSR.
Metric
DOVE
SeedVR2-3B
SwiftVR
FlashVSR Tiny
ReCaVSR
MOS-Q ↑
3.15
3.40
3.32
3.78
3.84
MOS-D ↑
3.17
3.35
3.15
3.72
3.81
MOS-T ↑
3.62
3.29
2.98
3.69
3.87
Appendix
Table C.2: Human evaluation on 35 VideoLQ videos with 15 raters. MOS-Q, MOS-D, and MOS-T measure overall visual quality, fine-detail quality, and temporal stability, respectively. Best results are shown in bold.
720p
1080p
1440p
4K
Method
FPS ↑
GB ↓
FPS ↑
GB ↓
FPS ↑
GB ↓
FPS ↑
GB ↓
RealViformer
31.88
5.44
14.87
12.36
8.69
21.66
3.88
39.03
DOVE
1.53
30.52
0.56
41.08
0.24
57.12
0.15 †
42.11 †
SeedVR2-3B
2.65
75.05
0.79 †
73.89 †
0.44 †
74.88 †
0.15 †
61.46 †
Stream-DiffVSR
1.76
10.24
0.85
25.92
0.47
61.67
0.17 †
44.39 †
FlashVSR Tiny
16.52
12.90
7.80
24.45
4.41
40.61
1.30
67.99
Appendix
Table C.3: Resolution scaling on one A100-80GB. GPU throughput and peak allocated memory for 201 input frames. Memory is reported in decimal GB.
Action
DiT layers
Layers
Slots per layer
∅
1, 2, 4, 6, 9, 22, 25, 27, 29, 30
10
0
W1
3, 28
2
1
W2
5, 7, 13, 17, 21, 24, 26
7
2
W4
8, 11, 20
3
4
A
15
1
2
W2+A
10, 19, 23
3
4
Appendix
Table C.4: Deployed layer-wise historical cache policy. Layer indices are one-based; reserved slots count latent positions per layer.
Figure C.2: Additional visual comparisons. Each example shows the input frame with a marked region and the corresponding crops from VSR methods. The scenes include building facades, masonry, and construction machinery.
Diffusion-based video super-resolution (VSR) methods deliver strong perceptual quality but are often unsuitable for latency-sensitive scenarios due to reliance on future frames and expensive multi-step denoising. We propose Stream-DiffVSR, a causally conditioned diffusion framework for efficient online VSR. Operating strictly on past frames, Stream-DiffVSR integrates a four-step distilled denoiser for fast inference, an Auto-regressive Temporal Guidance (ARTG) module that injects motion-aligned cues during latent denoising, and a lightweight temporal-aware decoder with a Temporal Processor Module (TPM) to enhance detail and temporal coherence. Unlike chunk-wise streaming inference, our strictly frame-by-frame causal design avoids sequence-level waiting, substantially reducing time-to-first-frame and end-to-end latency. Stream-DiffVSR processes 720p frames in 0.328 seconds on an RTX 4090 and consistently outperforms prior diffusion-based baselines. Compared with the online state-of-the-art TMP, it improves perceptual quality (LPIPS +0.095). Compared with prior diffusion-based VSR methods such as MGLD-VSR, it reduces per-frame runtime by over 130x. Moreover, Stream-DiffVSR substantially lowers time-to-first-frame for diffusion-based VSR, reducing initial delay from over 4600 seconds to 0.328 seconds, making diffusion-based VSR markedly more practical for low-latency online and streaming deployment. Project page: https://jamichss.github.io/stream-diffvsr-project-page/
Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang +4
National Yang Ming Chiao Tung University · Shanda AI Research Tokyo · MediaTek Inc.
Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at https://github.com/kopperx/RelayVSR.
Video super-resolution (VSR) using large-scale Diffusion Transformer (DiT) priors achieves exceptional perceptual quality but is often impractical due to the quadratic computational cost of processing dense spatio-temporal token sequences. Existing efficiency-oriented methods risk irreversible detail loss and temporal flickering, a vulnerability especially pronounced in one-step diffusion models. To address this, we propose TRaM-VSR, a Token Routing and Merging framework for adaptive token allocation, leveraging both context-aware video priors and network-level priors. First, token importance is estimated by fusing motion-sensitive temporal cues with semantic text similarity, isolating dynamic objects and structural boundaries. Next, this importance is further calibrated and adjusted by an offline planner to guide routing across optimally grouped network blocks. Technically, within each routed group, structurally critical tokens are processed in a high-fidelity local stream, while less informative tokens are aggregated into a compact global stream, both modulated by network depth and aligned with the multigranular nature of diffusion models. Extensive experiments show that TRaM-VSR accelerates inference significantly while preserving state-of-the-art reconstruction quality and robust temporal consistency. The code is available at https://github.com/Ree1s/TRaM-VSR.
Sicheng Gao, Zhuyun Zhou, Yixuan Liu +3
Advanced Micro Devices Inc. · Computer Vision Lab, CAIDAS & IFI, University of Würzburg