Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at https://github.com/kopperx/RelayVSR.
Figures & tables
Figure 1: Visual comparison and streaming schedule of RelayVSR. Top: RelayVSR outputs on an AI-generated video and a real-world video; the 29.3 FPS badge reports RelayVSR throughput at 1080p on one NVIDIA A100 80GB in dual-endpoint mode with a 15-frame keyframe interval. Bottom left: the sparse-reference schedule for first-keyframe and dual-endpoint inference, with a fixed playback delay in the latter.
Figure 2: RelayVSR framework and training. Top: the large generative model learns frame-wise target keyframe latents in Stage 1 and undergoes one-step GAN distillation in Stage 2. Bottom left: the lightweight VSR network super-resolves every frame by jointly attending to persistent keyframe and rolling video memories. Bottom right: VARO uses reinforcement learning to update the large generative model with a system-level reward on videos super-resolved by the fixed lightweight VSR network and a reference-level reward on decoded keyframes.
Dataset
Metric
RealViformer
STAR
DOVE
SeedVR2
SparkVSR
SwiftVR
FlashVSR
Ours
REDS30
PSNR ↑
23.32
22.14
23.28
22.17
23.29
21.33
21.41
22.77
SSIM ↑
0.5967
0.5432
0.6103
0.5860
0.6050
0.5234
0.5399
0.5874
LPIPS ↓
0.3043
0.4967
0.3773
0.3158
0.3909
0.3564
0.3311
0.3051
NIQE ↓
3.0804
5.2237
4.2282
3.5157
4.4496
3.4153
2.9504
3.5398
MUSIQ ↑
59.12
37.95
50.44
57.65
47.58
63.77
56.01
54.30
CLIP-IQA ↑
0.3236
0.2132
0.2904
0.3041
0.2755
0.3914
0.3160
0.3161
Table 1: Quantitative comparison on synthetic and real-world VSR benchmarks. The best and second-best results are marked in red and blue , respectively.
Figure 3: Qualitative VSR comparisons. The top two examples are real-world videos and the bottom example is an AIGC video. ∗ marks the keyframe sample.
Figure 4: Keyframe and non-keyframe comparison on a real-world video. The ∗ symbol marks keyframes; the other columns are non-keyframes.
Method
FPS ↑
Peak Mem.
First output
(GB) ↓
(s)
DOVE
0.56
41.08
356.801
SeedVR2-3B
0.79
73.89
255.199
Stream-DiffVSR
0.85
25.92
0.705
SwiftVR
23.48
31.36
1.107
FlashVSR
7.80
24.45
2.830
Table 2: Inference efficiency at 1080×1920 output resolution on one NVIDIA A100-80G using 201-frame clips. Timing excludes input wait and operations outside model inference.
Figure 5: Quality–efficiency trade-off. Left to right: Δ=5,10,15,20,25 .
Ref. memory
Video memory
MUSIQ ↑
DOVER ↑
MS ↑
✗
✓
56.80
0.4760
0.9911
✓
✗
63.65
0.5110
0.9769
✓
✓
63.72
0.5290
0.9830
Table 3: Memory ablation on UDM10 before VARO ( Δ=15 , dual-endpoint). ✓/✗ indicates enabled/disabled memory; MS is motion smoothness.
Variant
MUSIQ ↑
DOVER ↑
MS ↑
All
Key
w/o post-training
63.72
64.05
0.5290
0.9830
w/ reference-level reward only
64.20
64.95
0.5355
0.9838
w/ system-level reward only
64.86
64.42
0.5400
0.9851
Ours
65.32
65.18
0.5448
0.9900
Table 4: VARO ablation on UDM10 ( Δ=15 , dual-endpoint). MS is motion smoothness.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Large generative model
Pretrained backbone
Wan2.2-TI2V-5B
LoRA configuration
Rank r=512 , scaling factor α=512
LR projection hidden widths
768, 3,072
Temporal attention window
4 keyframes
Spatial attention window
22×40
Appendix
Table 5: Model configuration. Parameter counts include trainable parameters only; spatial attention windows are measured in tokens.
Setting
Flow matching
One-step GAN
Lightweight VSR
Initialization
Wan2.2
Adapted generator
Random
Video frames per clip
8
8
16
HR crop
704×1280
704×1280
1024×1536
Global video / image batch
32 / 96
32 / 96
64 / —
Updates
75K
5K
60K
GPUs
16
16
16
Appendix
Table 6: Base training settings. Batch sizes are global. GAN learning rate and weight decay apply to both generator and discriminator.
Setting
Value
Rollout
Training clip
46 frames at 704×1280 HR
Reference schedule
Δ=15 , dual-endpoint
Candidates per clip
4
Optimization
Optimizer and learning rate
AdamW, 5×10−6
Appendix
Table 7: VARO training settings. Batch size counts LR–HR clips. Weight tuples follow the component order in Eq. ( 8 ).
Method
720p
1080p
1440p
2160p
FPS ↑
GB ↓
FPS ↑
GB ↓
FPS ↑
GB ↓
FPS ↑
GB ↓
RealBasicVSR
41.758
11.04
18.502
24.81
10.671
38.04
4.888
66.54
RealViformer
31.884
5.44
14.873
12.36
8.689
21.66
3.880
39.03
Stream-DiffVSR
1.759
10.24
0.849
25.92
0.468
61.67
0.169
44.39
FlashVSR
16.518
12.90
7.799
24.45
4.410
40.61
1.302
67.99
SwiftVR
51.056
21.25
23.483
31.36
13.367
40.74
5.773
65.26
Appendix
Table 8: Inference efficiency across output resolutions. Model FPS and peak allocated memory (GB) for 201-frame 4× VSR on one A100-80G. RelayVSR uses Δ=15 . Best and second-best entries in each column are bold and underlined, respectively.
Mode
Δ
720p
1080p
1440p
2160p
Dual-endpoint
5
36.02
18.10
10.58
4.91
10
50.80
25.60
15.24
7.12
15
58.32
29.29
17.59
8.23
20
64.34
32.20
19.56
9.17
25
68.10
34.13
20.75
9.69
First-keyframe
5
40.55
20.42
11.89
5.55
Appendix
Table 9: RelayVSR throughput across reference intervals. Model FPS for 201-frame videos at four output resolutions on one A100-80G.
No-reference quality
Temporal indicators
Method
NIQE ↓
MUSIQ ↑
CLIP-IQA ↑
DOVER ↑
SC ↑
BC ↑
MS ↑
RealViformer
5.0770
46.17
0.3819
0.5053
0.8846
0.9198
0.9848
Stream-DiffVSR
4.0491
54.52
0.4694
0.5704
0.8830
0.9161
0.9813
SwiftVR
4.2822
53.89
0.4640
0.5512
0.8831
0.9177
0.9824
FlashVSR
3.8160
55.15
0.4747
0.5895
0.8829
0.9144
0.9802
RelayVSR
4.1396
54.72
0.4701
0.5911
0.8837
0.9157
0.9827
Appendix
Table 10: LongVSR60 quality. The benchmark contains 30 real-world and 30 AI-generated videos. NIQE, MUSIQ, and CLIP-IQA use 100 uniformly sampled frames per video; SC, BC, and MS are supplementary temporal indicators. Best and second-best values are bold and underlined, respectively.
Baseline
Overall quality ↑
Content fidelity ↑
Temporal stability ↑
DOVE
+53.6
-8.9
-8.7
SeedVR2
+46.8
+13.5
+13.1
SparkVSR
+43.9
+21.9
+5.9
SwiftVR
+26.7
+13.2
+11.9
FlashVSR
-3.3
+9.0
+17.0
Appendix
Table 11: User study. GSB scores (%) compare RelayVSR with each baseline. Positive scores favor RelayVSR; negative scores favor the baseline.
Figure 6: Detail propagation between sparse references. Left: the source frame locates the fixed crop. Right: keyframes 00 and 15 (marked ∗ ) and five intervening frames compare the same license-plate region in the LR input, RealViformer, FlashVSR, and RelayVSR.
Figure 7: Additional visual comparisons I. Three examples highlight printed backdrop text and a face, underwater coral structure, and a dog’s face and fur. Left: full-frame context with marked crops. Right: aligned crops from the input and six VSR methods, including RelayVSR.
Figure 8: Additional visual comparisons II. Aligned crops show lettering on a street sign, a squirrel, and a car partly seen through foliage. Each example includes full-frame context at left and comparisons with earlier and recent VSR methods at right.
Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At 1080×1920 output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72× faster while using 38.0% less peak allocated memory than FlashVSR Tiny. The code is available at https://github.com/kopperx/ReCaVSR.
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
Diffusion-based video super-resolution (VSR) methods deliver strong perceptual quality but are often unsuitable for latency-sensitive scenarios due to reliance on future frames and expensive multi-step denoising. We propose Stream-DiffVSR, a causally conditioned diffusion framework for efficient online VSR. Operating strictly on past frames, Stream-DiffVSR integrates a four-step distilled denoiser for fast inference, an Auto-regressive Temporal Guidance (ARTG) module that injects motion-aligned cues during latent denoising, and a lightweight temporal-aware decoder with a Temporal Processor Module (TPM) to enhance detail and temporal coherence. Unlike chunk-wise streaming inference, our strictly frame-by-frame causal design avoids sequence-level waiting, substantially reducing time-to-first-frame and end-to-end latency. Stream-DiffVSR processes 720p frames in 0.328 seconds on an RTX 4090 and consistently outperforms prior diffusion-based baselines. Compared with the online state-of-the-art TMP, it improves perceptual quality (LPIPS +0.095). Compared with prior diffusion-based VSR methods such as MGLD-VSR, it reduces per-frame runtime by over 130x. Moreover, Stream-DiffVSR substantially lowers time-to-first-frame for diffusion-based VSR, reducing initial delay from over 4600 seconds to 0.328 seconds, making diffusion-based VSR markedly more practical for low-latency online and streaming deployment. Project page: https://jamichss.github.io/stream-diffvsr-project-page/
Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang +4
National Yang Ming Chiao Tung University · Shanda AI Research Tokyo · MediaTek Inc.