Real-time video super-resolution requires high spatio-temporal fidelity under strict latency constraints, challenging diffusion models due to their iterative sampling cost and limited temporal coordination. We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams. Our Cross-Step Attention mechanism reuses intermediate denoising features across adjacent frames and diffusion steps, enabling temporal information exchange without explicit temporal modeling. We further introduce Trajectory-Coupled Diffusion Scheduling, which aligns adjacent diffusion states and provides cleaner intermediate representations for cross-step conditioning, improving temporal coherence. These components are integrated into a streaming inference pipeline that incrementally propagates latent states across frames, reducing the effective computational complexity from O(N⋅S) to O(N+S) for N frames and S diffusion steps. Experiments on REDS4 and YouHQ40-Test demonstrate improved perceptual quality and temporal realism while maintaining frame-wise stability. Our method achieves over 40 FPS at 512×512 resolution after cold start, enabling real-time VSR without explicit temporal modeling.
Figures & tables
Figure 1: Cross-Step Attention for stream-oriented diffusion. Frame n at step s attends to frame n−1 at step s+1 for efficient feature reuse and temporal consistency.
Method
PSNR
SSIM
LPIPS
DISTS
FID
tLPIPS
tDISTS
FVD
REDS4
Bicubic
26.13
0.73
0.453
0.186
5.80
0.023
0.003
18.89
Real-ESRGAN
23.35
0.67
0.242
0.105
2.97
0.006
0.007
23.48
Stable-VSR
28.15
0.80
0.101
0.046
0.06
0.004
0.006
2.59
LDM
24.66
0.69
0.214
0.097
2.20
0.030
0.025
11.71
LDMCS
24.07
0.65
0.198
0.092
1.43
0.044
0.027
8.84
Table 1: Quantitative comparison on REDS4 and YouHQ40-Test with 30 diffusion steps. LDMCS denotes Cross-Step Attention , while LDMCS∗ is fine-tuned on REDS for 15 epochs.
Figure 2: REDS4 examples with different methods.
Method
Runtime ↓
Steps/s ↑
FPS ↑
512 2
1024 2
512 2
1024 2
512 2
1024 2
Bicubic
1.5
3
–
–
20
10
Real-ESRGAN
3.3
9.1
–
–
9
3.3
StableVSR
120
750
–
–
–
–
LDM
13
36
67
25
2.2
0.8
LDMCS
20
75
44
12
1.5
0.4
Table 2: Inference efficiency at 5122 and 10242 resolutions.
Method
REDS4
YouHQ40-Test
LPIPS ↓
DISTS ↓
FID ↓
FVD ↓
LPIPS ↓
DISTS ↓
FID ↓
FVD ↓
LDM5
0.300
0.144
5.72
21.84
0.302
0.154
3.96
76.10
LDM10
0.263
0.121
3.74
16.44
0.275
0.137
2.31
71.06
LDM15
0.240
0.109
2.98
14.20
0.259
0.130
1.91
69.17
LDM20
0.228
0.103
2.61
13.18
0.250
0.126
1.71
67.95
LDM25
0.220
0.100
2.35
12.21
0.245
0.124
1.61
67.30
Table 3: Effect of diffusion step budgets on reconstruction and temporal quality.
Diffusion-based video super-resolution (VSR) methods deliver strong perceptual quality but are often unsuitable for latency-sensitive scenarios due to reliance on future frames and expensive multi-step denoising. We propose Stream-DiffVSR, a causally conditioned diffusion framework for efficient online VSR. Operating strictly on past frames, Stream-DiffVSR integrates a four-step distilled denoiser for fast inference, an Auto-regressive Temporal Guidance (ARTG) module that injects motion-aligned cues during latent denoising, and a lightweight temporal-aware decoder with a Temporal Processor Module (TPM) to enhance detail and temporal coherence. Unlike chunk-wise streaming inference, our strictly frame-by-frame causal design avoids sequence-level waiting, substantially reducing time-to-first-frame and end-to-end latency. Stream-DiffVSR processes 720p frames in 0.328 seconds on an RTX 4090 and consistently outperforms prior diffusion-based baselines. Compared with the online state-of-the-art TMP, it improves perceptual quality (LPIPS +0.095). Compared with prior diffusion-based VSR methods such as MGLD-VSR, it reduces per-frame runtime by over 130x. Moreover, Stream-DiffVSR substantially lowers time-to-first-frame for diffusion-based VSR, reducing initial delay from over 4600 seconds to 0.328 seconds, making diffusion-based VSR markedly more practical for low-latency online and streaming deployment. Project page: https://jamichss.github.io/stream-diffvsr-project-page/
Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang +4
National Yang Ming Chiao Tung University · Shanda AI Research Tokyo · MediaTek Inc.
Diffusion-based models have shown strong performance in video super-resolution (VSR) and video frame interpolation (VFI). However, their role in the coupled space-time video super-resolution (STVSR) setting remains limited. Existing diffusion-based STVSR approaches suffer from two issues: (1) low inference efficiency and (2) insufficient utilization of spatiotemporal information. These limitations impede deployment. To address these issues, we introduce DiffST, an efficient spatiotemporal-aware video diffusion framework for real-world STVSR. To improve efficiency, we adapt a pre-trained diffusion model for one-step sampling and process the entire video directly rather than operating on individual frames. Furthermore, to enhance spatiotemporal information utilization, we introduce cross-frame context aggregation (CFCA) and video representation guidance (VRG). The CFCA module aggregates information across multiple keyframes to produce intermediate frames. The VRG module extracts video-level global features to guide the diffusion process. Extensive experiments show that DiffST obtains leading results on real-world STVSR tasks. It also maintains high inference efficiency, running about 17× faster than previous diffusion-based STVSR methods. Code is available at: https://github.com/zhengchen1999/DiffST.
Zheng Chen, Ruofan Yang, Jin Han +5
Shanghai Jiao Tong University · Huawei Noah’s Ark Lab · Duke University +1
Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At 1080×1920 output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72× faster while using 38.0% less peak allocated memory than FlashVSR Tiny. The code is available at https://github.com/kopperx/ReCaVSR.