Real-time video super-resolution requires high spatio-temporal fidelity under strict latency constraints, challenging diffusion models due to their iterative sampling cost and limited temporal coordination. We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams. Our Cross-Step Attention mechanism reuses intermediate denoising features across adjacent frames and diffusion steps, enabling temporal information exchange without explicit temporal modeling. We further introduce Trajectory-Coupled Diffusion Scheduling, which aligns adjacent diffusion states and provides cleaner intermediate representations for cross-step conditioning, improving temporal coherence. These components are integrated into a streaming inference pipeline that incrementally propagates latent states across frames, reducing the effective computational complexity from O(N⋅S) to O(N+S) for N frames and S diffusion steps. Experiments on REDS4 and YouHQ40-Test demonstrate improved perceptual quality and temporal realism while maintaining frame-wise stability. Our method achieves over 40 FPS at 512×512 resolution after cold start, enabling real-time VSR without explicit temporal modeling.
Figures & tables
Figure 1: Cross-Step Attention for stream-oriented diffusion. Frame n at step s attends to frame n−1 at step s+1 for efficient feature reuse and temporal consistency.
Method
PSNR
SSIM
LPIPS
DISTS
FID
tLPIPS
tDISTS
FVD
REDS4
Bicubic
26.13
0.73
0.453
0.186
5.80
0.023
0.003
18.89
Real-ESRGAN
23.35
0.67
0.242
0.105
2.97
0.006
0.007
23.48
Stable-VSR
28.15
0.80
0.101
0.046
0.06
0.004
0.006
2.59
LDM
24.66
0.69
0.214
0.097
2.20
0.030
0.025
11.71
LDMCS
24.07
0.65
0.198
0.092
1.43
0.044
0.027
8.84
Table 1: Quantitative comparison on REDS4 and YouHQ40-Test with 30 diffusion steps. LDMCS denotes Cross-Step Attention , while LDMCS∗ is fine-tuned on REDS for 15 epochs.
Figure 2: REDS4 examples with different methods.
Method
Runtime ↓
Steps/s ↑
FPS ↑
512 2
1024 2
512 2
1024 2
512 2
1024 2
Bicubic
1.5
3
–
–
20
10
Real-ESRGAN
3.3
9.1
–
–
9
3.3
StableVSR
120
750
–
–
–
–
LDM
13
36
67
25
2.2
0.8
LDMCS
20
75
44
12
1.5
0.4
Table 2: Inference efficiency at 5122 and 10242 resolutions.
Method
REDS4
YouHQ40-Test
LPIPS ↓
DISTS ↓
FID ↓
FVD ↓
LPIPS ↓
DISTS ↓
FID ↓
FVD ↓
LDM5
0.300
0.144
5.72
21.84
0.302
0.154
3.96
76.10
LDM10
0.263
0.121
3.74
16.44
0.275
0.137
2.31
71.06
LDM15
0.240
0.109
2.98
14.20
0.259
0.130
1.91
69.17
LDM20
0.228
0.103
2.61
13.18
0.250
0.126
1.71
67.95
LDM25
0.220
0.100
2.35
12.21
0.245
0.124
1.61
67.30
Table 3: Effect of diffusion step budgets on reconstruction and temporal quality.