cs.CVOct 8, 2026

Streaming-Aware Diffusion for Real-Time Video Super-Resolution via Cross-Step Attention

Authors: Harris Partaourides, Sotirios Chatzis

Organizations: Ethical AI Novelties Limassol, Cyprus · Cyprus University of Technology Limassol, Cyprus

Abstract

Real-time video super-resolution requires high spatio-temporal fidelity under strict latency constraints, challenging diffusion models due to their iterative sampling cost and limited temporal coordination. We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams. Our Cross-Step Attention mechanism reuses intermediate denoising features across adjacent frames and diffusion steps, enabling temporal information exchange without explicit temporal modeling. We further introduce Trajectory-Coupled Diffusion Scheduling, which aligns adjacent diffusion states and provides cleaner intermediate representations for cross-step conditioning, improving temporal coherence. These components are integrated into a streaming inference pipeline that incrementally propagates latent states across frames, reducing the effective computational complexity from O(N⋅S)O(N \cdot S) to O(N+S)O(N + S) for NN frames and SS diffusion steps. Experiments on REDS4 and YouHQ40-Test demonstrate improved perceptual quality and temporal realism while maintaining frame-wise stability. Our method achieves over 40 FPS at 512×512512 \times 512 resolution after cold start, enabling real-time VSR without explicit temporal modeling.

Figures & tables

Explore similar work

CardsList
  1. Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion

    Dec 29, 2025Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang +4Video Super-ResolutionDiffusion Model Inference Acceleration

  2. DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution

    May 13, 2026Zheng Chen, Ruofan Yang, Jin Han +5Video Diffusion ModelsVideo Super-Resolution