cs.CVSep 30, 2026

Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement

Authors: Ziyin Huang, Sik-Ho Tsang, Xinyuan Qin, Yui-Lam Chan, Xueling Zhou, Feiyu Chen

Organizations: School of Artificial Intelligence, Shenzhen Polytechnic University, Shenzhen, China · Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong, China · Department of Computer Science, Hong Kong Chu Hai College, Hong Kong, China · College of Eilte Engineers, Dongguan University of Technology, Dongguan, China · Department of Computer Science, City University of Hong Kong, Hong Kong, China

Abstract

Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to the disruption of temporal correlations. To address these challenges, we propose the Spatial-Temporal Multi-scale Network (STM-Net), a novel framework specifically tailored for compressed SCV enhancement. Our approach integrates three complementary components: a Prior-Guided Spatio-Temporal Dispatcher (PG-STD) that routes input into three parallel streams to avoid feature contamination, a Bidirectional Temporal Feature Extraction (BTFE) module that adaptively handles abrupt transitions without explicit detection, and a Cascaded Multi-scale Feature Distillation (CMFD) module that preserves critical high-frequency details. Experimental results demonstrate that STM-Net outperforms state-of-the-art methods in both objective metrics and subjective visual quality, providing a robust solution for screen content artifacts. Code is available at https://github.com/HUANGZiyin1/STM-Net.

Figures & tables

Explore similar work

CardsList
  1. NeR-SC: Adapting Neural Video Representation to Screen Content

    May 26, 2026Ruohan Shi, Jiaoyan Zhao, Haogang FengVideo CompressionNeural Representations

  2. Frequency Decoupled Framework for Screen Content Image Super-Resolution

    Jun 8, 2026Xufei Wang, Qicheng Zhang, Qi Wu +2

  3. DiffCVE: Diffusion-based Compressed Video Enhancement

    Jul 8, 2026Wenqiang Xiao, Wenzhuo Ma, Junxi Zhang +1Video CompressionPerceptual Quality