Organizations: School of Artificial Intelligence, Shenzhen Polytechnic University, Shenzhen, China · Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong, China · Department of Computer Science, Hong Kong Chu Hai College, Hong Kong, China · College of Eilte Engineers, Dongguan University of Technology, Dongguan, China · Department of Computer Science, City University of Hong Kong, Hong Kong, China
Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to the disruption of temporal correlations. To address these challenges, we propose the Spatial-Temporal Multi-scale Network (STM-Net), a novel framework specifically tailored for compressed SCV enhancement. Our approach integrates three complementary components: a Prior-Guided Spatio-Temporal Dispatcher (PG-STD) that routes input into three parallel streams to avoid feature contamination, a Bidirectional Temporal Feature Extraction (BTFE) module that adaptively handles abrupt transitions without explicit detection, and a Cascaded Multi-scale Feature Distillation (CMFD) module that preserves critical high-frequency details. Experimental results demonstrate that STM-Net outperforms state-of-the-art methods in both objective metrics and subjective visual quality, providing a robust solution for screen content artifacts. Code is available at https://github.com/HUANGZiyin1/STM-Net.
TABLE I: Δ PSNR ( Δ P) (dB) and Δ SSIM ( Δ S) ( ×10−3 ) at QP=37, average results for QP=32, 27, 22, and BD-Rate comparison.
Fig. 2: Δ PSNR curves for mixvideo ; dashed lines mark scene switches.
Fig. 3: Subjective visual quality comparison at QP=37. Top: Textual content enhancement on Sephora . Bottom: Graphical content enhancement on ChineseEditing . STM-Net restores clearer texts and sharper edges compared to state-of-the-art methods.
Model Structure
STM-Net
STM-Nscale
STM-NHF
STM-Net-S
Multi-scale Feature Extraction
✓
×
✓
✓
High-Frequency Processing
✓
✓
×
✓
Number of RBs
3
3
3
2
Number of MSFDBs
3
3
3
2
Δ PSNR (dB)
0.875
0.839
0.834
0.798
TABLE II: Ablation Study of Different Architecture Components in STM-Net at QP=37
Fig. 4: Visualization of Intermediate Features on Scene Switch. ( Left : BTFE Features Suppressed for Previous Frames; Middle : BTFE Features Emphasized for Future Frames, e.g.: left menu changed; Right : CMFD Features Emphasized for Screen Content Edges at Current Frame)
Management school The University of Sheffield · Undergraduate School of Artificial Intelligence Shenzhen Polytechnic University · School of Artificial Intelligence Shenzhen University of Information Technology