Organizations: School of Artificial Intelligence, Shenzhen Polytechnic University, Shenzhen, China · Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong, China · Department of Computer Science, Hong Kong Chu Hai College, Hong Kong, China · College of Eilte Engineers, Dongguan University of Technology, Dongguan, China · Department of Computer Science, City University of Hong Kong, Hong Kong, China
Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to the disruption of temporal correlations. To address these challenges, we propose the Spatial-Temporal Multi-scale Network (STM-Net), a novel framework specifically tailored for compressed SCV enhancement. Our approach integrates three complementary components: a Prior-Guided Spatio-Temporal Dispatcher (PG-STD) that routes input into three parallel streams to avoid feature contamination, a Bidirectional Temporal Feature Extraction (BTFE) module that adaptively handles abrupt transitions without explicit detection, and a Cascaded Multi-scale Feature Distillation (CMFD) module that preserves critical high-frequency details. Experimental results demonstrate that STM-Net outperforms state-of-the-art methods in both objective metrics and subjective visual quality, providing a robust solution for screen content artifacts. Code is available at https://github.com/HUANGZiyin1/STM-Net.
TABLE I: Δ PSNR ( Δ P) (dB) and Δ SSIM ( Δ S) ( ×10−3 ) at QP=37, average results for QP=32, 27, 22, and BD-Rate comparison.
Fig. 2: Δ PSNR curves for mixvideo ; dashed lines mark scene switches.
Fig. 3: Subjective visual quality comparison at QP=37. Top: Textual content enhancement on Sephora . Bottom: Graphical content enhancement on ChineseEditing . STM-Net restores clearer texts and sharper edges compared to state-of-the-art methods.
Model Structure
STM-Net
STM-Nscale
STM-NHF
STM-Net-S
Multi-scale Feature Extraction
✓
×
✓
✓
High-Frequency Processing
✓
✓
×
✓
Number of RBs
3
3
3
2
Number of MSFDBs
3
3
3
2
Δ PSNR (dB)
0.875
0.839
0.834
0.798
TABLE II: Ablation Study of Different Architecture Components in STM-Net at QP=37
Fig. 4: Visualization of Intermediate Features on Scene Switch. ( Left : BTFE Features Suppressed for Previous Frames; Middle : BTFE Features Emphasized for Future Frames, e.g.: left menu changed; Right : CMFD Features Emphasized for Screen Content Edges at Current Frame)
Implicit neural representations have emerged as a promising paradigm for video compression, with recent methods achieving competitive performance on natural video. However, screen content video -- common in remote desktop, online education, and cloud gaming -- exhibits distinct statistics: sharp edges, limited color palettes, and strong temporal redundancy. Existing neural representation methods, designed for natural scenes, lack mechanisms to exploit these properties, leaving substantial room for improvement. In this paper, we propose NeR-SC, a neural representation framework tailored for screen content video. Building on the SNeRV backbone, NeR-SC introduces three screen-content-specific modules: (i) a learnable color palette that models the discrete color structure of screen content by restricting the low-frequency sub-band to a learned color set; (ii) a multi-gate dense fusion module that replaces sequential feature fusion with dense, attention-gated cross-stage interaction; and (iii) an embedding-level frame skip strategy that bypasses redundant decoder invocations for static frames, with zero training overhead. Experiments on DSCVC and VCD show that NeR-SC achieves 40.32dB and 41.73dB average PSNR, outperforming representative neural video representation methods and, at low bitrates, surpassing H.264 and H.265. The skip strategy enables real-time decoding with no loss in quality.
Ruohan Shi, Jiaoyan Zhao, Haogang Feng
Management school The University of Sheffield · Undergraduate School of Artificial Intelligence Shenzhen Polytechnic University · School of Artificial Intelligence Shenzhen University of Information Technology
Methods based on implicit neural representations have demonstrated superior performance in Screen Content Image Super-Resolution (SCISR) . However, they overlooked the inherent frequency characteristics, leading to suboptimal performance. We propose a frequency decoupled framework (FDF) that rethinks SCISR from a phasor perspective by capturing structured energy in amplitude and relational continuity in phase, and jointly exploiting them with bespoke implicit representations to faithfully recover the regular textures and global configuration of Screen Content Image (SCI). Amplitude-Phase Factorization Network (APFN) first separates images into amplitude and phase streams, where Amplitude Clustering Module (ACM) organizes sparse yet high-energy amplitude responses into representative prototypes for periodic pattern extraction, while Phase Consistency Self-Attention (PCSA) progressively reinforces configuration through continuous consistency propagation. And Oscillation-Anharmonic Implicit Fitting Network (OAIF-Net) integrates periodic and coherent implicit representations for efficient exploitation of the periodic patterns and coherent context embedded in SCI. Experimental results show FDF achieves state-of-the-art SCISR performance at multiple scales across four public SCI datasets. Ablation experiments further demonstrate the effectiveness of each component in extracting and exploiting periodic patterns and coherent context.
Xufei Wang, Qicheng Zhang, Qi Wu +2
School of Electronic and Information Engineering, Anhui University, Hefei 230039, China
Perceptual quality enhancement of severely compressed videos remains challenging due to complex artifact patterns and substantial information loss. Recent diffusion models have demonstrated strong generative capability for visual restoration, but directly applying them to compressed video often ignores compression degradation characteristics and may introduce structure-inconsistent hallucinations. To address this issue, this paper presents a diffusion-based compressed video enhancement method, named DiffCVE. Coding Prior-enhanced Dual Conditioning (CPDC) branches are designed to jointly model compressed video and coding prior conditions, where coding priors including residuals and motion vectors provide complementary structural and motion guidance during the diffusion denoising process. To make the diffusion process aware of compression severity, a Compression Degradation Semantic Prompting (CDSP) mechanism is introduced to leverage QP-conditioned textual prompts together with LoRA fine-tuning. In addition, a Coding Prior-guided Weighted Fusion (CPWF) module is incorporated into the VAE decoder to fuse VAE encoder and coding prior encoder features with QP-predicted weights. Extensive experiments demonstrate the effectiveness of the proposed method in improving perceptual quality, especially under severe compression settings. The project page with enhanced video demonstrations is available at https://wqmaker.github.io/projects/DiffCVE/.
Wenqiang Xiao, Wenzhuo Ma, Junxi Zhang +1
School of Remote Sensing and Information Engineering, Wuhan University