cs.LGSep 30, 2026

Semantic-Aware Joint Source-Channel Optimization for Encoder-Agnostic Digital Video Communication

Authors: Xiangben Zhu, Caili Guo, Yang Yang, Chuanhong Liu, Meiyi Zhu

Organizations: Beijing Key Laboratory of Network System Architecture and Convergence, School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China · China Mobile (Suzhou) Software Technology Company Limited, Suzhou 215011, China · Department of Engineering, King’s College London, London WC2R 2LS, U.K.

Abstract

Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most existing approaches rely on computationally intensive deep learning-based video encoders and decoders, which hinders their deployment in resource-constrained scenarios. To address this issue, we propose a lightweight semantic-aware joint source-channel optimization (SAJSCO) scheme that can be integrated into existing digital video communication systems as a plug-in module. Specifically, we develop a video communication system model in which the transmitter jointly optimizes source and channel coding parameters based on the inter-frame semantic importance of the input video and estimated channel state information. On this basis, we formulate an optimization problem that maximizes semantic importance weighted video reconstruction quality under a maximum bitrate constraint. To solve it, we first quantify inter-frame semantic importance using a cosine similarity-based metric with a shifted window mechanism. We then develop a multi-actor proximal policy optimization (MPPO) algorithm to solve the formulated problem by jointly adapting the source compression rate and channel coding rate. The learned policy can be directly applied to different video encoders without encoder-specific retraining or fine-tuning. SAJSCO achieves Bjøntegaard Delta rate reductions of 34.86% and 18.01% when integrated with H.265, a conventional video encoder, and DCVC-RT, a deep learning-based video encoder. Over-the-air experiments on a hardware testbed further demonstrate a PSNR gain of up to 1.448 dB with H.265 and an LPIPS reduction of up to 0.033 with DCVC-RT compared with the respective best-performing fixed-parameter baselines.

Figures & tables

Explore similar work

CardsList
  1. ChronoSC: Task-Oriented Semantic Communication via Temporal-to-Color Encoding

    May 11, 2026Phuc H. Nguyen, Trung T. Nguyen, Quy N. Duong +1Video CompressionSemantic Communications

  2. Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission

    Sep 14, 2026Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du +1Video CompressionFine-Grained Video Understanding

  3. Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance

    Dec 8, 2025Naifu Xue, Zhaoyang Jia, Jiahao Li +4Video CompressionVideo Coding