Organizations: Beijing Key Laboratory of Network System Architecture and Convergence, School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China · China Mobile (Suzhou) Software Technology Company Limited, Suzhou 215011, China · Department of Engineering, King’s College London, London WC2R 2LS, U.K.
Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most existing approaches rely on computationally intensive deep learning-based video encoders and decoders, which hinders their deployment in resource-constrained scenarios. To address this issue, we propose a lightweight semantic-aware joint source-channel optimization (SAJSCO) scheme that can be integrated into existing digital video communication systems as a plug-in module. Specifically, we develop a video communication system model in which the transmitter jointly optimizes source and channel coding parameters based on the inter-frame semantic importance of the input video and estimated channel state information. On this basis, we formulate an optimization problem that maximizes semantic importance weighted video reconstruction quality under a maximum bitrate constraint. To solve it, we first quantify inter-frame semantic importance using a cosine similarity-based metric with a shifted window mechanism. We then develop a multi-actor proximal policy optimization (MPPO) algorithm to solve the formulated problem by jointly adapting the source compression rate and channel coding rate. The learned policy can be directly applied to different video encoders without encoder-specific retraining or fine-tuning. SAJSCO achieves Bjøntegaard Delta rate reductions of 34.86% and 18.01% when integrated with H.265, a conventional video encoder, and DCVC-RT, a deep learning-based video encoder. Over-the-air experiments on a hardware testbed further demonstrate a PSNR gain of up to 1.448 dB with H.265 and an LPIPS reduction of up to 0.033 with DCVC-RT compared with the respective best-performing fixed-parameter baselines.
Figures & tables
Fig. 1: System model of semantic-aware joint source–channel optimization for video communication
Fig. 2: Illustration of the proposed inter-GOP semantic importance calculation method.
Fig. 3: Architecture of semantic-aware joint source–channel optimization for video communication
Parameter
Value
Parameter
Value
Number of GOPs, N
64
Shifted-window size, M
8
Training episodes
500
PPO update epochs, Ep
8
Discount factor, η
0.999
Clip parameter, ϵ
0.2
Optimizer
Adam
Penalty constant, ρ
50
Actor learning rate, δa
1×10−3
Critic learning rate, δc
3×10−3
Hyper-parameter, c1
10
Hyper-parameter, c2
0.3
TABLE I: Simulation Parameters
Fig. 4: Ablation study of the proposed MPPO algorithm.
Fig. 5: Performance comparison under different SNR conditions on the ActivityNet dataset with the H.265 encoder.
Fig. 6: PSNR and LPIPS performance versus bitrate on the ActivityNet dataset with the H.265 encoder under the RadioML channel setting.
Fig. 7: Performance comparison under the RadioML CSI setting on the HEVC testset.
Fig. 8: Visualization comparison of reconstructed video frames under different transmission schemes.
Method
MACs
Params
Speed
H.265
none
none
Enc. 42.41 fps, Dec. 25 fps
SAJSCO+H.265
318.97M
2.23M
Enc. 35.73 fps, Dec. 25 fps
DCVC-RT
385G
20.7M
Enc. 109 fps, Dec. 90 fps
SAJSCO+DCVC-RT
385.32G
22.93M
Enc. 66.73 fps, Dec. 90 fps
TABLE II: Comparison of Complexity
Fig. 9: Visualization of the state space and the corresponding action selection of the proposed MPPO algorithm.
Fig. 10: Hardware architecture of the proposed USRP-based communication testbed.
Parameter
Value
Parameter
Value
Carrier Frequency
5.5 GHz
LDPC Block Length
1800 bits
Sampling Rate
1 MSps
TX Gain (Seg. 1 & 4)
50 dB
Symbol Rate
62.5 kSps
TX Gain (Seg. 2 & 3)
60 dB
Bandwidth
125 kHz
RX Gain
56 dB
Modulation
BPSK
Antenna Model
VERT2450
Samples per Symbol
16
Antenna Placement
Vertical
TABLE III: Experimental Parameters
Method
PSNR
WPSNR
LPIPS
WLPIPS
Dec. Rate
H.265+SAJSCO
28.905
28.922
0.358
0.351
93.75%
H.265 ( r=1/3 )
25.358
25.578
0.518
0.508
84.38%
H.265 ( r=1/2 )
27.457
27.514
0.400
0.403
89.06%
H.265 ( r=2/3 )
27.262
27.425
0.527
0.544
92.19%
DCVC+SAJSCO
27.741
27.585
0.306
0.311
75.00%
DCVC ( r=1/3 )
26.455
26.084
0.339
0.351
70.31%
TABLE IV: Performance Comparison on the USRP Testbed