Organizations: Beijing Key Laboratory of Network System Architecture and Convergence, School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China · China Mobile (Suzhou) Software Technology Company Limited, Suzhou 215011, China · Department of Engineering, King’s College London, London WC2R 2LS, U.K.
Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most existing approaches rely on computationally intensive deep learning-based video encoders and decoders, which hinders their deployment in resource-constrained scenarios. To address this issue, we propose a lightweight semantic-aware joint source-channel optimization (SAJSCO) scheme that can be integrated into existing digital video communication systems as a plug-in module. Specifically, we develop a video communication system model in which the transmitter jointly optimizes source and channel coding parameters based on the inter-frame semantic importance of the input video and estimated channel state information. On this basis, we formulate an optimization problem that maximizes semantic importance weighted video reconstruction quality under a maximum bitrate constraint. To solve it, we first quantify inter-frame semantic importance using a cosine similarity-based metric with a shifted window mechanism. We then develop a multi-actor proximal policy optimization (MPPO) algorithm to solve the formulated problem by jointly adapting the source compression rate and channel coding rate. The learned policy can be directly applied to different video encoders without encoder-specific retraining or fine-tuning. SAJSCO achieves Bjøntegaard Delta rate reductions of 34.86% and 18.01% when integrated with H.265, a conventional video encoder, and DCVC-RT, a deep learning-based video encoder. Over-the-air experiments on a hardware testbed further demonstrate a PSNR gain of up to 1.448 dB with H.265 and an LPIPS reduction of up to 0.033 with DCVC-RT compared with the respective best-performing fixed-parameter baselines.
Figures & tables
Fig. 1: System model of semantic-aware joint source–channel optimization for video communication
Fig. 2: Illustration of the proposed inter-GOP semantic importance calculation method.
Fig. 3: Architecture of semantic-aware joint source–channel optimization for video communication
Parameter
Value
Parameter
Value
Number of GOPs, N
64
Shifted-window size, M
8
Training episodes
500
PPO update epochs, Ep
8
Discount factor, η
0.999
Clip parameter, ϵ
0.2
Optimizer
Adam
Penalty constant, ρ
50
Actor learning rate, δa
1×10−3
Critic learning rate, δc
3×10−3
Hyper-parameter, c1
10
Hyper-parameter, c2
0.3
TABLE I: Simulation Parameters
Fig. 4: Ablation study of the proposed MPPO algorithm.
Fig. 5: Performance comparison under different SNR conditions on the ActivityNet dataset with the H.265 encoder.
Fig. 6: PSNR and LPIPS performance versus bitrate on the ActivityNet dataset with the H.265 encoder under the RadioML channel setting.
Fig. 7: Performance comparison under the RadioML CSI setting on the HEVC testset.
Fig. 8: Visualization comparison of reconstructed video frames under different transmission schemes.
Method
MACs
Params
Speed
H.265
none
none
Enc. 42.41 fps, Dec. 25 fps
SAJSCO+H.265
318.97M
2.23M
Enc. 35.73 fps, Dec. 25 fps
DCVC-RT
385G
20.7M
Enc. 109 fps, Dec. 90 fps
SAJSCO+DCVC-RT
385.32G
22.93M
Enc. 66.73 fps, Dec. 90 fps
TABLE II: Comparison of Complexity
Fig. 9: Visualization of the state space and the corresponding action selection of the proposed MPPO algorithm.
Fig. 10: Hardware architecture of the proposed USRP-based communication testbed.
Parameter
Value
Parameter
Value
Carrier Frequency
5.5 GHz
LDPC Block Length
1800 bits
Sampling Rate
1 MSps
TX Gain (Seg. 1 & 4)
50 dB
Symbol Rate
62.5 kSps
TX Gain (Seg. 2 & 3)
60 dB
Bandwidth
125 kHz
RX Gain
56 dB
Modulation
BPSK
Antenna Model
VERT2450
Samples per Symbol
16
Antenna Placement
Vertical
TABLE III: Experimental Parameters
Method
PSNR
WPSNR
LPIPS
WLPIPS
Dec. Rate
H.265+SAJSCO
28.905
28.922
0.358
0.351
93.75%
H.265 ( r=1/3 )
25.358
25.578
0.518
0.508
84.38%
H.265 ( r=1/2 )
27.457
27.514
0.400
0.403
89.06%
H.265 ( r=2/3 )
27.262
27.425
0.527
0.544
92.19%
DCVC+SAJSCO
27.741
27.585
0.306
0.311
75.00%
DCVC ( r=1/3 )
26.455
26.084
0.339
0.351
70.31%
TABLE IV: Performance Comparison on the USRP Testbed
Semantic communication (SC) aims to reduce transmission overhead by conveying task-relevant information rather than raw data. However, existing SC approaches for video largely focus on pixel-level reconstruction or rely on complex spatiotemporal pipelines, leading to excessive bandwidth usage and latency that are unsuitable for low-resource deployments. In this paper, we propose ChronoSC, a task-oriented semantic communication framework for Video Question Answering (VideoQA). ChronoSC introduces Chrono-Color Stacking, a lightweight and lossless projection scheme that encodes temporal video dynamics into a single static image, enabling extreme temporal compression before transmission. This compact semantic representation is transmitted using a lightweight Deep Joint Source-Channel Coding (DeepJSCC) transceiver and explicitly reconstructed at the receiver. Unlike latent-space methods, explicit visual reconstruction enables the direct reuse of pre-trained vision-language models; specifically, a pre-trained BLIP model is employed to infer answers from noisy, reconstructed chrono-images. Experiments on the CLEVRER dataset show that ChronoSC achieves up to 192 times bandwidth reduction compared to raw video transmission while maintaining high VideoQA accuracy.
Phuc H. Nguyen, Trung T. Nguyen, Quy N. Duong +1
Smart Green Transformation Center (GREEN-X), VinUniversity, Vietnam. · School of Computer Science and Statistics, Trinity College Dublin, Ireland
Emerging physical AI systems require low-latency, task-oriented video communication over unreliable channels. We propose a semantic-aware multi-level neural video coding method for robust low-latency video transmission over unreliable channels that are abstracted as multi-level packet erasure channels. Built upon the real-time DCVC-RT neural video codec, the proposed framework introduces a semantic- and feature-aware coding strategy that partitions encoded representations into packets carrying different levels of semantic and latent-feature importance and assigns these packets to different streams, each associated with a priority level when transmitted over unreliable communication channels. We also developed an error-resilient entropy model that removes inter-packet dependencies, allowing each packet to be decoded independently under packet losses. The complete system is trained end-to-end over the abstracted multi-level packet erasure channels, enabling learning of channel-aware representations together with importance-aware packet assignment while facilitating the network for differentiated packet prioritization. Experiments show that the proposed framework significantly improves robustness over baseline DCVC-RT under packet erasures, achieving graceful degradation in less important regions while better preserving task-relevant visual content.
While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates remains challenging. Some NVCs incorporate perceptual or adversarial objectives but still suffer from artifacts due to limited generation capacity, whereas others leverage pretrained diffusion models to improve quality at the cost of heavy sampling complexity. To overcome these challenges, we propose S2VC, a Single-Step diffusion based Video Codec that integrates a conditional coding framework with an efficient single-step diffusion generator, enabling realistic reconstruction at low bitrates with reduced sampling cost. Recognizing the importance of semantic conditioning in single-step diffusion, we introduce Contextual Semantic Guidance to extract frame-adaptive semantics from buffered features. It replaces text captions with efficient, fine-grained conditioning, thereby improving generation realism. In addition, Temporal Consistency Guidance is incorporated into the diffusion U-Net to enforce temporal coherence across frames and ensure stable generation. Extensive experiments show that S2VC delivers state-of-the-art perceptual quality with an average 52.73% bitrate saving over prior perceptual methods, underscoring the promise of single-step diffusion for efficient, high-quality video compression. Project: https://onedc-codec.github.io/s2vc/
Naifu Xue, Zhaoyang Jia, Jiahao Li +4
Communication University of China · University of Science and Technology of China · Microsoft Research Asia