Diffusion-aided Task-oriented Semantic Communications with Model Inversion Attack
Organizations: School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, Guangdong 518172, China
Abstract
Semantic communication enhances transmission efficiency by conveying semantic information rather than raw input symbol sequences. Task-oriented semantic communication further aims to retain only task-specific information, thereby achieving greater bandwidth savings. However, these neural-network-based communication systems are vulnerable to model inversion attacks, in which adversaries attempt to recover sensitive input information from intercepted semantic features. The key challenge is therefore to preserve privacy while maintaining task accuracy and robustness. We consider a task-confidential setting in which the adversary attempts to reconstruct the original input from intercepted features without knowing the legitimate receiver's task or model. Although PSNR and SSIM are commonly used to assess reconstruction quality, we find that an external classifier can still perform the legitimate receiver's task with nontrivial accuracy on reconstructions with low PSNR or SSIM, indicating that these reconstructions still contain task-level semantic leakage. We therefore propose DiffSem, which splits the diffusion process between controlled transmitter-side self-noising and matched receiver-side reverse denoising. Experiments on the MNIST, CIFAR-10, and CelebA datasets show that DiffSem improves the legitimate receiver's task accuracy without increasing either the transmitted feature size or information leakage.
Figures & tables
| Dataset | Image Latent | Encoder | Receiver Classifier | Third-party Classifier | Adversary |
| MNIST | Conv2d , stride 2 + LReLU Conv2d , stride 2 + LReLU Conv2d + LayerNorm | Concat Linear + ReLU Linear ( classes) | Linear + Tanh Linear + Tanh Linear + Sigmoid | Flatten Linear + Tanh Linear Reshape to | |
| CIFAR-10 | Conv2d + LReLU Conv2d , stride 2 + LReLU Residual Block Conv2d , stride 2 + LReLU + LayerNorm | Concat Conv + Tanh Strided Conv + Tanh Residual Block Global Pool + Linear ( classes) | Strided Conv + ReLU Strided Conv + ReLU Residual Block Global Pool Linear + ReLU Linear ( classes) | Conv + Tanh Deconv + Tanh Residual Block Deconv + Tanh Residual Block Conv + Tanh | |
| CelebA | Conv2d + BN + ReLU Conv2d + BN + ReLU Conv2d , stride 2 + BN + ReLU Conv2d , stride 2 + BN + ReLU Conv2d + BN + ReLU Conv2d + LayerNorm | MobileNetV2 Input latent Linear head ( logits) | ResNet-18 Input image Linear head ( logits) | Conv2d + BN + ReLU Residual Block Upsample + Conv2d + BN + ReLU Residual Block Upsample + Conv2d + BN + ReLU Conv2d + ReLU Conv2d + Tanh |
| Component | MNIST | CIFAR-10 | CelebA |
| Latent Domain | |||
| DiffSem U-Net | Time EmbedFC ( ) ResidualConvBlock UnetDown UnetDown AvgPool2d + GELU ConvTranspose2d + GN + ReLU UnetUp UnetUp Conv2d + GroupNorm + ReLU + Conv2d | Time EmbedFC ( ) ResidualConvBlock UnetDown UnetDown AvgPool2d + GELU ConvTranspose2d + GN + ReLU UnetUp UnetUp Conv2d + GroupNorm + ReLU + Conv2d | Time embedding (sinusoidal + MLP) Conv2d ResBlock Downsample + ResBlock Downsample + ResBlock Mid ResBlock + Self-Attention + ResBlock Upsample + ResBlock Upsample + ResBlock GroupNorm + SiLU + Conv2d |
| E-DiffSem U-Net | Same U-Net backbone as DiffSem Time EmbedFC ( ) Context EmbedFC ( ) | Same U-Net backbone as DiffSem Time EmbedFC ( ) Context EmbedFC ( ) | Same U-Net backbone as DiffSem Time embedding (sinusoidal + MLP) Attribute embedding MLP ( ) Time and condition fused in every ResBlock |