Detection of Adversarial Attacks on Super-Resolvers Using Spectral Features
Authors: Emma J. Reid, Haley Duba-Sullivan, Tony G. Allen
Organizations: Human Analysis and Biometrics Oak Ridge Laboratory Oak Ridge, TN, USA · Radar and Computational Imaging Oak Ridge Laboratory Oak Ridge, TN, USA
The integration of deep learning models into image preprocessing pipelines such as super-resolution introduces a largely unexplored attack vector for adversaries targeting downstream tasks. To ensure trustworthiness of critical imaging pipelines, we must be able to detect adversarial behavior within preprocessing models. In this paper, we propose a spectral-based detection method for identifying adversarial attacks embedded in super-resolution model weights. More specifically, we use the radially-averaged power spectral density as a discriminative feature to train an extreme gradient boosting (XGBoost) detector, demonstrating detectability of model-level threats in super-resolution networks. We further benchmark our detector against magnitude- and phase-based Fourier spectrum detectors, evaluating each method across a range of training and cross-architecture scenarios. Our proposed detector out-performs the comparison detectors in most of these scenarios and indicates that high-frequency features are most informative for detecting AdvSR attacks across SR architectures.
Figures & tables
Fig. 1: Example of AdvSR attacks for SRCNN, EDSR, and SwinIR architectures targeting misclassification of “war planes” to “trailer trucks” for YOLOv11. These attacks preserve image quality while causing misclassification with high confidence.
Fig. 2: Pipeline for training and testing detection methods. We apply clean and adversarial super-resolvers to LR images, apply feature extractors to the resulting SR images, and input these features to detectors for both training and testing.
Ours
InputMFS
InputPFS
SR model
Acc.
F1
AUC
Acc.
F1
AUC
Acc.
F1
AUC
SRCNN
0.845
0.840
0.926
0.890
0.888
0.954
0.481
0.445
0.498
EDSR
0.970
0.969
0.997
0.886
0.887
0.952
0.485
0.456
0.488
SwinIR
0.962
0.962
0.997
0.917
0.919
0.978
0.523
0.526
0.524
TABLE I: Detector performance across SR architectures. Each detector is trained and tested using outputs from the same SR architecture. The best result for each architecture and metric is shown in bold.
Ours
InputMFS
InputPFS
Training
Test SR
Acc.
F1
AUC
Acc.
F1
AUC
Acc.
F1
AUC
Leave-one-out
SRCNN
0.500
0.000
0.564
0.576
0.446
0.649
0.462
0.462
0.472
Leave-one-out
EDSR
0.966
0.967
0.999
0.723
0.700
0.788
0.470
0.496
0.481
Leave-one-out
SwinIR
0.864
0.846
0.936
0.769
0.775
0.844
0.504
0.513
0.525
Unified
SRCNN
0.856
0.854
0.934
0.848
0.856
0.921
0.511
0.524
0.491
Unified
EDSR
0.985
0.985
1.000
0.818
0.824
0.886
0.492
0.538
0.487
TABLE II: Detector performance under leave-one-SR-model-out and unified training. The unified detectors are trained using all three SR architectures; “All SR” is their pooled test set. The best result for each test set and metric is shown in bold.
Fig. 4: Shapley analysis of the proposed radially-averaged PSD detector across SR architectures. Feature color represents radial spectral power, while the horizontal Shapley value represents the contribution toward a clean or adversarial prediction.
Single image super-resolution aims to reconstruct high-resolution images from low-resolution inputs. This paper proposes FreeTransformSR, a novel lightweight super-resolution network based on a channel-wise free low-rank learnable transform. The transform learns task-adaptive basis functions in a data-driven manner, enabling adaptive feature modulation with minimal parameter overhead. To further enhance high-frequency detail recovery, we introduce a local feature modulation branch that complements transform-domain processing with depthwise convolution. In addition, a soft complexity adaptive module dynamically fuses the outputs of local convolution and window self-attention branches through a lightweight gating network, adaptively adjusting the fusion ratio based on regional texture characteristics. An adaptive intensity modulation strategy is also incorporated to adjust transform-domain response strength at the sample level, enabling the network to dynamically adjust processing intensity according to input features. Extensive experiments on five benchmark datasets demonstrate that FreeTransformSR achieves competitive PSNR/SSIM performance with significantly fewer parameters and FLOPs. Specifically, FreeTransformSR achieves 32.41 dB on BSD100 x2 and 27.00 dB on Urban100 x4 with only 595K parameters, while delivering faster inference speed than competing methods, making it well-suited for deployment in resource-constrained scenarios. Source code is available at: https://github.com/HJiLi/FreeTransformSR.
Hongji Li, Yunhui Li
Chinese Academy of Sciences Changchun Institute of Optics Fine Mechanics and Physics, Changchun, Jilin 130033, China · University of the Chinese Academy of Sciences, Beijing 100049, China
Super-Resolution (SR) is a time-hallowed image processing problem that aims to improve the quality of a Low-Resolution (LR) sample up to the standard of its High-Resolution (HR) counterpart. We aim to address this by introducing Super-Resolution Generator (SuRGe), a fully-convolutional Generative Adversarial Network (GAN)-based architecture for SR. We show that distinct convolutional features obtained at increasing depths of a GAN generator can be optimally combined by a set of learnable convex weights to improve the quality of generated SR samples. In the process, we employ the Jensen-Shannon and the Gromov-Wasserstein losses respectively between the SR-HR and LR-SR pairs of distributions to further aid the generator of SuRGe to better exploit the available information in an attempt to improve SR. Moreover, we train the discriminator of SuRGe with the Wasserstein loss with gradient penalty, to primarily prevent mode collapse. The proposed SuRGe, as an end-to-end GAN workflow tailor-made for super-resolution, offers improved performance while maintaining low inference time. The efficacy of SuRGe is substantiated by its superior performance compared to 28 state-of-the-art contenders on 10 benchmark datasets.
Electronics and Communication Sciences Unit (ECSU), Dolby Laboratories, India · Statistics and Mathematics Unit (SMU), Indian Statistical Institute, Kolkata, India
Diffusion prior-based methods have shown impressive results in real-world image super-resolution (ISR), yet two key challenges persist: balancing pixel-level fidelity with semantic quality, and adapting to diverse degradations. Existing dual-branch approaches freeze the pixel module during semantic training, but the semantic branch can still expand capacity within the pixel subspace, precluding genuine perceptual improvement. Moreover, using a single static adapter cannot generalize across heterogeneous real-world corruptions. To address both issues, we propose FreqOrtho-SR, which comprises: Frequency-guided Mixture of LoRA Experts (FreqMoE), it routes inputs to specialized experts via a non-parametric FFT-based degradation-feature extractor that encodes frequency-domain signatures, enabling stable and interpretable specialization across corruption types; and Orthogonal Gradient Projection (OGP), which reframes the dual-objective optimization as a subspace-constrained problem: by extracting the pixel-fidelity subspace via SVD on combined expert weight deltas and projecting semantic gradients onto its null space, OGP guarantees orthogonality between the two objectives, enabling genuinely complementary learning without mutual interference. Experiments show that FreqOrtho-SR achieves competitive overall performance and a strong fidelity-perception trade-off across multiple benchmarks with efficient single-step inference. The source code of our method can be found at \href.
Minh Son Hoang, Dinh Phu Tran, Quyen Nguyen Duc +2