Organizations: School of Computer Science and Technology, Xi’an Jiaotong University · School of Computer Science and Information Engineering, Hefei University of Technology
Hyperspectral image super-resolution (HSI-SR) aims to recover spatial detail while preserving the spectral shape on which quantitative analysis relies. Recent HSI-SR methods, from repurposed RGB super-resolution backbones to dedicated spectral-spatial architectures, have greatly improved spatial reconstruction. However, overlooking the compact spectral structure of hyperspectral data leaves residual spectral errors, while binding the spectral treatment to each architecture forces it to be rebuilt for every new backbone. Yet the low-dimensional structure of spectra belongs to the data, not to any backbone. Backbones differ in the errors they leave, but not in the structure of the true spectra. One rectifier design can therefore serve any backbone. Building on this insight, we propose the \textbf{S}pectral \textbf{R}ectification \textbf{S}uper-\textbf{R}esolution Network (\textbf{SR2-Net}), a model-agnostic rectifier that needs nothing from the backbone but its output, leaves its internal architecture untouched, and is trained per backbone. SR2-Net follows an \emph{enhance-then-rectify} pipeline in which Hierarchical Spectral-Spatial Synergy Attention (\textbf{H-S3A}) reinforces cross-band interactions, while Mode-Constrained Rectification (\textbf{MCR}) confines the correction to a learned compact spectral subspace. A degradation-consistency constraint further ties the output to the observed low-resolution input. Experiments with five backbones spanning CNN, Transformer, and diffusion families show that one fixed configuration improves spectral fidelity in every reported setting while preserving or improving spatial quality. Averaged over thirty in-domain settings, SR2-Net removes 22.6% of the residual spectral error at a backbone-independent cost of 0.048M parameters.
Figures & tables
Figure 1: Motivation. (a) Strong SR models, general-purpose and HSI-specific alike, introduce spectral distortion during reconstruction; (b) the model-agnostic SR 2 -Net rectifier restores spectral consistency.
Figure 2: Overview of SR 2 -Net . Given the LR input ILR , an SR backbone produces an initial reconstruction I~SR , which SR 2 -Net rectifies into the output I^SR . (a) H-S 3 A enhances cross-band interaction via spectral grouping and adjacent shuffle; (b) TSA fuses spectral, height, and width views to recalibrate features; (c) MCR, drawn with N stages ( N=1 in our configuration), projects features onto a compact spectral subspace and aggregates the back-projected corrections.
Figure 3: Error maps at ×4 (in-domain, CAVE and ARAD-1K). Pseudo-RGB uses bands 25 / 15 / 5 as R–G–B, and band-averaged absolute error maps (darker is better) pair each +Ours result with its Base. Red boxes mark where SR 2 -Net sharpens edges and fine structures; the cross-domain ICVL comparison and more galleries are in Appendix I .
Figure 4: Per-pixel spectra at ×4 on CAVE across CNN-, Transformer-, and diffusion-based backbones. For two representative pixels (marked in the pseudo-RGB images), the baseline spectra oscillate away from the ground truth while SR 2 -Net follows it; additional scenes and datasets are in Appendix J .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Framework
PyTorch
Hardware
NVIDIA RTX 3090 GPU
Optimizer
AdamW ( β1=0.9 , β2=0.99 )
Learning rate
1×10−4→1×10−5 (cosine)
Epochs
1000
Batch size
8
Appendix
Table A1: Training hyperparameters, shared across all backbones.
Component
Configuration
Spectral groups
4 contiguous groups of ⌈S/4⌉ bands (bands zero-padded to a multiple of four; padding removed after back-projection)
H-S 3 A blocks
4 , applied in sequence with independent parameters and a residual connection around each: F(l)=F(l−1)+Blockl(F(l−1)) , where F(0) is the zero-padded input and F(4)=Fs enters MCR
▹ per-group stems
identity / 3×3 / 5×5 depthwise / 3×3 dilated (one per group); 1×1 mixing after the shuffle
▹ TSA views
height, width, spectral ( h,w,s )
MCR blocks (refinement stages N )
1 ; ablations that drop MCR remove the module entirely (projection and back-projection included), so the rectified output is ΠS(Fs) , giving the H-S 3 A-only variant of the component ablation
▹ projection P(⋅)
1×1 conv, S′→r , GeLU
Appendix
Table A2: Configuration of the SR 2 -Net rectifier, where S is the number of spectral bands, S′=4⌈S/4⌉ the zero-padded band count, and r the MCR bottleneck rank.
Tag
Type
Setting
K1
Kernel
area-based downsampling
K2
Kernel
alternative bicubic implementation
B1
Blur
Gaussian σ=0.6 , kernel 5×5
B2
Blur
Gaussian σ=1.0 , kernel 7×7
N1
Noise
Gaussian σ=1/255
N2
Noise
Gaussian σ=2/255
Appendix
Table A3: Test-time degradation settings. Each is applied only at test time to a model trained under standard bicubic downsampling.
Method
Scale
mPSNR ↑
mSAM ↓
mSSIM ↑
CC ↑
Base (SwinIR)
×2
50.4662
0.7727
0.9972
0.9990
×4
39.5717
1.3950
0.9709
0.9902
×8
33.4185
2.4682
0.8869
0.9551
Base + SG
×2
46.1893
1.9059
0.9959
0.9975
×4
38.7834
2.4512
0.9698
0.9889
×8
33.2419
3.2638
0.8859
0.9541
Appendix
Table A6: Spectral-rectifier comparison on ARAD-1K with SwinIR. SG, PCA, and IBP are fixed post-hoc methods; MCR and SR 2 -Net are jointly trained. Best in bold , second best underlined .
Variant (SwinIR, ×4 )
mPSNR ↑
mSAM ↓
mSSIM ↑
Base
39.58±.03
1.397±.011
0.9710±.0005
+ Ldeg only
39.70±.03
1.371±.010
0.9710±.0007
+ Generic head
39.99±.04
1.362±.012
0.9713±.0006
+ MCR only
40.25±.03
1.319±.010
0.9716±.0006
+ SR 2 -Net
40.96±.03
1.280±.009
0.9753±.0006
Appendix
Table A7: Seed stability on ARAD-1K ( ×4 , SwinIR): mean ± std over three seeds {42,123,2024} . The architecture gap (cheap controls versus MCR/ SR 2 -Net ) is 0.043 to 0.091 mSAM, at least 3.5× the across-seed std.
Scale
Metric
Base
+Ours
Δ
×4
Params (M)
8.228
8.276
+0.048
FLOPs (G)
19.302
20.785
+1.482
Appendix
Table A8: Complexity of SR 2 -Net using SwinIR on ARAD-1K at ×4 . Parameters and FLOPs are computed for the 322 LR training patch ( 1282 HR); the rectifier adds 0.048 M parameters and 1.482 G FLOPs.
Residual statistic
Base
+ SR 2 -Net
Improved
p
Spectral MAE
0.00783
0.00534
50/50
<10−6
Gradient ∣Δe∣
0.000838
0.000781
44/50
<10−3
Curvature ∣Δ2e∣
0.000697
0.000618
42/50
<10−3
Appendix
Table A10: Dataset-level residual spectral-error analysis over 50 ARAD-1K test scenes (SwinIR, ×4 ). “Improved” counts scenes on which SR 2 -Net lowers the statistic; p is from a one-sided Wilcoxon signed-rank test. Per-scene mSAM improves on all 50 scenes ( p<10−6 ).
Hyperspectral image super-resolution is essential for enhancing the spatial fidelity of HSI data, yet existing deep learning methods often struggle with substantial spectral redundancy and the limited non-linear modeling capacity of standard feed-forward networks (FFNs). To address these challenges, we propose Spectral Dynamic Attention Network (SDANet), a framework designed to adaptively suppress redundant spectral interactions. SDANet integrates two key components: 1) Dynamic Channel Sparse Attention (DCSA) module that computes channel-wise correlations and selectively preserves the most informative attention responses through dynamic and data-dependent sparsification. 2) Frequency-Enhanced Feed-Forward Network (FE-FFN) that jointly models spatial and frequency-domain representations to enhance non-linear expressiveness. Extensive experiments on two benchmark datasets demonstrate that SDANet achieves state-of-the-art HISR performance while maintaining competitive efficiency. The code will be made publicly available at https://github.com/oucailab/SDANet.
Tengya Zhang, Feng Gao, Lin Qi +2
Sanya Oceanographic Institution, Ocean University of China, Sanya, China · State Key Laboratory of Physical Oceanography, Ocean University of China, Qingdao 266100, China · Department of Electrical and Computer Engineering, Mississippi State University, Starkville, MS 39762 USA
Hyperspectral super-resolution (HSR) reconstructs a high-spatial-resolution hyperspectral image by fusing a low-resolution hyperspectral image (LR-HSI) with a high-resolution multispectral image (HR-MSI). In the absence of real-world paired data, HSR methods are evaluated almost exclusively on synthetic experiments derived from hyperspectral datasets through Wald's protocol. Despite the protocol's widespread adoption, its practical implementation varies markedly across research works, typically relying on a single (usually Gaussian) or very few point spread functions (PSFs), one or two spectral response functions (SRFs), and a couple of spatial downsampling factors. As a result, reported performance figures are difficult to compare across the literature, in addition to being often difficult to reproduce; furthermore, they may not generalize across realistic sensing conditions. We introduce HyperBench, a unified and extensible framework that standardizes synthetic experimentation for HSR. HyperBench supports diverse degradation configurations spanning ten PSFs, four SRFs derived from operational multispectral sensors, configurable spatial downsampling factors, and matched additive white Gaussian noise; its goal is to automate large-scale evaluation and structured logging. By decoupling model development from experimental design, the framework enables reproducible, apples-to-apples cross-method comparison with minimal friction. We use HyperBench to evaluate six recently proposed HSR methods across a 70-configuration sweep on four widely used hyperspectral scenes and observe that the inter-method PSNR spread widens from approximately 5 dB on the easiest PSF to over 13 dB on the hardest - a fragility that is structurally invisible to the prevailing single-configuration evaluation protocol. HyperBench code is available at https://github.com/ritikgshah/HyperBench .
Achieving cross-sensor generalization and arbitrary-scale reconstruction with a single model remains challenging in hyperspectral super-resolution (HSR). Although recent methods support arbitrary-scale reconstruction, applying them to new sensors or scales beyond the training range often requires additional data and computation to maintain reconstruction quality. To address these challenges, we propose OmniHSR, which predicts band-shared spatial operators rather than spectral values. Cross-Spectral Mapping (CSM) resamples inputs with any number of bands to fixed reference positions and predicts local operators with Gaussian supports. Continuous Operator-Field Reconstruction (COFR) composes these operators into a continuous field and applies them to all original bands for arbitrary-scale reconstruction. Experiments demonstrate that operator prediction outperforms direct spectral-value prediction on all seven datasets. Trained solely on ARAD with only 0.538M parameters, OmniHSR outperforms all directly transferred baselines on six unseen datasets without target-domain training data or adaptation. Across twelve upsampling factors from ×2 to ×48, it improves average PSNR on Pavia U and Chikusei by 0.55 dB over the strongest baseline. It also surpasses baselines trained from scratch or adapted on the target sensor and achieves up to 36× faster inference. Our code will be publicly released soon.
Ji-Xuan He, Guohang Zhuang, Bo Junge +6
Xi’an Jiaotong University · Hefei University of Technology · National University of Singapore +1