Organizations: School of Computer Science and Technology, Xi’an Jiaotong University · School of Computer Science and Information Engineering, Hefei University of Technology
Hyperspectral image super-resolution (HSI-SR) aims to recover spatial detail while preserving the spectral shape on which quantitative analysis relies. Recent HSI-SR methods, from repurposed RGB super-resolution backbones to dedicated spectral-spatial architectures, have greatly improved spatial reconstruction. However, overlooking the compact spectral structure of hyperspectral data leaves residual spectral errors, while binding the spectral treatment to each architecture forces it to be rebuilt for every new backbone. Yet the low-dimensional structure of spectra belongs to the data, not to any backbone. Backbones differ in the errors they leave, but not in the structure of the true spectra. One rectifier design can therefore serve any backbone. Building on this insight, we propose the \textbf{S}pectral \textbf{R}ectification \textbf{S}uper-\textbf{R}esolution Network (\textbf{SR2-Net}), a model-agnostic rectifier that needs nothing from the backbone but its output, leaves its internal architecture untouched, and is trained per backbone. SR2-Net follows an \emph{enhance-then-rectify} pipeline in which Hierarchical Spectral-Spatial Synergy Attention (\textbf{H-S3A}) reinforces cross-band interactions, while Mode-Constrained Rectification (\textbf{MCR}) confines the correction to a learned compact spectral subspace. A degradation-consistency constraint further ties the output to the observed low-resolution input. Experiments with five backbones spanning CNN, Transformer, and diffusion families show that one fixed configuration improves spectral fidelity in every reported setting while preserving or improving spatial quality. Averaged over thirty in-domain settings, SR2-Net removes 22.6% of the residual spectral error at a backbone-independent cost of 0.048M parameters.
Figures & tables
Figure 1: Motivation. (a) Strong SR models, general-purpose and HSI-specific alike, introduce spectral distortion during reconstruction; (b) the model-agnostic SR 2 -Net rectifier restores spectral consistency.
Figure 2: Overview of SR 2 -Net . Given the LR input ILR , an SR backbone produces an initial reconstruction I~SR , which SR 2 -Net rectifies into the output I^SR . (a) H-S 3 A enhances cross-band interaction via spectral grouping and adjacent shuffle; (b) TSA fuses spectral, height, and width views to recalibrate features; (c) MCR, drawn with N stages ( N=1 in our configuration), projects features onto a compact spectral subspace and aggregates the back-projected corrections.
Figure 3: Error maps at ×4 (in-domain, CAVE and ARAD-1K). Pseudo-RGB uses bands 25 / 15 / 5 as R–G–B, and band-averaged absolute error maps (darker is better) pair each +Ours result with its Base. Red boxes mark where SR 2 -Net sharpens edges and fine structures; the cross-domain ICVL comparison and more galleries are in Appendix I .
Figure 4: Per-pixel spectra at ×4 on CAVE across CNN-, Transformer-, and diffusion-based backbones. For two representative pixels (marked in the pseudo-RGB images), the baseline spectra oscillate away from the ground truth while SR 2 -Net follows it; additional scenes and datasets are in Appendix J .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Framework
PyTorch
Hardware
NVIDIA RTX 3090 GPU
Optimizer
AdamW ( β1=0.9 , β2=0.99 )
Learning rate
1×10−4→1×10−5 (cosine)
Epochs
1000
Batch size
8
Appendix
Table A1: Training hyperparameters, shared across all backbones.
Component
Configuration
Spectral groups
4 contiguous groups of ⌈S/4⌉ bands (bands zero-padded to a multiple of four; padding removed after back-projection)
H-S 3 A blocks
4 , applied in sequence with independent parameters and a residual connection around each: F(l)=F(l−1)+Blockl(F(l−1)) , where F(0) is the zero-padded input and F(4)=Fs enters MCR
▹ per-group stems
identity / 3×3 / 5×5 depthwise / 3×3 dilated (one per group); 1×1 mixing after the shuffle
▹ TSA views
height, width, spectral ( h,w,s )
MCR blocks (refinement stages N )
1 ; ablations that drop MCR remove the module entirely (projection and back-projection included), so the rectified output is ΠS(Fs) , giving the H-S 3 A-only variant of the component ablation
▹ projection P(⋅)
1×1 conv, S′→r , GeLU
Appendix
Table A2: Configuration of the SR 2 -Net rectifier, where S is the number of spectral bands, S′=4⌈S/4⌉ the zero-padded band count, and r the MCR bottleneck rank.
Tag
Type
Setting
K1
Kernel
area-based downsampling
K2
Kernel
alternative bicubic implementation
B1
Blur
Gaussian σ=0.6 , kernel 5×5
B2
Blur
Gaussian σ=1.0 , kernel 7×7
N1
Noise
Gaussian σ=1/255
N2
Noise
Gaussian σ=2/255
Appendix
Table A3: Test-time degradation settings. Each is applied only at test time to a model trained under standard bicubic downsampling.
Method
Scale
mPSNR ↑
mSAM ↓
mSSIM ↑
CC ↑
Base (SwinIR)
×2
50.4662
0.7727
0.9972
0.9990
×4
39.5717
1.3950
0.9709
0.9902
×8
33.4185
2.4682
0.8869
0.9551
Base + SG
×2
46.1893
1.9059
0.9959
0.9975
×4
38.7834
2.4512
0.9698
0.9889
×8
33.2419
3.2638
0.8859
0.9541
Appendix
Table A6: Spectral-rectifier comparison on ARAD-1K with SwinIR. SG, PCA, and IBP are fixed post-hoc methods; MCR and SR 2 -Net are jointly trained. Best in bold , second best underlined .
Variant (SwinIR, ×4 )
mPSNR ↑
mSAM ↓
mSSIM ↑
Base
39.58±.03
1.397±.011
0.9710±.0005
+ Ldeg only
39.70±.03
1.371±.010
0.9710±.0007
+ Generic head
39.99±.04
1.362±.012
0.9713±.0006
+ MCR only
40.25±.03
1.319±.010
0.9716±.0006
+ SR 2 -Net
40.96±.03
1.280±.009
0.9753±.0006
Appendix
Table A7: Seed stability on ARAD-1K ( ×4 , SwinIR): mean ± std over three seeds {42,123,2024} . The architecture gap (cheap controls versus MCR/ SR 2 -Net ) is 0.043 to 0.091 mSAM, at least 3.5× the across-seed std.
Scale
Metric
Base
+Ours
Δ
×4
Params (M)
8.228
8.276
+0.048
FLOPs (G)
19.302
20.785
+1.482
Appendix
Table A8: Complexity of SR 2 -Net using SwinIR on ARAD-1K at ×4 . Parameters and FLOPs are computed for the 322 LR training patch ( 1282 HR); the rectifier adds 0.048 M parameters and 1.482 G FLOPs.
Residual statistic
Base
+ SR 2 -Net
Improved
p
Spectral MAE
0.00783
0.00534
50/50
<10−6
Gradient ∣Δe∣
0.000838
0.000781
44/50
<10−3
Curvature ∣Δ2e∣
0.000697
0.000618
42/50
<10−3
Appendix
Table A10: Dataset-level residual spectral-error analysis over 50 ARAD-1K test scenes (SwinIR, ×4 ). “Improved” counts scenes on which SR 2 -Net lowers the statistic; p is from a one-sided Wilcoxon signed-rank test. Per-scene mSAM improves on all 50 scenes ( p<10−6 ).
Sanya Oceanographic Institution, Ocean University of China, Sanya, China · State Key Laboratory of Physical Oceanography, Ocean University of China, Qingdao 266100, China · Department of Electrical and Computer Engineering, Mississippi State University, Starkville, MS 39762 USA