Achieving cross-sensor generalization and arbitrary-scale reconstruction with a single model remains challenging in hyperspectral super-resolution (HSR). Although recent methods support arbitrary-scale reconstruction, applying them to new sensors or scales beyond the training range often requires additional data and computation to maintain reconstruction quality. To address these challenges, we propose OmniHSR, which predicts band-shared spatial operators rather than spectral values. Cross-Spectral Mapping (CSM) resamples inputs with any number of bands to fixed reference positions and predicts local operators with Gaussian supports. Continuous Operator-Field Reconstruction (COFR) composes these operators into a continuous field and applies them to all original bands for arbitrary-scale reconstruction. Experiments demonstrate that operator prediction outperforms direct spectral-value prediction on all seven datasets. Trained solely on ARAD with only 0.538M parameters, OmniHSR outperforms all directly transferred baselines on six unseen datasets without target-domain training data or adaptation. Across twelve upsampling factors from ×2 to ×48, it improves average PSNR on Pavia U and Chikusei by 0.55 dB over the strongest baseline. It also surpasses baselines trained from scratch or adapted on the target sensor and achieves up to 36× faster inference. Our code will be publicly released soon.
Figures & tables
Figure 1: Cross-sensor, arbitrary-scale reconstruction with OmniHSR. (a) Existing methods, here CiaoSR, decode the spectral values of every band from features and repeat querying and decoding per requested scale. (b) OmniHSR predicts local spatial operators with CSM and applies them to all observed bands with COFR. (c) A new sensor requires adapting or retraining existing methods, and a new scale requires retraining fixed-scale methods. (d) At ×16 , OmniHSR reaches a higher zero-shot PSNR than the baselines on Pavia U and Chikusei and runs 36.5× faster than CiaoSR on the 145 -band Botswana. N is the number of requested scales; times are measured on Botswana.
Figure 2: Overview of OmniHSR. (a) CSM resamples B input bands to M fixed normalized band positions and predicts scale-conditioned spatial supports and zero-sum operators. (b) COFR composes the operators with normalized Gaussian weights into one operator at target positions spaced 1/s apart and applies it to all raw B bands.
Figure 3: Visual comparison at × 4 on ARAD (in-domain) and on CAVE, Chikusei and Pavia U (unseen sensors, zero-shot). For each dataset, the top row shows a false-color zoom, Chikusei in the near-infrared composite of bands 70, 100 and 36 and Pavia U in that of about 810, 650 and 550 nm, and the bottom row the absolute error averaged over bands; the first column gives the full 256×256 crop with the zoom window. Numbers are PSNR / SAM on the crop. The right column plots the spectrum at the circled pixel with its RMSE in the legend.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Zero-shot OmniHSR against the best of the six baselines under each way of reusing them on a new sensor, mean PSNR (dB) over seven factors from × 3.5 to × 16 on test tiles of Pavia U and Chikusei. Orange: lead of OmniHSR over the best of the three ways.
Figure 5: Zero-shot reconstructions at × 4 on the other three unseen datasets, laid out as Figure 3 without the spectra. Pavia C and Botswana use a near-infrared composite of about 810, 650 and 550 nm. Numbers are PSNR / SAM on the crop.
Figure 6: RMSE of every band over the crop at × 4 on the seven datasets, with the crops chosen as in Appendix B.2 .
Figure 7: Zero-shot reconstructions at × 24 on Pavia U and Chikusei, laid out as Figure 3 . The input has 11×11 pixels, so the zoom shows the whole crop. Numbers are PSNR / SAM on the crop, and the spectrum is taken at the edge pixel with the median bicubic RMSE.
Figure 8: Output paths at × 4 with the models of Table 12 . Spectral values: the network outputs the value of every band. The right column gives the RMSE of every band over the crop on a log scale. Numbers are PSNR / SAM on the crop.
Figure 9: What the network predicts at × 4. Left: the supports in the window, colored by the direction of the long axis (grey: axis ratio below 1.2). Middle: the norm of every 5×5 operator over the crop. Right: the three edge operators in the window with the largest norm.
Figure 10: Latency against magnification factor on one RTX 4090D (24 GB). The same 32×32 low-resolution input is magnified by ×2 to ×48 (outputs of 64 to 1536 pixels), left on the 31 -band ARAD and right on the 128 -band Chikusei, log-log axes. The x-axis is the magnification factor of a single output and every point is one independent inference, not the accumulated time of rendering many scales from one encoding. A baseline stops where all three inputs run out of memory (OOM). The annotation gives the latency of OmniHSR at ×48 and its ratio to the fastest and the slowest baseline that reach this factor.
Spectral super-resolution of multispectral satellite images can enable high temporal- and spatial-resolution hyperspectral satellite imagery at a modest cost, significantly increasing the applicability of hyperspectral remote sensing. This task is inherently ill-posed, making it well-suited for deep learning-based methods. In this study, the spectral super-resolution task is framed as an operator learning problem, and SSRON is proposed as a Deep Operator Network that effectively learns function-to-function mappings from downsampled spectra to continuous spectra. The model is trained to super-resolve Sentinel-2A-like multispectral imagery to EMIT images. Compared to baseline models, SSRON achieves superior performance across all metrics. The model also demonstrates zero-shot spectral super-resolution capability by predicting bands unseen during training. Furthermore, its continuous-output formulation suggests the potential to estimate spectra at finer wavelength intervals than the native sensor. These results suggest the potential of SSRON and establishes operator learning as a promising direction for spectral super-resolution.
Seokhyun Chin
California Institute of Technology 1200 E California Boulevard, Pasadena, CA, USA
Hyperspectral super-resolution (HSR) reconstructs a high-spatial-resolution hyperspectral image by fusing a low-resolution hyperspectral image (LR-HSI) with a high-resolution multispectral image (HR-MSI). In the absence of real-world paired data, HSR methods are evaluated almost exclusively on synthetic experiments derived from hyperspectral datasets through Wald's protocol. Despite the protocol's widespread adoption, its practical implementation varies markedly across research works, typically relying on a single (usually Gaussian) or very few point spread functions (PSFs), one or two spectral response functions (SRFs), and a couple of spatial downsampling factors. As a result, reported performance figures are difficult to compare across the literature, in addition to being often difficult to reproduce; furthermore, they may not generalize across realistic sensing conditions. We introduce HyperBench, a unified and extensible framework that standardizes synthetic experimentation for HSR. HyperBench supports diverse degradation configurations spanning ten PSFs, four SRFs derived from operational multispectral sensors, configurable spatial downsampling factors, and matched additive white Gaussian noise; its goal is to automate large-scale evaluation and structured logging. By decoupling model development from experimental design, the framework enables reproducible, apples-to-apples cross-method comparison with minimal friction. We use HyperBench to evaluate six recently proposed HSR methods across a 70-configuration sweep on four widely used hyperspectral scenes and observe that the inter-method PSNR spread widens from approximately 5 dB on the easiest PSF to over 13 dB on the hardest - a fragility that is structurally invisible to the prevailing single-configuration evaluation protocol. HyperBench code is available at https://github.com/ritikgshah/HyperBench .
Hyperspectral imaging provides rich spectral information for quantitative remote sensing, yet hyperspectral sensors remain costly and thus unavailable in many UAV deployments. Spectral super-resolution (SSR) seeks to reconstruct hyperspectral images (HSIs) from multispectral images (MSIs). Most existing SSR methods assume a fixed and known spectral response function (SRF) and are therefore limited to single-sensor settings. In practical cross-sensor scenarios, the spectral degradation from HSI to MSI is unknown and varies with sensor characteristics and scene content, which renders HSI reconstruction ill-posed. This paper proposes a physics-guided deep unfolding network, termed PGU-Net, to address blind cross-sensor SSR by jointly estimating the HSI and a learnable spectral transformation function (STF). PGU-Net unrolls an alternating optimization procedure into an end-to-end trainable architecture with stages, where each stage sequentially updates the HSI and the STF. Both modules combine learnable proximal networks with differentiable closed-form solvers, enabling physical interpretability while retaining strong representation capacity. Experiments on benchmark datasets (CAVE and NTIRE 2022) with multiple SRFs demonstrate accurate recovery of the STF (degradation operator) and improved reconstruction performance over state-of-the-art SSR methods. Furthermore, evaluations on a real UAV cross-sensor dataset (Headwall Nano HSI and DJI P4 Multispectral MSI) verify the effectiveness and robustness of PGU-Net under truly blind conditions, and suggest that the estimated STF may exhibit land-cover-related differences.
Zhaolin Li, Jinsong Chen, Shanxin Guo +3
a Center for Geo-Spatial Information, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, 518055, PR China · c Shenzhen Engineering Laboratory of Ocean Environmental Big Data Analysis and Application, Shenzhen, 518055, PR China · b University of Chinese Academy of Sciences, Beijing 100049, China