cs.CVMar 27, 2026

HyVIC: A Metric-Driven Spatio-Spectral Hyperspectral Image Compression Architecture Based on Variational Autoencoders

Authors: Martin Hermann Paul Fuchs, Behnood Rasti, Begüm Demir

Organizations: Faculty of Electrical Engineering and Computer Science, Technische Universität Berlin, 10587 Berlin, Germany

Abstract

The rapid growth of hyperspectral data archives in remote sensing (RS) necessitates effective compression methods for storage and transmission. Recent advances in learning-based hyperspectral image (HSI) compression have significantly enhanced both reconstruction fidelity and compression efficiency. However, existing methods typically adapt variational image compression models designed for natural images, without adequately accounting for the distinct spatio-spectral redundancies inherent in HSIs. In particular, they lack explicit architectural designs to balance spatial and spectral feature learning, limiting their ability to effectively leverage the unique characteristics of hyperspectral data in RS. To address this issue, in this paper, we aim to study the effects of spatio-spectral feature learning on the rate-distortion (RD) performance of variational HSI compression as a first time in RS. To this end, we propose to use configurable spatial and spectral feature learning blocks within variational HSI compression. To achieve this, we introduce spatio-spectral variational hyperspectral image compression architecture (HyVIC), a configurable variational autoencoder (VAE) for HSI compression. Extensive experiments on two benchmark datasets demonstrate that the trade-off between spatial and spectral feature learning is crucial for the reconstruction fidelity. Motivated by this, we also present a metric-driven strategy to systematically select the hyperparameters of the proposed model. In detail, HyVIC achieves high spatial and spectral reconstruction fidelity across a wide range of compression ratios (CRs) and improves the state of the art by up to 4.66dB in terms of BD-PSNR. Our code and pre-trained model weights are publicly available at https://git.tu-berlin.de/rsim/hyvic .

Explore similar work

May 18, 2026cs.CV

LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift

Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling, and channel dimensionality. As a result, models trained under a fixed spectral configuration often fail to generalize to other sensors. Existing Vision Transformer (ViT) approaches either rely on implicit spectral modeling with fixed channel assumptions or adopt explicit spatial-spectral attention with prohibitive computational cost, leading to a fundamental trade-off between efficiency and expressiveness. In this work, we introduce Low-rank Efficient Spatial-Spectral ViT (LESSViT), a sensor-flexible architecture for cross-spectral generalization. LESSViT is built on LESS Attention, a structured low-rank factorization that models joint spatial-spectral interactions through separable spatial and spectral components, reducing the complexity of full spatial-spectral attention from O(N2C2)O(N^2 C^2) to O(rNC)O(rNC), where NN is the number of spatial tokens, CC is the number of spectral channels, and rr is the rank of the low-rank approximation. We further incorporate channel-agnostic patch embedding and wavelength-aware positional encoding to support flexible spectral inputs. To enable efficient and robust pretraining, we introduce a hyperspectral masked autoencoder (HyperMAE) with decoupled spatial-spectral masking and hierarchical channel sampling. We evaluate LESSViT under a cross-spectral generalization setting that simulates cross-sensor variability. Experiments on the SpectralEarth benchmark demonstrate that LESSViT improves robustness under spectral shifts while remaining competitive in-distribution, and explicit and efficient spatial-spectral modeling is essential for scalable and generalizable hyperspectral representation learning.
Mar 23, 2026cs.CV

A Variational Latent-Space Framework for Uncertainty-Aware Spectral Image Emulation

Synthetic spectral image generation is essential for remote sensing simulation and mission design, yet physically based radiative transfer models (RTMs) remain computationally expensive. Existing learning-based emulators reduce this cost, but are mostly deterministic parameter-to-spectrum regressors with limited spatial modeling and uncertainty information. We formulate spectral image emulation as a parameter-conditioned latent-variable problem and propose a variational autoencoder (VAE)-based framework combining nonlinear spectral-image representations, fast inference, and per-pixel uncertainty estimates. The framework is instantiated at spectrum and spatial--spectral levels through two-step VAE pretraining and latent mapping. We evaluate it on PROSAIL-simulated hyperspectral vegetation cubes (211 bands) and real Sentinel-3 OLCI multispectral ocean-colour imagery (21 bands) against classical regression emulators and a deep CNN baseline. Results show that no single architecture is optimal: pixel-to-pixel models perform best on controlled hyperspectral simulations, whereas a fully convolutional VAE is more robust on noisy real observations with missing or contaminated pixels. VAE-based emulators also achieve high throughput for large-scale generation. On Sentinel-3, the spatial--spectral VAE provides predictive intervals closer to empirical errors than pixel-wise neural and classical emulators, although absolute calibration remains incomplete. A look-up-table-based retrieval experiment further shows that reconstruction fidelity alone does not ensure reliable leaf area index and chlorophyll retrieval. Emulators should therefore be evaluated in representative remote-sensing end-use scenarios, not by reconstruction metrics alone.
Apr 30, 2026eess.IV

Spectral Dynamic Attention Network for Hyperspectral Image Super-Resolution

Hyperspectral image super-resolution is essential for enhancing the spatial fidelity of HSI data, yet existing deep learning methods often struggle with substantial spectral redundancy and the limited non-linear modeling capacity of standard feed-forward networks (FFNs). To address these challenges, we propose Spectral Dynamic Attention Network (SDANet), a framework designed to adaptively suppress redundant spectral interactions. SDANet integrates two key components: 1) Dynamic Channel Sparse Attention (DCSA) module that computes channel-wise correlations and selectively preserves the most informative attention responses through dynamic and data-dependent sparsification. 2) Frequency-Enhanced Feed-Forward Network (FE-FFN) that jointly models spatial and frequency-domain representations to enhance non-linear expressiveness. Extensive experiments on two benchmark datasets demonstrate that SDANet achieves state-of-the-art HISR performance while maintaining competitive efficiency. The code will be made publicly available at https://github.com/oucailab/SDANet.