Hyperspectral Image Models: Technical Report
Organizations: Department of Information Technology Siddhartha Academy of Higher Education Vijayawada, Andhra Pradesh 521108, India · Department of Computer Science and Engineering Vellore Institute of Technology Bhopal, Madhya Pradesh 466114, India · Department of Computer and Information Sciences Indira Gandhi National Open University New Delhi 110068, India · Department of Computer Science and Engineering Tezpur University Tezpur, Assam 784028, India
Abstract
Hyperspectral remote sensing has advanced across diverse deep learning paradigms, including spectral spatial CNNs, Vision Transformers, Mamba, graph neural networks, Kolmogorov Arnold networks, and self supervised masked autoencoding. Yet progress remains hindered by fragmented repositories, incompatible tensor conventions, and non standardized evaluation. Hyperspectral Image Models addresses these challenges through a modular framework unifying 55 representative models across six paradigms with a common registry, automatic 4D/5D tensor adaptation, and standardized constructors. It integrates 24 benchmark scenes from Airborne, Spaceborne, UAV, and Mars CRISM sensors, with caching, label remapping, PCA, explicit band selection or raw spectra, optional spatial max pooling, and arbitrary PxP patch extraction. To prevent inflated accuracy from overlapping windows, it supports class balanced random partitioning and spatially disjoint regional blocking with Chebyshev guard bands that eliminate train test pixel overlap. Experiments use a single config.yaml with deterministic seeds and complete provenance, generating LaTeX benchmark tables and classification maps. Across 1,320 model scene evaluations and 6,600 seeded runs, scene difficulty dominates architecture, with mean accuracy ranging from 96.40% on Botswana to 56.70% on Houston 2018, versus a 15 point spread across paradigm means. No paradigm universally dominates, while sub 1 M parameter models can match architectures two orders of magnitude larger. Code is publicly available at https://github.com/Tanishq251/Hyperspectral-Image-Models.
Figures & tables
| Framework | Year | Models | Parad. | Scenes (Platform) | Key Architectural Coverage | Spatial Split | Auto-Download |
|---|---|---|---|---|---|---|---|
| DeepHyperX [ 76 ] | 2019 | 13 | 2 | 5 (Airborne) | Classical (SVM), 1D/2D/3D CNN | Fold-based | Script-based |
| HyperSTAR [ 80 ] | 2020 | 5 | 2 | 3 (Airborne) | Classical (SVM), 1D/2D/3D CNN | Predefined Disjoint | Manual |
| TorchGeo [ 77 ] | 2022 | Generic | Vision | Many (General EO) | Generic backbones (ResNet, ViT) | Grid-based | API-based |
| HyTAS [ 78 ] | 2024 | 12 | 1 | 5 (Airborne) | Vision & Spectral Transformers | Random only | Manual |
| Hyperspectral-Image-Models (ours) | 2026 | 55 | 6 | 24 (4 platforms) | CNN, Transformer, Mamba/SSM, GCN, KAN, Self-Supervised | Disjoint & Random | Hugging Face Hub |
| Scene | Size ( ) | Bands | Cls | Labelled px | Min | Max | Imb. | Train % | Sensor |
| Airborne (14) | |||||||||
| Augsburg | 332 485 | 180 | 7 | 78,294 | 575 | 30,329 | 53 | 0.27 | DAS Specim |
| Berlin | 1723 476 | 244 | 8 | 464,671 | 6,672 | 268,642 | 40 | 0.05 | HyMap |
| Chikusei | 2517 2335 | 128 | 19 | 77,592 | – | – | – | 0.73 | Headwall Photonics |
| Dioni | 250 1376 | 176 | 12 | 20,024 | 150 | 6,374 | 42 | 1.80 | AVIRIS-NG |
| Houston 2013 | 349 1905 | 144 | 15 | 15,029 | 325 | 1,268 | 4 | 2.99 | ITRES CASI-1500 |
| Subsystem | Core Component | Modular Capabilities |
|---|---|---|
| Dataset Layer | DatasetLoader | Unified retrieval and caching across 24 multi-platform scenes (Airborne, Spaceborne, UAV, Mars CRISM); automatic scaling, label remapping, background masking. |
| Preprocessing | HyperspectralDataset | Arbitrary spatial patch extraction ( to ); PCA dimensionality reduction to components, band pooling, explicit selection, or full raw spectra. |
| Split Engine | data_split.py | Class-balanced random sampling (fixed counts or fractional percentage ratios) and spatially disjoint component-level partitioning with a guard band that removes evaluation patches overlapping any training patch. |
| Model Registry | @register_model | Unified model ecosystem hosting 55 architectures across 6 paradigms (CNN, ViT, SSM/Mamba, GCN, KAN, MAE); auto-discovery; dynamic 4D/5D InputShapeWrapper . |
| Training | trainer.py | Shared loop with 6 optimizers (Adam, AdamW, SGD, RMSprop, Adagrad, Adadelta), plateau LR decay, validation early stopping, atomic checkpointing. |
| Evaluation | metrics.py | Strict held-out evaluation: OA, AA, Cohen’s , full confusion matrices, per-class accuracies, and cross-seed aggregation ( ). |
| Setting | Value | Source |
|---|---|---|
| Input | ||
| Patch size | dataset.patch_size | |
| Patch stride | 1 | dataset.stride |
| Dimensionality reduction | PCA | preprocessing.dim_reduction_method |
| Retained components | 30 | preprocessing.num_pca_bands |
| Tensor layout | preprocessing.use_channel_dim | |
| Paradigm | Impl. | Years | Par. (M) | MACs (M) | Core mechanism |
|---|---|---|---|---|---|
| CNN | 10 | 2017–2026 | 0.012–1.34 | 1.7–102 | local weight sharing, 2D/3D spectral-spatial convolution |
| Transformer | 16 | 2021–2026 | 0.003–4.44 | 1.8–615 | multi-head self-attention over spectral or patch tokens |
| Mamba / SSM | 20 | 2024–2026 | 0.052–4.25 | 0.2–345 | input-dependent selective state-space scan, linear in length |
| Graph / GCN | 4 | 2024–2026 | 0.022–0.33 | 1.9–28 | message passing over a superpixel adjacency graph |
| KAN | 2 | 2024–2024 | 0.344–0.43 | 1.2–11 | learnable univariate B-spline functions on network edges |
| Self-supervised | 3 | 2024–2024 | 0.513–34.23 | 16.3–4195 | masked spectral-spatial reconstruction, then fine-tuning |
| Indian Pines | Pavia University | WHU-Hi-LongKou | Nili Fossae | |||||
| Model | OA | AA | OA | AA | OA | AA | OA | AA |
| CNN | ||||||||
| SSRN | 84.71 3.9 | 91.99 1.4 | 97.81 0.5 | 98.09 0.4 | 98.36 0.3 | 97.44 0.4 | 96.69 1.0 | 95.29 1.8 |
| DBDA | 80.73 3.9 | 89.34 1.7 | 97.46 0.3 | 97.29 0.5 | 98.05 0.2 | 96.78 0.2 | 96.45 0.5 | 95.76 0.7 |
| DKDMN | 84.82 1.8 | 91.64 1.1 | 95.44 1.2 | 95.82 1.7 | 97.64 0.5 | 95.90 0.8 | 93.73 1.4 | 94.46 0.9 |
| ENL_FCN | 80.73 2.2 | 90.37 1.3 | 94.43 1.7 | 95.32 0.6 | 96.93 0.6 | 95.58 1.4 | 92.22 3.6 | 93.26 2.4 |
| Off target (pp) | |||||||
| Scene | Cls | Labelled px | Kept (%) | Split (%) | Mean | Worst | Overlap (%) |
| Airborne (14) | |||||||
| Augsburg | 7 | 78,294 | 94.4 | 49.5/31.9/18.6 | 2.6 | 10.8 | 58.9 |
| Berlin | 8 | 464,671 | 99.7 | 50.7/30.1/19.2 | 1.1 | 4.6 | 31.9 |
| Chikusei | 19 | 77,592 | 100.0 | 47.2/32.8/20.0 | 5.1 | 16.2 | 6.7 |
| Dioni | 12 | 20,024 | 98.9 | 49.8/29.6/20.6 | 1.8 | 6.8 | 27.6 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Yr | Venue | Par. (M) | MACs (M) | Paper | |
| CNN (10 implemented) | ||||||
| SSRN | 2017 | TGRS | 0.136 | 42.00 | 11 | 10.1109/TGRS.2017.2755542 |
| HybridSN | 2019 | GRSL | 0.535 | 16.01 | 11 | 10.1109/LGRS.2019.2918719 |
| pResNet | 2019 | TGRS | 0.510 | 15.87 | 11 | 10.1109/TGRS.2018.2860125 |
| DBDA | 2020 | Remote Sens. | 0.113 | 47.29 | 11 | 10.3390/rs12030582 |
| ENL_FCN | 2020 | TGRS | 0.089 | 11.18 | 11 | 10.1109/TGRS.2020.3014286 |
| Scene | Classes |
|---|---|
| Augsburg | Forest; Residential Area; Industrial Area; Low Plants; Allotment; Commercial Area; Water |
| Berlin | Forest; Residential; Industrial; Low Plants; Soil; Allotment; Commercial; Water |
| Botswana | Water; Hippo grass; Floodplain grasses 1; Floodplain grasses 2; Reeds; Riparian; Firescar; Island interior; Acacia woodlands; Acacia shrublands; Acacia grasslands; Short mopane; Mixed mopane; Exposed soils |
| Chikusei | Water; Bare soil (farmland); Bare soil (park); Bare soil (roadside); Pavement (asphalt); Pavement (brick); Farmland (paddy); Farmland (other); Greenhouse; Grass (park); Grass (roadside); Tree (park); Tree (roadside); Forest; Building (low-rise); Building (high-rise); Building (factory); Power line; Swimming pool |
| Dioni | Dense Urban Fabric; Mineral Extraction Sites; Non Irrigated Arable Land; Fruit Trees; Olive Groves; Coniferous Forest; Dense Sclerophyllous Vegetation; Sparce Sclerophyllous Vegetation; Sparsely Vegetated Areas; Rocks and Sand; Water; Coastal Water |
| Holden | Analcime; Plagioclase; Prehnite; High-Ca Pyroxene; Serpentine; Margarite |
| Scene | Platform | Mean OA (%) | OA range (min–max) | Highest (model) | |
|---|---|---|---|---|---|
| Botswana | Spaceborne | 96.40 | 11.67 | 37.6–99.8 | Graph ( MS2GCAN ) |
| Pavia Centre | Airborne | 96.31 | 9.35 | 40.0–99.5 | Graph ( MS2GCAN ) |
| WHU-Hi-LongKou | UAV | 95.13 | 7.37 | 51.7–98.8 | Transformer ( MMFormer ) |
| Chikusei | Airborne | 95.03 | 13.48 | 11.8–99.4 | Mamba ( R2Mamba ) |
| Holden | Orbital | 94.82 | 9.29 | 39.5–99.1 | CNN ( SSRN ) |
| Utopia | Orbital | 94.30 | 9.89 | 38.2–98.8 | Mamba ( HyperMamba ) |