HyperSAM: A Promptable Foundation Model for Hyperspectral Remote Sensing
Organizations: School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an 710049, China · Faculty of Electronic and Information Engineering, Xi’an Jiaotong University, Xi’an 710049, China · State Key Laboratory of Remote Sensing and Digital Earth, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China · Helmholtz-Zentrum Dresden-Rossendorf, 09599 Freiberg, Germany · Faculty of Electrical and Computer Engineering, University of Iceland, 101 Reykjavik, Iceland · School of Information and Communication Technology, Griffith University, Nathan, QLD 4111, Australia · School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an 710049, China
Abstract
Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. First, large hyperspectral corpora rarely provide high spatial resolution together with reliable dense annotations. Second, many hyperspectral models are still trained almost from scratch, so the geometric and interactive priors learned by modern vision foundation models are not fully reused. To alleviate these issues, we \highlight{present} \textbf{HyperSAM}, a promptable hyperspectral foundation model that couples a data-centric hyperspectral synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). On the data side, HyperSAM synthesizes full-spectrum hyperspectral cubes from high-resolution SpaceNet multispectral imagery through a physics-informed abundance-transfer generator, while SAM3-derived pseudo-masks provide object-centric supervision. On the model side, the latest implementation uses a frozen SAM3 RGB image branch, a trainable hyperspectral side encoder initialized from the RGB vision transformer (ViT), ControlNet-style zero-initialized feature injection, and a lightweight mixture-of-experts mask refiner. To enhance training robustness against noisy pseudo-labels, Cross-modal Sample Selection (CromSS)-style confidence selection is incorporated for noisy-label weighting. Extensive experiments show that HyperSAM obtains strong generalization on diverse hyperspectral tasks (e.g., classification, anomaly detection, change detection, target detection, and airborne oil-spill mapping) and that high-quality synthetic hyperspectral data can be more effective than simply scaling noisy hyperspectral supervision.
Figures & tables
| Task | Supervision or prompt | Mask generation | Feature or score construction | Decision rule |
| HC | One randomly sampled support pixel per class from a seeded pool | 32 points/side; IoU 0.3; stability 0.4; NMS 0.7 | Select the smallest covering mask and average its features; use the single-pixel embedding if no mask covers the support | Assign candidate masks to the nearest prototype by cosine similarity |
| HAD | No labels, prompts, or target spectra | 96 points/side; IoU 0.4; stability 0.4; NMS 0.7 | Suppress large background masks and retain compact proposals; default area-ratio threshold is 0.0009 | Convert retained mask proposals into an anomaly response map |
| HCD | No changed-pixel labels | 32 points/side for each date; IoU 0.3; stability 0.4; NMS 0.7 | Average dense features inside masks and compute cosine-distance change scores for both temporal directions | Merge the two directional maps by maximum response and threshold at quantile 0.74 |
| HTD | One target spectrum or one support prompt | 32 points/side; IoU 0.3; stability 0.4; NMS 0.7 | Locate a target-like support region and use its mean feature as the target prototype | Score candidate masks by cosine similarity to the target prototype and apply light spatial post-processing |
| HC (Indian Pines / PaviaU) | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | SSFTT [ 43 ] | TGRS-ViT [ 44 ] | HyperSIGMA-LP [ 6 ] | DOFA-LP [ 12 ] | HyperFree [ 7 ] | SpectralEarth [ 16 ] | SAM3 [ 15 ] | Ours | SSFTT [ 43 ] | TGRS-ViT [ 44 ] | HyperSIGMA-LP [ 6 ] | DOFA-LP [ 12 ] | HyperFree [ 7 ] | SpectralEarth [ 16 ] | SAM3 [ 15 ] | Ours |
| OA (%) | 58.43 | 54.14 | 41.74 | 60.01 | 51.35 | 57.22 | 62.49 | 69.37 | 69.34 | 63.99 | 65.27 | 64.22 | 75.23 | 65.48 | 84.61 | 89.12 |
| AA (%) | 70.07 | 65.28 | 46.99 | 73.64 | 62.07 | 69.90 | 75.80 | 77.46 | 71.59 | 75.85 | 72.34 | 63.09 | 62.49 | 56.02 | 85.04 | 87.38 |
| Kappa (%) | 53.41 | 48.24 | 33.50 | 56.09 | 47.40 | 52.28 | 58.60 | 65.11 | 60.65 | 56.01 | 54.67 | 54.36 | 64.53 | 54.57 | 80.45 | 85.81 |
| HAD (Beach-1 / Beach-2) | ||||||||||||||||
| Metric | RXD [ 4 ] | Auto-AD [ 45 ] | HyperSIGMA [ 6 ] | DOFA [ 12 ] | TDD [ 46 ] | HyperFree [ 7 ] | SpectralEarth [ 16 ] | Ours | RXD [ 4 ] | Auto-AD [ 45 ] | HyperSIGMA [ 6 ] | DOFA [ 12 ] | ADLR [ 47 ] | HyperFree [ 7 ] | SpectralEarth [ 16 ] | Ours |
| Configuration | Component | Downstream performance | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Training data | Network structure | Training loss | PaviaU OA | Beach-2 ODP | Hermiston IoU | Airport ODP | |||||
| HyperFree | Raw-MS | Syn-HSI | RGB | HSI | MoE | Conf. | |||||
| HyperFree training data | ✓ | – | – | ✓ | ✓ | ✓ | ✓ | 79.33 | 1.3903 | 68.40 | 1.2500 |
| Raw multispectral input | – | ✓ | – | ✓ | ✓ | ✓ | ✓ | 85.84 | 1.3576 | 66.73 | 1.3088 |
| Synthetic hyperspectral input | – | – | ✓ | ✓ | ✓ | – | ✓ | 84.92 | 1.4452 | 66.61 | 1.3544 |
| Without frozen RGB branch | – | – | ✓ | – | ✓ | ✓ | ✓ | 84.89 | 1.3126 | 66.45 | 1.3208 |
| Initialization | PaviaU OA | Beach-2 ODP | Hermiston IoU | Airport ODP |
|---|---|---|---|---|
| Random | 83.59 | 1.3967 | 68.43 | 1.3126 |
| Zero | 89.12 | 1.4641 | 70.31 | 1.3297 |
| Configuration | PaviaU OA | Beach-2 ODP | Hermiston IoU | Airport ODP |
|---|---|---|---|---|
| w/o reconstruction | 81.64 | 1.3846 | 63.88 | 1.2715 |
| w/o cosine consistency | 80.26 | 1.2681 | 61.15 | 1.2194 |
| w/o projection consistency | 87.91 | 1.4218 | 68.74 | 1.3312 |
| w/o abundance sparsity | 83.52 | 1.3467 | 65.21 | 1.2916 |
| w/o total variation | 88.73 | 1.4479 | 69.42 | 1.3421 |
| Full generator objective | 89.12 | 1.4641 | 70.31 | 1.3297 |
| Configuration | PaviaU OA | Beach-2 ODP | Hermiston IoU | Airport ODP |
|---|---|---|---|---|
| w/o Dice | 85.06 | 1.4159 | 67.16 | 1.3226 |
| w/o BCE | 86.24 | 1.4273 | 67.81 | 1.3396 |
| w/o confidence weighting | 87.48 | 1.4524 | 68.96 | 1.3587 |
| w/o | 87.17 | 1.4592 | 68.27 | 1.3472 |
| Full training objective | 89.12 | 1.4641 | 70.31 | 1.3297 |
| Method | OA (%) | AA (%) | Kappa (%) | Time (s) | Energy (Wh) | CO 2 e (g) |
|---|---|---|---|---|---|---|
| HyperSAM | 66.91 | 74.09 | 62.02 | 21.11 1.98 | 0.624 0.223 | 0.2964 0.1059 |
| HyperFree | 50.36 | 58.06 | 43.62 | 23.40 18.94 | 0.244 0.111 | 0.1159 0.0527 |
| SSFTT | 54.85 | 65.07 | 49.51 | 8.30 0.06 | 0.087 0.001 | 0.0413 0.0005 |
| TGRS-ViT | 49.73 | 63.49 | 44.92 | 8.16 0.13 | 0.120 0.002 | 0.0570 0.0010 |
| HyperSIGMA-LP | 40.69 | 46.99 | 33.53 | 27.01 37.63 | 0.274 0.401 | 0.1302 0.1905 |
| DOFA-LP | 59.32 | 64.48 | 55.34 | 13.69 9.88 | 0.181 0.168 | 0.0860 0.0798 |