cs.CVSep 29, 2026

HyperSAM: A Promptable Foundation Model for Hyperspectral Remote Sensing

Authors: Li Pang, Xinqiao Wu, Jing Yao, Pedram Ghamisi, Jun Zhou, Zhengchao Chen, Deyu Meng, Xiangyong Cao

Organizations: School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an 710049, China · Faculty of Electronic and Information Engineering, Xi’an Jiaotong University, Xi’an 710049, China · State Key Laboratory of Remote Sensing and Digital Earth, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China · Helmholtz-Zentrum Dresden-Rossendorf, 09599 Freiberg, Germany · Faculty of Electrical and Computer Engineering, University of Iceland, 101 Reykjavik, Iceland · School of Information and Communication Technology, Griffith University, Nathan, QLD 4111, Australia · School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an 710049, China

Abstract

Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. First, large hyperspectral corpora rarely provide high spatial resolution together with reliable dense annotations. Second, many hyperspectral models are still trained almost from scratch, so the geometric and interactive priors learned by modern vision foundation models are not fully reused. To alleviate these issues, we \highlight{present} \textbf{HyperSAM}, a promptable hyperspectral foundation model that couples a data-centric hyperspectral synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). On the data side, HyperSAM synthesizes full-spectrum hyperspectral cubes from high-resolution SpaceNet multispectral imagery through a physics-informed abundance-transfer generator, while SAM3-derived pseudo-masks provide object-centric supervision. On the model side, the latest implementation uses a frozen SAM3 RGB image branch, a trainable hyperspectral side encoder initialized from the RGB vision transformer (ViT), ControlNet-style zero-initialized feature injection, and a lightweight mixture-of-experts mask refiner. To enhance training robustness against noisy pseudo-labels, Cross-modal Sample Selection (CromSS)-style confidence selection is incorporated for noisy-label weighting. Extensive experiments show that HyperSAM obtains strong generalization on diverse hyperspectral tasks (e.g., classification, anomaly detection, change detection, target detection, and airborne oil-spill mapping) and that high-quality synthetic hyperspectral data can be more effective than simply scaling noisy hyperspectral supervision.

Figures & tables

Explore similar work

CardsList
  1. Hyperspectral Image Models: Technical Report

    Sep 30, 2026Tanishq Rachamalla, Aryan Das, Srishti Kaushik +1Hyperspectral ImageDeep Learning

  2. HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

    May 17, 2026Guanyiman Fu, Jingtao Li, Zihang Cheng +8Hyperspectral Image

  3. Benchmarking Hyperspectral Foundation Models for Hyperspectral Unmixing

    Sep 23, 2026Edgard Dabier, Christophe Kervazo, Pietro Gori +1Hyperspectral ImageWireless Foundation Models