Token Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification
Authors: Yimin Zhu, Mahmood Elahi, Lincoln Linlin Xu
Organizations: Department of Geomatics Engineering, University of Calgary, Canada · Department of Electrical and Software Engineering, University of Calgary, Canada
Although hyperspectral images (HSIs) provide rich spectral-spatial information, accurate pixel-level classification remains challenging because of spectral-spatial heterogeneity and complex spatial structures. Existing vision state space models (Mamba) typically construct sequences according to predefined spatial neighborhoods, without explicitly accounting for semantic similarity or spatial non-stationarity. To address this limitation, we propose Token Clustering and Semantic Sequence Mamba (STMamba), which organizes sparse tokens into semantically coherent sequences for hyperspectral image classification with the following features. First, at the macro level, a hierarchical encoder decoder progressively selects semantic tokens with the Token Clustering Module (TCM) and restores dense features using a parameter-free Cross-scale Neighborhood Attention (CNA) Upsampler. Second, at the micro level, TCM first identifies representative cluster centers through density-aware clustering and estimates soft memberships based on feature similarity. A quadtree-based dynamic selection strategy then retains sparse and spatially distributed tokens from each semantic cluster, forming coherent semantic-token sequences while reducing redundant pixel-wise representations. Third, parallel Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture complementary long-range spatial and spectral dependencies within homogeneous semantic token sequences while suppressing irrelevant interactions across heterogeneous regions. Experimental results on three large-scale benchmark datasets demonstrate that STMamba outperforms the SOTA methods with respect to quantitative and qualitative results.
Figures & tables
Figure 1: (a) Traditional Vision Mamba Token Sequencing organizes an entire patch or image using predefined scanning paths despite heterogeneous spatial patterns [ 6 ] . (b) Our method adaptively selects sparse representative tokens through density-aware clustering and organizes them into homogeneous semantic sequences for Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM).
Figure 2: Overview of the proposed STMamba architecture. It consists of four hierarchical stages with Token Clustering Modules (TCM) for semantic token grouping, quadtree-based sparse token selection, and Spatial and Spectral Semantic-wise Sequencing Mamba modules (Spa-SWSM and Spe-SWSM) for spatial–spectral modeling. Hierarchical features are progressively encoded and subsequently recovered by a cross-scale Neighborhood Attention upsampler to generate the final prediction map.
PU
No.
Class
Color
Tr./Val./Te.
No.
Class
Color
Tr./Val./Te.
1
Asphalt
5/5/6621
6
Bare soil
5/5/5019
2
Meadows
5/5/18639
7
Bitumen
5/5/1320
3
Gravel
5/5/2089
8
Self-blocking bricks
5/5/3672
4
Trees
5/5/3054
9
Shadows
5/5/937
5
Painted metal sheets
5/5/1335
Table I: Land-cover classes and sample numbers of the three datasets.
Class
Color
Train Num.
Traditional
Transformer
CNN
Mamba-based
Clustering-based
Mamba + Clustering
Random Forest
CPFormer [ 54 ]
S 2 VNet [ 5 ]
SDMamba [ 55 ]
MambaLG [ 56 ]
MambaHSI [ 10 ]
MambaHSI+ [ 6 ]
S 2 Mamba [ 57 ]
DSCC [ 31 ]
PSFormer [ 58 ]
DMSGer [ 59 ]
Ours
1
5
66.24±5.52
72.44±6.70
74.06±12.31
78.40±6.29
77.88±5.42
77.35±13.27
81.18±11.34
76.20±7.29
69.66±8.59
74.85±8.81
55.81± 4.09
80.49±6.36
2
5
51.31±14.13
74.04±16.19
60.40±16.92
64.91±14.65
31.18±9.20
76.61±10.62
71.56±13.94
61.69±15.09
68.96±8.97
62.70±17.64
84.39± 4.76
78.86±10.04
3
5
48.29±11.90
74.51±11.23
84.30±11.58
73.45±12.64
92.62±2.17
66.57±14.81
77.20±12.81
75.25±12.42
80.64±10.54
89.23±13.45
91.52± 6.25
89.82±13.04
4
5
81.81±7.67
84.32±6.61
85.65±14.14
87.91±6.50
96.59±0.90
72.32±7.78
88.91±5.65
94.55±3.40
68.95±6.84
96.47± 2.61
68.76±6.82
96.61±2.89
5
5
98.96±0.39
99.56±0.82
99.23±0.98
98.76±0.98
99.87±0.27
99.67±0.51
98.50±2.88
99.70±0.80
96.04±4.03
100.00± 0.00
99.38±0.24
99.89±0.15
Table II: Quantitative classification results on the PU dataset.
Class
Color
Train Num.
Traditional
Transformer
CNN
Mamba-based
Clustering-based
Mamba + Clustering
Random Forest
CPFormer [ 54 ]
S 2 VNet [ 5 ]
SDMamba [ 55 ]
MambaLG [ 56 ]
MambaHSI [ 10 ]
MambaHSI+ [ 6 ]
S 2 Mamba [ 57 ]
DSCC [ 31 ]
PSFormer [ 58 ]
DMSGer [ 59 ]
Ours
1
30
93.71±2.73
90.26±4.58
95.91±4.82
91.47±6.73
95.24±4.43
83.46±1.78
85.16± 1.57
89.90±5.22
82.80±6.85
94.57±5.36
96.26±1.58
94.32±5.31
2
30
93.58±4.54
93.82±4.88
92.61±5.58
90.53±4.93
97.22±3.43
92.54±4.53
95.10±2.69
92.85±6.51
84.18±6.26
93.97±4.88
97.72± 0.95
97.19±1.73
3
30
94.15±2.12
98.50±2.25
99.72±0.44
99.88±0.15
99.43±0.56
92.52±4.05
92.16±8.33
99.81±0.11
95.00±4.99
100.00± 0.00
100.00± 0.00
100.00±0.00
4
30
94.24±2.41
94.29±4.33
93.44±2.05
93.42±2.44
96.57±2.78
96.91±2.08
96.66±2.29
93.51± 1.76
81.15±5.34
97.71±2.27
89.53±2.56
97.58±2.34
5
30
91.95±4.73
99.06±0.84
99.32±1.29
100.00± 0.00
99.31±1.83
99.30±0.60
99.58±0.32
100.00± 0.00
97.18±3.29
99.61±1.16
100.00± 0.00
98.14±5.05
Table III: Quantitative classification results on the HU13 dataset.
Class
Color
Train Num.
Traditional
Transformer
CNN
Mamba-based
Clustering-based
Mamba + Clustering
Random Forest
CPFormer [ 54 ]
S 2 VNet [ 5 ]
SDMamba [ 55 ]
MambaLG [ 56 ]
MambaHSI [ 10 ]
MambaHSI+ [ 6 ]
S 2 Mamba [ 57 ]
DSCC [ 31 ]
PSFormer [ 58 ]
DMSGer [ 59 ]
Ours
1
30
91.70±2.38
83.75±7.07
90.61±5.74
91.95±3.22
89.82±3.18
82.31±5.77
82.34±6.12
91.37±2.92
72.70±6.26
90.29±4.46
89.72± 1.98
91.14±2.52
2
30
84.00±4.39
71.97±6.20
81.09±8.07
82.06±3.71
79.98± 2.99
79.80±6.95
81.74±4.62
75.13±5.48
71.08±5.27
71.55±24.26
60.96±4.01
81.38±4.29
3
30
98.64±0.89
99.83±0.27
99.63±0.39
99.80±0.24
99.98±0.05
96.17±5.89
99.91±0.16
99.96±0.09
98.86±1.44
90.00±30.00
100.00± 0.00
100.00± 0.00
4
30
88.78±2.19
94.08±3.83
95.90±2.06
96.42±1.13
94.34±1.12
92.10±1.62
96.81±1.26
96.36± 1.09
85.59±4.09
82.00±27.47
87.85±2.33
94.51±1.88
5
30
69.09±4.46
90.28±4.80
81.15±2.88
82.63±2.73
83.05±2.84
84.94± 1.39
89.16±3.14
80.13±3.94
79.25±5.93
77.10±26.07
82.53±2.50
90.77±3.54
Table IV: Quantitative classification results on the HU18 dataset.
Figure 3: Visual comparison on the PU dataset. (a) Random Forest, (b) CPFormer, (c) S 2 VNet, (d) SDMamba, (e) MambaLG, (f) MambaHSI, (g) MambaHSI+, (h) S 2 Mamba, (i) DSCC, (j) PSFormer, (k) DMSGer, (l) Ours, (m) Ground Truth, and (n) ICA Map.
Figure 4: Visual comparison on the HU13 dataset. (a) Random Forest, (b) CPFormer, (c) S 2 VNet, (d) SDMamba, (e) MambaLG, (f) MambaHSI, (g) MambaHSI+, (h) S 2 Mamba, (i) DSCC, (j) PSFormer, (k) DMSGer, (l) Ours, (m) Ground Truth, and (n) ICA Map.
Figure 5: Visual comparison on the HU18 dataset. (a) Random Forest, (b) CPFormer, (c) S 2 VNet, (d) SDMamba, (e) MambaLG, (f) MambaHSI, (g) MambaHSI+, (h) S 2 Mamba, (i) DSCC, (j) PSFormer, (k) DMSGer, (l) Ours, (m) Ground Truth, and (n) ICA Map.
Figure 6: Selected semantic tokens and their scanning order across four hierarchical stages.
Figure 7: Visualization of clustering centers across four STMamba stages on PU, including sampled anchors, aggregated centers, their spatial distributions, and semantic memberships.
Setting
Pos.
SWSM
Quad.
CNA
Ldiv
PU
HU13
HU18
OA
AA
κ
OA
AA
κ
OA
AA
κ
w/o Pos.
×
✓
✓
✓
✓
84.10
93.57
80.17
95.77
96.46
95.42
71.14
85.94
65.04
w/o SWSM
×
×
✓
✓
×
80.80
89.08
76.10
92.42
93.32
91.80
70.02
85.55
63.88
w/o Quad-Tree
✓
✓
×
✓
✓
85.94
93.85
82.36
95.30
96.06
94.92
70.65
85.73
64.47
w/o CNA
✓
✓
✓
×
✓
83.57
88.95
79.22
95.03
95.92
94.63
66.64
84.62
60.23
w/o Ldiv
✓
✓
✓
✓
×
85.88
93.30
82.23
95.61
96.35
95.25
70.40
85.79
64.24
Table V: Ablation Study of the Main Components of the Proposed STMamba on Three Datasets.
Figure 8: PCA visualization of decoder features on the HU13 dataset by using CNA upsampler (left) and bilinear-interpolation (right).
Hyperparameter
Value
PU
HU13
HU18
OA
AA
κ
OA
AA
κ
OA
AA
κ
λs
0.01
86.49
93.26
82.91
95.57
96.31
95.21
71.44
86.73
65.49
0.05
87.67
94.95
84.37
96.63
97.21
96.36
72.09
86.91
66.23
0.10
88.59
94.65
85.47
97.11
97.65
96.88
70.44
86.19
64.32
0.20
88.52
93.13
85.27
96.96
97.50
95.71
71.48
86.49
65.35
0.30
88.28
93.87
85.01
95.89
96.69
95.56
71.97
86.45
66.11
Table VI: Sensitivity Analysis of Hyperparameters.
Setting
Spa- SWSM
Spe- SWSM
Membership Score Order Scan
Membership Gate
PU
HU13
HU18
OA
AA
κ
OA
AA
κ
OA
AA
κ
w/o Spa-SWSM
×
✓
✓
✓
84.72 (-2.26)
93.76 (-0.20)
80.96 (-2.61)
95.13 (-1.00)
95.90 (-0.83)
94.74 (-1.07)
68.80 (-2.60)
84.91 (-1.18)
62.55 (-2.74)
w/o Spe-SWSM
✓
×
✓
✓
84.00 (-2.98)
93.45 (-0.51)
80.13 (-3.44)
95.27 (-0.86)
95.96 (-0.77)
94.88 (-0.93)
70.95 (-0.45)
86.09 (0.00)
64.89 (-0.40)
Spatial-order Scan
✓
✓
×
✓
85.57 (-1.41)
93.74 (-0.22)
81.92 (-1.65)
95.42 (-0.71)
96.28 (-0.45)
95.05 (-0.76)
70.43 (-0.97)
85.49 (-0.60)
64.25 (-1.04)
w/o Membership Gate
✓
✓
✓
×
86.30 (-0.68)
93.76 (-0.20)
82.84 (-0.73)
95.66 (-0.47)
96.36 (-0.37)
95.30 (-0.51)
69.78 (-1.62)
86.18 (+0.09)
63.66 (-1.63)
Full Model
✓
✓
✓
✓
86.98
93.96
83.57
96.13
96.73
95.81
71.40
86.09
65.29
Table VII: Ablation study of the internal designs of SWSM. Values in parentheses indicate the performance difference relative to the full model.
Figure 9: Visualization of clustering membership and hard clustering map on the HU18 dataset at (a) stage1, (b) stage2, (c) stage3, and (d) stage4.
Figure 10: Performance of the proposed method under different numbers of training samples per class on the PU, HU13, and HU18 datasets. Results are reported over 10 independent runs.
Method
PU
HU13
HU18
Params (M)
GFLOPs
Params (M)
GFLOPs
Params (M)
GFLOPs
MambaLG
0.293
80.273
0.297
262.881
0.289
543.230
MambaHSI
0.412
5.938
0.418
26.078
0.407
21.108
MambaHSI+
1.631
9.703
1.643
38.652
1.636
49.168
DSCC
24.580
73.570
24.616
76.682
24.537
69.397
Ours
0.602
31.952
0.605
106.367
0.599
212.674
Table VIII: Comparison of model parameters and computational complexity.
Figure 11: Visualization of the hidden influence matrices in SWSM. At stage 1, (a) shows the forward semantic-topology scanning influence matrix (left) and the corresponding permutation-restored token influence matrix (right) for cluster 1, while (b) shows the same visualization for cluster 8. At stage 3, (c) and (d) present the corresponding results for clusters 1 and 8, respectively.
Beijing Institute of Technology, Beijing 100081, China · Beijing Institute of Technology Chongqing Innovation Center, Chongqing 401135, China · Key Laboratory of Photoelectronic Imaging Technology and System, Ministry of Education of China, Beijing 100081, China