Organizations: Center for Sustainability and the Global Environment (SAGE), University of Wisconsin–Madison USA · Portsmouth AI and Data Science Centre (PAIDS), School of Computing, University of Portsmouth Portsmouth, UK · ESA, ESRIN, 𝜑-lab, Frascati Italy
Geo-Foundation Models (GFMs) have been evaluated across diverse Earth observation tasks and domains, showing strong potential to produce reliable maps even with sparse labels. However, systematic benchmarking of GFMs for Cryosphere applications remains limited, primarily because suitable evaluation datasets are scarce. We address this gap by introducing Cryo-Bench, a benchmark comprising six semantic segmentation datasets covering five cryospheric components: supraglacial debris, glacial lakes under two sensing configurations, sea ice, calving fronts and Antarctic ice-shelf extent. The benchmark includes multispectral, RGB, and synthetic aperture radar observations from regions underrepresented in existing pretraining archives. We evaluate thirteen GFMs alongside U-Net and Vision Transformer baselines trained from scratch under a unified evaluation protocol. With frozen encoders, the U-Net achieves the highest six-dataset average mean intersection over union (mIoU) of 68.22%, slightly exceeding TerraMind (67.86%). However, the paired difference of 0.36 percentage points has a 95% confidence interval of [-0.15, +0.89], indicating that the observed ordering is not statistically significant. Fine-tuning with a fixed learning rate produces mixed outcomes, with leading models such as TerraMind and DOFA experiencing declines in average mIoU (-1.2 and -1.7). In contrast, learning-rate optimization substantially improves fine-tuning performance: DOFA reaches 93.97% mIoU on the RGB glacial lake task, ranking the U-Net fourth, while Scale-MAE and GFM-Swin surpass U-Net on calving fronts. In the few-shot setting, five GFMs outperform U-Net, retaining 91.1% of their full-label accuracy compared with 86.1% for U-Net.
Figures & tables
Fig. 1: Cryo-Bench covers six datasets and five cryospheric components across multispectral, RGB, and radar imagery. Every encoder is paired with the same UPerNet decoder and evaluated under the same three regimes. Dataset colors indicate the input modality.
TABLE I: Cryo-Bench compared with existing benchmarks used to evaluate geospatial foundation models.
Dataset
Component
Sensor
Classes
Train
Val
Test
10% tiles
GSDD
Supraglacial debris
Sentinel-2 + terrain
2
1,496
186
198
~150
GLID
Glacial lakes
Multi-source RGB
2
14,400
1,600
2,367
~1,440
GLB
Glacial lakes
S2 + S1 + terrain (11 band)
2
15,300
1,912
1,913
~1,530
SICD
Sea ice
Sentinel-1 SAR
6
461
51
20
~46
CaFFe
Calving fronts
Multi-mission SAR
4
504
55
122
~50
Shelf-Bench
Ice-shelf extent
SAR
2
42,424
6,974
6,545
~4,242
TABLE II: Datasets associated with component, sensor, class count and split sizes in Cryo-Bench.
Fig. 2: Sample pairs of images and labels for each of the six Cryo-Bench datasets. When the dataset permitted, we chose scenes to show more than one class. When optical inputs make the target class more visible than true color, the inputs are shown as false-color composites: GSDD uses B5, B4, B3 and GLB uses B8, B4, B3 (it lacks B5); GLID is true color (B4, B3, B2), matching Figures 13 to 15. Binary tasks are drawn with a black background and a white target class, as in Figures 13 to 18 ; multi-class tasks maintain the categorical palette and use grey marks to indicate pixels outside the labeled area/region.
Dataset
Class composition of the test split
GSDD
Debris 18.8%, background 81.2%
GLID
Lake 1.9%, background 98.1%
GLB
Lake 4.5%, background 95.5%
SICD
Old ice 45.1%, open water 33.1%, thick FY 16.8%, young ice 2.3%, thin FY 1.5%, new ice 1.1%
CaFFe
Glacier 49.4%, rock 35.7%, N/A 8.0%, ocean/ice melange 6.9%
Shelf-Bench
Ice shelf 48.7%, non-shelf 51.3%
TABLE III: Class composition of the test splits. Foreground share is 1.9% on GLID and 4.5% on GLB, against 48.7% on Shelf-Bench.
Situation
Handling
Dataset bands match pretraining
Matching channels are used directly
Optical-only encoder on a SAR dataset
SAR replicated into a three-band proxy
Single-channel SAR (CaFFe), radar encoder
Fed directly
Single-channel SAR (CaFFe), optical encoder
Repeated three times
Multi-band dataset, three-channel encoder
Terrain and derived layers dropped
U-Net and ViT baselines
All available bands, full tile for U-Net
TABLE IV: Band adaptation per dataset and encoder family.
Model
GSDD
GLID
GLB
SICD
CaFFe
Shelf-Bench
Avg mIoU
Rank
U-Net
73.89
91.58
80.46
20.61
59.82
82.97
68.22
1
TerraMind
74.63
88.26
79.91
33.27
46.64
84.46
67.86
2
DOFA
72.96
92.61
79.39
20.41
50.71
87.08
67.19
3
RemoteCLIP
73.42
90.89
78.09
14.51
56.64
83.59
66.19
4
GFM-Swin
73.00
89.69
78.42
11.55
58.12
85.63
66.07
5
Scale-MAE
73.47
90.91
79.16
6.02
58.19
83.28
65.17
6
TABLE V: Frozen-encoder results with 100% of the available labels, reported as test mIoU (%). Baselines are trained from scratch.
Fig. 3: Frozen-encoder results. (a) Six-dataset average mIoU, ranked, with 95% bootstrap intervals from resampling test scenes; the shaded band and dashed line mark the leading interval and estimate, and TerraMind’s interval overlaps it, so the lead is not resolved. (b) Distribution of those scores across the fifteen models within each dataset, showing how much each task separates representations overall.
Model
GSDD
GLID
GLB
SICD
CaFFe
Shelf-Bench
Avg mIoU
Rank
DOFA
71.36
88.22
73.12
13.67
47.22
86.50
63.35
1
RemoteCLIP
65.56
86.74
70.92
18.37
55.00
80.80
62.90
2
TerraMind
69.90
79.82
74.88
21.44
38.15
82.67
61.14
3
GFM-Swin
64.86
84.79
71.19
9.97
53.04
82.01
60.98
4
Scale-MAE
66.40
84.86
71.25
8.81
53.06
80.29
60.78
5
U-Net
72.00
81.42
75.42
8.51
36.48
78.40
58.71
6
TABLE VI: Few-shot results with 10% of labels and frozen encoders.
Fig. 4: Model robustness under changing evaluation regimes, measured relative to each model’s frozen performance: (a) the proportion of full-label accuracy retained when only 10% of the labels are available; and (b) the change in average mIoU after fine-tuning with the fixed default learning rate of 10−4 . Since the baseline models do not have a frozen–fine-tuned distinction, they are assigned a value of zero in (b).
Model
GSDD
GLID
GLB
SICD
CaFFe
Shelf-Bench
Avg mIoU
Rank
U-Net
73.89
91.58
80.46
20.61
59.82
82.97
68.22
1
GFM-Swin
73.77
83.76
75.98
25.21
57.28
86.53
67.09
2
TerraMind
75.89
87.14
79.04
18.57
56.14
82.93
66.62
3
DOFA
77.22
84.33
79.50
10.55
56.25
84.87
65.45
4
S12-MoCo
75.51
86.08
77.92
22.36
45.11
82.30
64.88
5
SatlasNet
76.74
85.83
77.58
12.61
52.04
83.72
64.75
6
TABLE VII: Fine-tuning at a fixed learning rate of 10−4 , on all six datasets. This is the sensitivity reference against which per-model learning-rate selection (Table VIII-A , Table VIII-B ) is compared.
GLID
CaFFe
GSDD
Model
LR
mIoU
vs 10−4
LR
mIoU
vs 10−4
LR
mIoU
vs 10−4
DOFA
10 -5
93.97
+9.63
10 -5
60.15
+3.90
10 -5
77.08
-0.14
TerraMind
10 -5
93.32
+6.18
10 -5
60.31
+4.18
10 -5
77.08
+1.19
GFM-Swin
10 -5
91.75
+7.98
10 -5
63.49
+6.21
10 -5
74.87
+1.10
U-Net
10 -3
90.44
-1.15
10 -4
59.82
+0.00
10 -4
73.89
+0.00
RemoteCLIP
10 -5
90.28
+12.44
10 -5
58.86
+33.55
10 -5
72.52
+20.35
TABLE VIII-A: Fine-tuning with the learning rate selected per model on the validation split; test mIoU is reported for the selected configuration. RGB glacial lake, calving-front and supraglacial debris datasets, the first group swept; continued in Table VIII-B for the multi-source lake, sea-ice and ice-shelf datasets.
GLB
SICD
Shelf-Bench
Model
LR
mIoU
vs 10−4
LR
mIoU
vs 10−4
LR
mIoU
vs 10−4
TerraMind
10 -5
83.97
+4.93
10 -5
22.19
-2.61
10 -5
88.47
+5.54
CROMA
10 -5
82.38
+5.28
10 -5
21.88
+5.13
10 -5
84.09
+1.55
SatlasNet
10 -5
82.01
+4.42
10 -5
22.22
-9.37
10 -5
84.72
+0.99
GFM-Swin
10 -5
81.38
+5.40
10 -5
20.79
-9.75
10 -5
88.03
+1.50
Scale-MAE
10 -5
80.69
+7.53
10 -5
19.25
+4.07
10 -5
86.85
+2.89
TABLE VIII-B: Fine-tuning with the learning rate selected per model on the validation split; test mIoU is reported for the selected configuration. Multi-source lake, sea-ice and ice-shelf datasets, the second group swept, continuing Table VIII-A .
Fig. 5: Learning rate sensitivity. (a) Change in test mIoU from selecting the learning rate per model, relative to the fixed default; models with no visible bar selected the default. (b) Rank correlation between regimes across the fifteen models.
Fig. 6: Average mIoU by pretraining objective across the three regimes. Error bars are the standard deviation across models within a group; groups with one member have none.
Fig. 7: Left: parameter count against frozen-encoder accuracy. Right: parameter count against the learning rate selected on validation, for the three swept datasets (GLID filled, GSDD half-filled, CaFFe hollow).
Fig. 8: Frozen-encoder performance on the three radar tasks, grouped by whether the encoder saw radar during pretraining. Error bars are standard deviations across models.
Model
Open water
New ice
Young ice
Thin FY ice
Thick FY ice
Old ice
mIoU
TerraMind
77.82
5.04
3.13
5.39
36.35
71.87
33.27
CROMA
68.24
1.73
2.92
3.90
20.83
33.27
21.81
S12-MoCo
45.65
0.19
2.73
0.78
12.05
62.57
20.66
U-Net
49.10
0.37
0.00
0.00
9.28
64.91
20.61
S12-DINO
53.42
0.12
2.25
6.74
8.17
52.84
20.59
DOFA
73.44
3.92
2.99
5.44
19.28
17.36
20.41
TABLE IX: Per-class IoU on the sea-ice task with frozen encoders and full-scene evaluation.
Model
mIoU (%)
Weighted IoU (%)
Change (pts)
TerraMind
33.27
64.64
+31.38
U-Net
20.61
47.31
+26.70
S12-MoCo
20.66
45.65
+24.99
SatlasNet
19.93
44.03
+24.10
S12-DINO
20.59
43.16
+22.57
CROMA
21.81
41.02
+19.20
TABLE X: Frozen-encoder sea-ice results under unweighted and frequency-weighted aggregation, ordered by weighted IoU.
Model
GSDD
GLID
GLB
Shelf-Bench
Avg foreground IoU
DOFA
57.20
85.58
60.65
86.24
72.42
U-Net
58.77
83.56
62.83
82.05
71.80
Scale-MAE
57.83
82.26
60.25
82.25
70.65
TerraMind
60.11
77.11
61.61
83.31
70.53
GFM-Swin
57.66
79.87
58.77
85.00
70.32
RemoteCLIP
57.95
82.22
58.12
82.59
70.22
TABLE XI: Foreground-class IoU on the four binary tasks with frozen encoders.
Encoder
8-bit statistics
Matched statistics
Change (pts)
ssl4eo_moco
5.76
20.66
+14.90
satlasnet_si
6.34
19.93
+13.59
unet_encoder
10.07
20.61
+10.54
prithvi
8.07
15.37
+7.30
spectralgpt
5.43
12.57
+7.13
ssl4eo_data2vec
12.74
17.25
+4.50
TABLE XII: Frozen-encoder sea-ice mIoU under 8-bit and matched input standardization, with full-scene evaluation in both cases.
Fig. 9: Per-class IoU on sea ice, with the reported mIoU, the mean over all six classes, in the final column. The three rare classes are close to zero for every model.
Fig. 10: Per-class IoU on the four calving-front labels with frozen encoders, with the reported mIoU, the mean over all four, in the final column. The N/A region is not a surface class yet scores highest.
Model
Params (M)
Window (px)
GFLOPs/window
Windows/tile
GFLOPs/tile
ms/window
U-Net
14.8
full tile
32.4
1
32.4
1.7
RemoteCLIP
126.8
224
14.0
4
56.1
4.2
SatlasNet
121.2
128
18.8
4
75.1
15.3
GFM-Swin
120.4
192
19.2
4
76.8
11.6
S12-MAE
53.5
224
45.9
4
183.4
4.4
S12-MoCo
53.5
224
45.9
4
183.4
4.4
TABLE XIII: Parameters, per-window and per-tile computational cost, and inference latency, measured on 256 by 256 tiles.
Fig. 11: Frozen-encoder six-dataset average mIoU versus the computational cost of segmenting one 256 by 256 tile, shown on a logarithmic axis. The frozen regime is used to ensure that the vertical axis is directly comparable with Table V and Figure 3 .
Fig. 12: Per-dataset accuracy after learning-rate selection plotted against per-tile computational cost, for the four datasets with a fixed native tile size (RGB and multi-source glacial lakes, supraglacial debris, ice-shelf extent). Calving fronts and sea ice are omitted because their tiles are assembled from variable-sized raw scenes, making per-tile cost less directly comparable across models there. Unlike Figure 11 , the vertical axis represents performance on an individual dataset under tuned fine-tuning, resulting in a different efficiency frontier.
Fig. 13: Predictions on the RGB glacial lake task (GLID). In the third scene, all models except the U-Net fail to detect the western lake.
Fig. 14: Predictions on supraglacial debris (GSDD), the least discriminative task in the benchmark. The input is displayed as a false-color composite (B5, B4, B3), which clearly separates debris from ice better than the natural-color rendering in Figure 2 .
Fig. 15: Predictions on glacial lakes in the multi-source (GLB) configuration. The input is a false-color composite (B8, B4, B3); Since GLB does not include B5, B8 substitutes the standard NIR band, causing water to appear near-black due to strong NIR absorption.
Fig. 16: Predictions on Antarctic ice-shelf extent (Shelf-Bench).
Fig. 17: Predictions on calving-front zones (CaFFe), coloured by zone. Rectangular discontinuities in the foundation model predictions correspond to sliding-window boundaries.
Fig. 18: Predictions on sea ice (SICD). Prithvi assigns a single class to the entire first scene.