Organizations: Space Applications Centre, ISRO, Ahmedabad, India · Indian Institute of Science Education and Research Bhopal, Bhopal, India · Centre of Studies in Resources Engineering (CSRE), Indian Institute of Technology Bombay, India
Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in RSFMs. FairRSFM maps georeferenced samples from 14 terrestrial biome classes into six ecologically meaningful macro-groups and evaluates models under a unified frozen-backbone evaluation protocol. The benchmark covers four downstream datasets: m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K with Dynamic World label maps. Using Prithvi-EO-2.0, SatMAE, and DOFA across three random seeds, we show that aggregate performance consistently masks biome-dependent disparities across architectures and tasks. For example, Prithvi-EO-2.0 reaches 90.98% overall macro-F1 on m-EuroSAT but a mean worst-group score of only 83.72%, while m-SA-Crop-Type drops from 27.30% overall mIoU to 18.47% in the Xeric and Mineralogical group. We further evaluate Biome-Orthogonal Linear Probing (BOLP), Dynamic Biome Reweighting (DBR), and GroupDRO as complementary mitigation baselines. Their effectiveness is model- and task-dependent; for example, BOLP improves Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without updating the RSFM backbone. FairRSFM provides a reusable protocol for diagnosing and mitigating ecological robustness gaps in remote sensing foundation models. Code and datasets are available at: https://github.com/aminurhossain/FairRSFM.
Figures & tables
Figure 1 : Overview of FairRSFM. Georeferenced samples are mapped from 14 terrestrial biomes to six macro-groups and evaluated under a frozen-backbone downstream evaluation protocol. BOLP removes biome-associated directions from frozen RSFM embeddings before training the task head.
Dataset
Task
Classes
Train
Val
Test
Notes
m-EuroSAT [ 12 ]
Classification
10
2,000
1,000
1,000
Sentinel-2, land cover
m-BigEarthNet [ 12 , 25 ]
Multi-label classification
43
20,000
1,000
1,000
Sentinel-2, multi-label land cover
m-SA-Crop-Type [ 12 ]
Segmentation
10
3,000
1,000
1,000
Sentinel-2, crop type
MMEarth20K [ 17 , 2 ]
Segmentation
9
16,000
2,000
2,000
Dynamic World label maps
Table 1 : FairRSFM dataset suite. GEO-Bench datasets are selected for classification and segmentation; the MMEarth20K subset increases dense land-cover segmentation diversity. Sample counts for GEO-Bench tasks follow the modified benchmark setting.
Configuration
Prithvi-EO-2.0 [ 27 ]
SatMAE [ 3 ]
DOFA [ 33 ]
Backbone
Spatio-Temporal ViT (300M)
ViT-Large (304M)
ViT-Base (86M)
Embedding Dim
1024 ( 16×16 patch)
1024 ( 16×16 patch)
768 ( 16×16 patch)
Input Bands
6 HLS Bands
10 Sentinel-2 Bands
9–12 Wavelength Bands
Encoder Status
Frozen (0 updates)
Frozen (0 updates)
Frozen (0 updates)
Classification Head
Linear / MLP probe
Linear probe
BatchNorm1d + Linear
Segmentation Head
FCN decoder
Conv decoder
Multi-level UPerNet [ 33 ]
Table 2 : Unified benchmarking configurations across foundation models.
Dataset
Split
Pheno.
Biomass
Trans.
Cryo.
Hydro.
Xeric
Unk.
m-EuroSAT
Train
1,430
50
481
–
–
–
39
Valid
741
15
227
–
–
–
17
Test
724
25
239
–
–
–
12
m-BigEarthNet
Train
11,269
356
6,291
–
–
–
2,084
Valid
518
17
335
–
–
–
130
Test
535
14
336
–
–
–
115
Table 3 : Per-macro-group sample counts across Train/Validation/Test splits for all benchmark datasets. Empty cells indicate absent groups. Unknown or unmatched samples are excluded from worst-group scores.
Dataset
Method
Test metric
Mworst
ECE ↓
NFR ↓
EOdd ↓
DPM ↓
m-EuroSAT
ERM
90.98 ± 2.36
83.72 ± 0.87
4.02 ± 1.64
5.96 ± 3.28
20.50 ± 1.75
9.91 ± 0.21
BOLP
93.42 ± 0.35
84.34 ± 0.00
2.37 ± 0.04
9.82 ± 0.30
17.06 ± 0.16
9.61 ± 0.03
DBR
91.36 ± 1.01
84.35 ± 0.50
5.50 ± 2.29
5.85 ± 0.70
21.23 ± 1.39
10.90 ± 0.17
GroupDRO
94.14 ± 0.47
84.34 ± 0.00
3.17 ± 0.40
10.34 ± 0.51
17.31 ± 0.12
10.04 ± 0.10
m-BigEarthNet
ERM
60.09 ± 0.99
46.12 ± 1.51
2.12 ± 0.28
31.14 ± 1.99
42.93 ± 1.90
12.66 ± 0.27
BOLP
62.75 ± 0.16
50.27 ± 0.63
2.07 ± 0.08
20.92 ± 0.39
45.93 ± 0.76
13.29 ± 0.11
Table 4 : Prithvi-EO-2.0: Mean and Std over seeds across datasets.
Model
Method
m-EuroSAT (Macro-F1)
m-BigEarthNet (F1@opt)
m-SA-Crop-Type (mIoU)
Phenological
High Biomass
Transitional
Phenological
High Biomass
Transitional
Transitional
Xeric
Prithvi-EO-2.0
ERM
88.96 ± 3.38
84.56 ± 0.31
86.71 ± 3.79
59.67 ± 0.48
61.14 ± 6.05
46.04 ± 1.85
28.55 ± 0.19
18.47 ± 0.16
BOLP
92.28 ± 0.53
84.34 ± 0.00
93.02 ± 0.41
61.85 ± 0.72
54.03 ± 1.68
50.27 ± 0.77
27.70 ± 0.67
19.38 ± 0.24
DBR
88.57 ± 0.96
85.00 ± 0.00
85.96 ± 2.74
55.59 ± 0.18
64.53 ± 0.13
48.03 ± 0.78
26.49 ± 1.93
19.81 ± 0.42
GroupDRO
92.82 ± 0.70
84.34 ± 0.00
93.68 ± 0.50
49.59 ± 0.67
56.78 ± 5.15
46.82 ± 0.60
28.52 ± 0.11
19.09 ± 0.26
SatMAE
ERM
93.99 ± 0.14
74.45 ± 0.51
93.62 ± 0.13
52.06 ± 0.22
47.97 ± 3.69
41.33 ± 0.36
27.45 ± 0.48
18.07 ± 0.23
Table 7 : Biome-wise performance for m-EuroSAT (Macro-F1), m-BigEarthNet (F1@opt), and m-SA-Crop-Type (mIoU) across foundation models and debiasing methods.
Model
Method
Pheno.
Biomass
Trans.
Cryo.
Hydro.
Xeric
Prithvi-EO-2.0
ERM
44.19 ± 1.39
41.20 ± 1.35
44.54 ± 1.04
43.74 ± 0.56
37.28 ± 2.14
35.32 ± 2.69
BOLP
40.57 ± 0.93
35.96 ± 2.17
41.38 ± 2.26
39.71 ± 0.63
32.43 ± 3.10
32.82 ± 1.91
DBR
43.58 ± 1.24
40.52 ± 1.05
44.03 ± 1.56
43.78 ± 1.23
38.49 ± 3.13
36.87 ± 2.85
GroupDRO
42.77 ± 0.33
40.27 ± 0.76
41.53 ± 0.08
43.44 ± 0.56
37.16 ± 0.26
40.30 ± 0.24
SatMAE
ERM
40.89 ± 0.81
34.89 ± 0.55
35.86 ± 0.35
37.42 ± 0.50
30.78 ± 0.57
34.11 ± 1.08
BOLP
39.91 ± 1.31
33.13 ± 2.08
35.29 ± 1.23
38.09 ± 0.93
30.27 ± 0.14
33.76 ± 0.58
Table 8 : Biome-wise performance for MMEarth20K (mIoU) across foundation models and debiasing methods.
Figure S1 : Global map of the six biome macro-groups used as primary evaluation strata in FairRSFM. Raw 14-class terrestrial biomes are pooled into spectral, phenological, hydrological, cryospheric, and surface-property regimes.
Forested regions with strong seasonal or phenological variation.
3
Transitional Herbaceous and Scrub
7 Tropical and Subtropical Grasslands, Savannas and Shrublands; 8 Temperate Grasslands, Savannas and Shrublands; 12 Mediterranean Forests, Woodlands and Scrub
Open or mixed vegetation regimes with strong grass, shrub, soil, and canopy heterogeneity.
4
Cryospheric and Short-Cycle
10 Montane Grasslands and Shrublands; 11 Tundra
Temperature-restricted ecosystems with short growing periods or cryospheric influence.
5
Xeric and Mineralogical
13 Deserts and Xeric Shrublands
Vegetation-sparse surfaces dominated by albedo, exposed soil, and mineral background.
6
Hydrologically Modulated
9 Flooded Grasslands and Savannas; 14 Mangroves
Water-influenced ecosystems where inundation strongly affects spectral response.
Table S1 : Mapping from 14 terrestrial biome classes to six biome macro-groups used for ecological robustness evaluation and mitigation in FairRSFM.
Biome Macro-Group
Biome IDs
Quota / Biome
Total (%)
1. Aseasonal High-Biomass
1, 3, 5
≈ 1,867
5,600 (28.0%)
2. High-Amplitude Phenol.
2, 4, 6
1,700
5,100 (25.5%)
3. Transitional Herbaceous
7, 8, 12
1,200
3,600 (18.0%)
4. Cryospheric & Short-Cycle
10, 11
1,400
2,800 (14.0%)
5. Xeric & Mineralogical
13
1,700
1,700 (8.5%)
6. Hydrologically Modulated
9, 14
600
1,200 (6.0%)
Table S2 : MMEarth20K sample allocation across the 6 macro-groups and 14 terrestrial biomes ( N=20,000 ). Biome IDs correspond directly to Table S1 .
Hyperparams
m-EuroSAT (Macro-F1)
MMEarth20K (mIoU)
α
Tupd
Overall ↑
Worst ↑
NFR ↓
Overall ↑
Worst ↑
NFR ↓
0.2
1
91.73 ± 0.84
83.89 ± 0.62
7.18 ± 0.85
48.02 ± 0.28
33.91 ± 0.51
27.84 ± 0.72
0.4
1
91.52 ± 0.92
84.26 ± 0.48
6.13 ± 0.68
47.91 ± 0.24
34.53 ± 0.46
26.68 ± 0.65
0.5
1
91.36 ± 1.01
84.35 ± 0.50
5.85 ± 0.70
47.84 ± 0.22
34.76 ± 0.42
26.42 ± 0.62
0.6
1
90.82 ± 1.15
83.74 ± 0.64
6.48 ± 0.82
47.36 ± 0.34
34.18 ± 0.48
27.02 ± 0.71
0.8
1
88.94 ± 1.95
82.38 ± 1.62
8.64 ± 1.85
46.12 ± 0.94
33.15 ± 1.12
29.48 ± 1.76
Table S3 : DBR sensitivity sweep across feedback strength α and update frequency Tupd on Prithvi-EO-2.0 test splits ( mean±std across S=3 random seeds).
Dataset
Rank ( k )
Overall ↑
Worst ↑
NFR ↓
m-EuroSAT ( ∣Gd∣=3 )
k=0 (ERM)
90.98 ± 2.36
83.72 ± 0.87
5.96 ± 3.28
k=1
92.47 ± 0.48
84.13 ± 0.18
8.14 ± 0.42
k=2 (Canonical)
93.42 ± 0.35
84.34 ± 0.00
9.82 ± 0.30
m-BigEarthNet ( ∣Gd∣=3 )
k=0 (ERM)
60.09 ± 0.99
46.12 ± 1.51
31.14 ± 1.99
k=1
61.24 ± 0.34
48.16 ± 0.85
24.53 ± 0.61
k=2 (Canonical)
62.75 ± 0.16
50.27 ± 0.63
20.92 ± 0.39
Table S4 : BOLP projection rank ( k ) ablation across datasets on Prithvi-EO-2.0 ( mean±std across S=3 seeds; k=0 denotes unprojected ERM).