While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability and expert interaction, they are either too expensive to train for the remote sensing domain or perform poorly without annotation. We posit that in expert domains like remote sensing, such training-free models require both fine details in both image and concept space. In image space, we propose a multiscale concept bottleneck using greedy quadtree routing to locate small concepts. In concept space, we replace contrastive vision language models with pre-trained MLLMs and present a way to get reliable concept scores from them. We introduce APERTURE that blends concept scores at the global image and native concept-scale level to give state-ofthe-art training-free model performance. To test these models, introduce SiFC, a fine-grained concept-centric dataset across three countries, with human-reviewed class-level concept maps. On SiFC, APERTURE outperforms the best training-free baselines by more than 10 percentage points in macro F1-score, and notably also outperforms supervised concept bottleneck models. Targeted component-removal tests examine whether concept scores respond to changes in visual evidence, while temporal experiments show that descriptor updates improve recognition of technological changes without retraining.
Figures & tables
Figure 1: Multi-scale concept bottleneck: Resizing the 4096 -pixel coal plant (spanning 4km2 ) shrinks the coal yard and chimney to 112 and 14 pixels (top right). Aperture instead crops concepts at home scales, enabling it to recognize both parts and correctly classify the site, unlike CLIP.
Figure 2: SiFC at a glance. (Left) sample facility classes, labeled with their country and the number of sites. Each tile covers about 4 km2 (4096 × 4096 px at zoom level 18, World Imagery Wayback ( Esri, ) ); the white bar is 500 m. (Right) 800 sites by country (inner) and classes (outer).
Figure 3: Overview of Aperture . For each class- and country-conditioned descriptor dk∣y,r , a frozen MLLM follows one quadtree branch to its home scale hk while pruning the remaining quadrants. Global and home-scale answer scores are blended into the concept score ckB . The concept scores are then averaged into the class score Sy , and the highest-scoring class is predicted.
Table 1: Macro-F1 (%; ↑ ) on SiFC and the fMoW facility subset. The highest and second-highest scores in each column across all rows are bold and underlined , respectively. All Gemma settings use the same frozen MLLM. TF denotes no task-specific training.
Figure 5
Table 3: Cumulative part-removal interventions. Mean true-class score drop from the intact image over 15 images per class. S1–S3 cumulatively remove one, two, and three parts, respectively (listed below). The framed block averages six classes; bars end at S3, with cuts marking S1 and S2. Bold : largest drop per column; blue : our settings.
Table 7
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
SiFC (ours)
fMoW
Model
Setting
Descriptor
Score
TF
India
USA
China
All
USA
France
Russia
All
Baseline
CLIP zero-shot
CND
cos.
✓
53.03
67.19
53.29
62.16
71.29
57.71
51.91
65.85
CBD
CBD
cos.
✓
59.37
66.60
57.73
66.42
70.57
54.11
51.02
62.74
LaZSL
AD
OT
✓
69.38
72.83
58.78
69.76
78.63
59.86
57.17
69.90
CF-CBM (high)
HCD
cos.
✗
40.25
25.67
48.60
39.35
12.60
25.89
24.11
23.18
CF-CBM (low)
HCD
cos.
✗
53.92
40.87
60.69
49.88
41.30
72.52
58.17
55.92
Appendix
Table 7: Macro F1 (%) on SiFC (six classes, 800 images; India, USA, China) and on the fMoW subset (USA, France, Russia). All rows share one MLLM and are training-free (TF ✓) except the CBM rows, which fit a linear head with leave-one-country-out folds (✗). Within each model block, bold is the best and underline the second best per column. The row in blue is our full method.
Class
USA
France
Russia
Total
Shipyard
15
6
5
26
Solar farm
24
20
6
50
Storage tank
74
66
47
187
Water treatment facility
26
43
27
96
Wind farm
22
14
5
41
Total
161
149
90
400
Appendix
Table 8: Proposed class and country distribution of the fMoW-FC facility subset.
Facility class
India
USA
China
Total
Iron and steel plant
48
28
104
180
LNG terminal
8
13
28
49
Nuclear power plant
7
52
15
74
Oil refinery
24
70
65
159
Power plant
101
42
37
180
Sewage treatment plant
22
90
46
158
Appendix
Table 9: SiFC facility sites by class and country.
Class
India
USA
China
Iron and steel
PIB
EPA
MEE
LNG terminals
PNGRB
FERC
MEE
Nuclear plants
NPCIL
EIA-860
NNSA
Oil refineries
PPAC
EIA-820
MEE
Power plants
CEA
EIA-860
MEE
Sewage treatment
CPCB
EPA ECHO
MEE
Appendix
Table 10: Government references for SiFC facility identity and locality. individual verification.
Top-1
Beam-2
Beam-3
Beam-4
Region
Home
Blend
Home
Blend
Home
Blend
Home
Blend
India
41.74
72.37
49.13
71.26
53.04
71.55
52.11
70.41
USA
66.10
75.54
73.13
77.91
76.40
78.27
74.00
77.94
China
45.56
79.62
53.17
80.48
53.31
79.01
53.84
79.67
All
55.09
80.19
61.58
80.42
64.13
80.11
63.10
80.12
Appendix
Table 11: Search width on SiFC. Macro-F1 (%); All : pooled.
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a localization tool for segmentation and grounding. To train this heterogeneous behaviour, we employ a multi-task reinforcement learning framework with adaptive task rewards covering multiple-choice VQA, free-form VQA, captioning, detection, and segmentation across a large variety of input types. Our approach achieves competitive results across a broad set of benchmarks, including high-resolution, multi-temporal, multi-modal and multi-view tasks. Further, as training data scales, our experiments show consistent improvements across most tasks both in and out of distribution, which correlate with per-task data diversity. These findings suggest that, for remote sensing VLMs, data scale is sufficient even without architectural novelty.
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2
A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Remote Sensing MLLMs fail to meet this requirement. Trained on carefully curated clean datasets, they learn brittle mappings that do not generalize to noisy conditions in operational Earth Observation. Consequently, their performance degrades when confronted with imperfect inputs in deployment. To quantify this vulnerability, we construct a realistic set of multimodal perturbations, including visual degradations such as cloud and fog cover, together with diverse human-centric textual variations ranging from colloquialisms to vague or omitted instructions. Empirical evaluations show that these perturbations significantly impair the visual-semantic reasoning capabilities of leading RS foundation models. To address this limitation, we introduce RemoteShield, a robust Remote Sensing MLLM trained to maintain consistent outputs across realistic input variations. During training, each clean sample is paired with its image-text perturbed variants to form a semantic equivalence cluster. Rather than directly fitting noisy samples, RemoteShield is optimized through preference learning over clean and perturbed conditions within the same cluster. By comparing model responses to clean and corrupted inputs, the model is encouraged to favor stable responses over perturbation-induced failures. This cross-condition alignment helps the model focus on underlying task semantics despite visual degradations and textual noise. Experiments on three Earth Observation tasks show that RemoteShield consistently delivers stronger robustness and cross-condition consistency than representative baselines under realistic multimodal perturbations.
Rui Min, Liang Yao, Shiyu Miao +5
1Hohai University · 2Nanjing University · 3Southeast University
Concept Bottleneck Models (CBMs) provide an intrinsically interpretable alternative to post-hoc explanations. However, existing CBMs often rely on predefined concept vocabularies or supervised annotations, lack explicit concept grounding, and summarize each concept with a single image-level score -- discarding spatial recurrence and inter-concept dependencies. We propose a Graph-based Concept Bottleneck Model (G-CBM), an intrinsically interpretable framework that performs unsupervised concept discovery via Non-negative Matrix Factorization (NMF) and represents the discovered concepts as nodes in a per-image concept-graph representation. G-CBM matches region-level features to these concept nodes -- providing concept grounding and capturing concept recurrence across the image -- and applies a \emph{tunable concept filtering threshold} τ to suppress weak region-level features. A Graph Attention Network (GAT) then performs concept-level reasoning by modeling nonlinear dependencies across nodes. Across ImageNet, HAM10000, PH2, and Derm7pt, G-CBM achieves an average relative AUC improvement of 3.7% over a ResNet-50 baseline. Concept filtering frequently improves predictive performance while inducing selective concept use, achieving peak AUC of 0.96 on PH2 with only 2 of 10 concepts and 0.92 on HAM10000 with 3.8 of 9 concepts. On dermoscopy benchmarks, G-CBM is competitive with supervised approaches requiring external annotations. Deletion/insertion analyses with random ablation controls show that the learned concept ranking faithfully reflects model predictions.
Md Mohasin Hossain, Anar Amirli, Robert Leist +2
German Research Center for Artificial Intelligence · Saarland University, Saarbrücken, Germany · 1German Research Center for Artificial Intelligence (DFKI), Saarbrücken, Germany +5