cs.CVMay 8, 2026

LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

Authors: Jun WangFengpeng LiHang DongTianjin HuangWei Han

Organizations: School of Computer Science, China University of Geosciences · PRADA Lab, King Abdullah University of Science and Technology · Department of Computer Science, University of Exeter

Abstract

Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general land-cover recognition, lithology interpretation is a knowledge-intensive task that requires experts to infer rock types from various features, e.g., subtle visual, spectral, textural, geomorphological, and contextual cues, making reliable automated interpretation highly challenging. Geological knowledge-guided large multimodal models offer new opportunities, yet their evaluation remains constrained by the lack of benchmarks that capture lithological annotations, multi-level geological semantics, and expert-informed assessment. Here, we propose LithoBench, a multi-level benchmark for evaluating geological semantic understanding in remote sensing lithology interpretation. LithoBench contains 10,000 expert-annotated interpretation instances across 12 representative lithological categories, including 4,000 multiple-choice and 6,000 open-ended tasks organized into five cognitive levels: Identification and Description, Comparative Analysis, Mechanism Explanation, Practical Application, and Comprehensive Reasoning. We further develop an expert-in-the-loop, knowledge-grounded semi-automated construction pipeline, coupling multi sub-processes, e.g., structured geological image descriptions, to enhance geological validity and evaluation reliability. Experiments with multiple large vision-language models eveal substantial limitations in geological semantic understanding, particularly on higher-order explanation, application, and reasoning tasks.

Explore similar work

May 5, 2026cs.AI

GeoDecider: A Coarse-to-Fine Agentic Workflow for Explainable Lithology Classification

Lithology classification aims to infer subsurface rock types from well-logging signals, supporting downstream applications like reservoir characterization. Despite substantial progress, most existing methods still treat lithology classification as a single-pass classification task. In contrast, practical experts incorporate geological principles, external knowledge, and tool-use capabilities to perform accurate classification. In this work, we propose GeoDecider, a coarse-to-fine agentic workflow that enables accurate and explainable lithology classification through training-free use of large language models (LLMs). GeoDecider reformulates lithology classification as an expert-like structured process and organizes it into a multi-stage workflow involving coarse-to-fine reasoning. Specifically, GeoDecider includes the following stages: (1) base classifier-guided coarse classification, which uses a pre-trained classifier to provide a rough reference for downstream tasks, thus reducing the overall cost of downstream reasoning, (2) tool-augmented reasoning, which utilizes several tools such as contextual analysis and neighbor retrieval to achieve finer and more precise classifications, (3) geological refinement, which post-processes the final results to enforce geological consistency. Experiments on four benchmarks show that GeoDecider outperforms representative baselines. Further analysis demonstrates that the proposed framework produces geologically interpretable predictions while achieving a better trade-off between classification performance and inference efficiency.
Jiahao Wang, Mingyue Cheng, Yitong Zhou +4
May 25, 2026cs.CV

Beyond Appearance: Can Multimodal Large Language Models Exploit Vertical Structure for Remote Sensing Natural Scene Understanding?

Multimodal large language models (MLLMs) have advanced rapidly in remote-sensing analysis, yet existing evaluations remain predominantly 2D-centric. Because spectrally confused regions can appear nearly identical yet differ substantially in vertical structure, appearance alone is often insufficient for reliable semantic interpretation in natural scenes. Vertical structure therefore provides decision-critical physical evidence, yet whether current MLLMs can effectively perceive, ground, and utilize such geometric evidence remains underexplored. To bridge this gap, we introduce VertiCue-Bench, the first diagnostic benchmark that uses controlled interventions to probe whether vertical height evidence is actually perceived, grounded, and utilized, and we establish a three-stage evidence-utilization framework of Perception--Grounding--Utilization. By constructing a Representation Intervention Spectrum spanning multiple presentation and interaction modalities, including Raw Visual, Tool-assisted, and Oracle Text conditions, together with controlled counterfactual tests, we conduct an in-depth disentangled diagnosis across 10 state-of-the-art models. Our experiments reveal and formally characterize the Vertical Structure Utilization Gap. Although current models exhibit emerging geometric perception capabilities, they still struggle to accurately ground vertical evidence to relevant spatial entities and integrate it into high-level semantic decisions. This finding identifies a critical bottleneck in developing physically grounded and geometry-aware remote-sensing MLLMs.
Jing Huang, Duanchu Wang, Junjie Yang +5
May 12, 2026cs.CV

GeoR-Bench: Evaluating Geoscience Visual Reasoning

Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, climate adaptation and environmental protection. Although current research has shown promising progress on specific geoscience tasks, such as remote sensing interpretation, geographic question-answering, existing benchmarks remain largely task-specific which failing to capture the open-ended real world geoscience problems. As a result, it remains unclear how far current AI systems are from achieving genuine geoscience intelligence. To address this gap, we present \textbf{GeoR-Bench}, a \underline{Bench}mark for evaluating \underline{Geo}science visual \underline{R}easoning through reasoning informed visual editing tasks. GeoR-Bench contains 440 curated samples spanning 6 geoscience categories and 24 task types, covering earth observation imagery and structured scientific representations such as maps and diagrams. We evaluate outputs along three dimensions, including reasoning, consistency, and quality. Benchmark results of 21 closed- and open-source multimodal models reveal that geoscience reasoning remains a critical bottleneck. The highest-performing model achieves 42.7% overall strict accuracy, while the best open-source models only get 10.3%. Notably, the visual consistency and image quality of the outputs frequently surpass their scientific accuracy. Ultimately, these findings indicate that current models generate superficially plausible results but fail to capture underlying earth science processes.
Yushuo Zheng, Zicheng Zhang, Huiyu Duan +7