RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models
Organizations: School of Computer Science, Wuhan University, Wuhan 430072, China · State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China
Abstract
Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradigms: task-specific visual classification, vision-language similarity matching, and autoregressive multimodal generation. However, visual classifiers rely on predefined label spaces, CLIP-based methods perform recognition through static image-text alignment, and multimodal large language models (MLLMs) introduce unnecessary token-level generation for classification tasks with explicit candidate categories. To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification. Unlike conventional MLLMs that formulate classification as autoregressive text generation, RSJEV reformulates scene classification as a candidate-conditioned multimodal discriminative decision process, where visual representations, task instructions, and candidate category semantics are jointly modeled. Specifically, we introduce a OnePass Decider that extracts multimodal decision states and directly estimates category probabilities within the candidate category space, eliminating autoregressive decoding while preserving vision-language interactions. Extensive experiments on three widely used remote sensing scene classification benchmarks, including UC Merced, AID, and NWPU-RESISC45, demonstrate that RSJEV achieves superior classification performance compared with representative CNN-, Transformer-, Mamba-, CLIP-, and MLLM-based methods. Moreover, RSJEV significantly reduces inference costs and achieves a better accuracy-efficiency trade-off with only a compact 0.8B-parameter model. These results demonstrate the effectiveness of state-conditioned multimodal decision making for efficient remote sensing image understanding. The code will be available at https://github.com/Dongtcs/RSJEV.
Figures & tables
| Method | Param | UCM | AID | NWPU | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | OA | P | R | F1 | OA | P | R | F1 | OA | ||
| CNN-based Models | |||||||||||||
| ResNet-101 [ 7 ] | 42.6 M | 95.91 | 95.71 | 95.70 | 95.71 | 92.33 | 91.79 | 91.89 | 92.23 | 90.34 | 90.34 | 90.25 | 90.34 |
| DenseNet-161 [ 8 ] | 28.7 M | 97.18 | 97.05 | 97.02 | 97.05 | 93.90 | 93.48 | 93.57 | 93.90 | 92.77 | 92.73 | 92.69 | 92.73 |
| EfficientNet [ 18 ] | 87.4 M | 97.12 | 96.76 | 96.80 | 96.76 | 91.48 | 91.39 | 91.33 | 91.76 | 92.49 | 92.45 | 92.43 | 92.45 |
| Transformer-based Models | |||||||||||||
| Backbone | Method | UCM | AID | NWPU |
|---|---|---|---|---|
| Qwen3.5-0.8B | Baseline | 87.96 | 91.29 | 88.24 |
| Ours | 97.80 | 96.44 | 94.61 | |
| InternVL3.5-1B | Baseline | 93.21 | 92.26 | 93.54 |
| Ours | 96.68 | 96.27 | 94.93 |