cs.CVSep 11, 2026

Evaluation of Vision-Language Models Across Diverse Coastal Environments

Authors: Seth KnoopChad R. SamuelsonGabriel R. SladeBrady MoonJoshua G. Mangelson

Abstract

Vision-language models (VLMs) enable robotic per- ception by associating visual observations with natural-language concepts. Yet their performance in coastal environments remains largely unexplored. We introduce a densely labeled coastal dataset containing more than 1,000 images collected across seven missions in three regions of Oahu, Hawaii, with 18 semantic classes and over 7,400 annotated instances. We evaluate seven modern VLMs through three complementary experiments mea- suring text-to-mask, mask-to-mask, and mask-to-text alignment. Broad landscape classes are generally recognized more accurately than conventional object and coastal classes, with coastal con- cepts presenting the greatest challenge. However, comparisons of shared conventional classes across coastal and terrestrial datasets reveal no consistent performance difference attributable solely to environmental context. Mask-to-mask matching also remains similar across conventional and coastal classes, while alternative textual labels substantially improve recognition of several coastal concepts. These results suggest that lower performance on coastal classes (at least on the objects/query categories evaluated) is heavily influenced by segmentation and linguistic representation.

Explore similar work

Date pendingcs.CV

Geometric Coastline Localization using Vision-Language Models

Coastline detection in remotely sensed imagery is commonly formulated as pixel-wise segmentation, even though coastlines used in coastal monitoring are ultimately represented as geometric curves and defined by geomorphic proxies such as vegetation lines, dune toes, or cliff edges. We revisit coastline extraction from a representation perspective and formulate the task as geometric boundary localization, where a thin coastline boundary is localized directly as a curve rather than derived from a segmentation mask. Using the New Zealand Coastal Change Dataset (NZCCD) and LINZ aerial imagery, we develop CoastlineVLM-7B, a vision-language model built on the GeoChat-7B/LLaVA-1.5 architecture that jointly performs coastline presence detection, proxy-type classification, and direct coastline grounding as an ordered polyline. We compare it against U-Net, UNet++, DeepLabV3+, and SegFormer trained using one-pixel-wide coastline masks and evaluate localization using tolerance- and distance-based geometric metrics. On the West Coast test set, U-Net provides stronger local boundary proximity, showing better tolerance scores and lower Chamfer and Modified Average Hausdorff distances, while CoastlineVLM-7B achieves the lowest Hausdorff and Earth Mover's distances, indicating reduced worst-case deviation and stronger global structural correspondence. Ablation studies show that geometric localization is driven primarily by direct coastline-grounding supervision and that GeoChat initialization provides a stronger starting representation than generic LLaVA-1.5 initialization. Zero-shot evaluation on the Australian VCMP dataset shows reduced performance for U-Net and CoastlineVLM-7B, but CoastlineVLM-7B achieves better geometric localization than U-Net. These results show that direct ordered-polyline grounding can be used as an alternative formulation for representing and localizing coastline geometry.
Rafia Malik, Bernhard Pfahringer, Karin Bryan +2
May 15, 2026cs.CV

Neutral-Reference Prompting for Vision-Language Models

Efficient transfer learning of vision-language models (VLMs) commonly suffers from a Base-New Trade-off (BNT): improving performance on unseen (new) classes often degrades accuracy on known (base) classes. Addressing how to boost recognition of unseen classes without sacrificing known-class performance remains a central challenge. Existing work often simplistically attributes the BNT to overfitting on known classes. We observe an interesting phenomenon: VLMs frequently exhibit asymmetric confusion on certain downstream data, i.e., samples of class A are systematically mispredicted as class B, while the reverse confusion (B to A) rarely occurs. For known classes, this kind of bias can be mitigated by tuning using a cross-entropy loss, but for unseen classes, such pretraining-induced bias persists and harms generalization. Motivated by this, we propose NeRP, a plug-and-play prompting correction strategy that improves discrimination on unseen classes without modifying model parameters. NeRP leverages neutral text prompts and reference images to measure class-wise prior preferences along the pre-trained inter-class geometry, and combines them with the sample likelihood to obtain the model's surrogate score. If, for a given sample, the prior strongly favors the current prediction while the observed evidence is clearly insufficient, we perform a local flip between easily confusable class pairs, thereby correcting prior-dominated mispredictions. Extensive experiments across multiple backbones and 15 few-shot and cross-domain benchmarks show that NeRP substantially improves accuracy on unseen classes while preserving known-class prediction performance.
Senmao Tian, Xiang Wei, Shunli Zhang
Sep 17, 2026cs.CV

Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the question affects VLMs in two opposite ways. Verbose questions make VLMs substantially more robust---e.g., rephrasing "Is there a cat?" into "Please look carefully and answer: is there a cat?". Conversely, VLMs become more fragile under corruption when the question is semantically complex or finer-grained, e.g., "what colour is the cup left of the chair?" instead of "is there a cup?". Both effects stem from question-conditioned cross-modal attention, which induces a spectral filter over image patches: verbose questions broaden its frequency support, while fine-grained questions concentrate it onto fewer visual scales. The model's answer drifts most when this filter and the corruption sit on the same spatial frequencies. We test the filter view on Qwen3-VL and LLaVA-OneVision across GQA and CLEVR; verbose paraphrasing reduces drift variance by 70--81% on the 8B models. The practical recipe---pad the prompt---further yields measurable gains in accuracy, even under image corruption.
Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza +5