cs.CVSep 27, 2026

Seeing and Solving Are Not Enough for Vision-Language Models

Authors: Ziheng Wang, Mingxuan Xie, Yilin Liu, Dayan Wu, Yang Li, Pengwen Dai

Organizations: Sun Yat-sen University · Zhejiang University · The Hong Kong University of Science and Technology · Institute of Information Engineering · Hunan University

Abstract

Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

    Sep 28, 2026Rong Yu Xu, Prayag Tiwari, Shaolei ZhangRecent Vision-Language ModelsVisual Reasoning

  2. Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models

    Jul 6, 2026Khang Nhat Hoang Vo, Artem Vazhentsev, Artem Shelmanov +2Uncertainty Quantification

  3. Composition, Not Conversation: VLMs Lose the Scene, Not the Thread

    Sep 29, 2026L. D. M. S. Sai Teja, Ufaq Khan, N. Siva Gopala Krishna +5TextvqaRecent Vision-Language Models