cs.CVAug 24, 2026

Investigating Relational Reasoning in VLMs

Authors: Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap

Abstract

Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we use the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths. For this, we propose a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues. Furthermore, the dataset is modified to test causal reliance on visual evidence. Our results show that current VLMs combine genuine visual reasoning with shortcut strategies primarily rooted in language cues.

Explore similar work

CardsList
  1. Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

    Oct 6, 2026Gouki Minegishi, Hiroki Furuta, Takeshi Kojima +2

  2. What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning

    Jun 1, 2025Zhaotian Weng, Haoxuan Li, Xin Eric Wang +2VLM EvaluationVLM Reasoning

  3. See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection

    Apr 27, 2026Zhiheng Wu, Tong Wang, Shuning Wang +2Tool Use in VLMsVLM Reasoning