cs.CVMay 5, 2026

Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation

Authors: Quanxing XuLing ZhouXian ZhongXiaohua HuangRubing HuangChia-Wen Lin

Organizations: School of Computer Science and Engineering, Macau University of Science and Technology, Macao SAR 999078, China · Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan, Hubei 430070, China · State Key Laboratory of Maritime Technology and Safety, Wuhan University of Technology, Wuhan 430063, China · Oulu School, Nanjing Institute of Technology, Nanjing 210096, China · Zhuhai MUST Science and Technology Research Institute, Macau University of Science and Technology, Zhuhai, Guangdong 519099, China · Department of Electrical Engineering, National Tsing Hua University, Hsinchu 30013, Taiwan

Abstract

With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of multimodal tasks. As a core problem in vision-language research, Visual Question Answering (VQA) has increasingly employed MLLMs to improve performance, particularly in open-domain settings where external knowledge is essential. In this work, we aim to further enhance retrieval-based VQA by more effectively integrating MLLMs with structured reasoning and knowledge acquisition. We introduce a logical prompting strategy that fuses Chain-of-Thought (CoT) reasoning with Visual Question Decomposition (VQD), termed CoVQD, to guide retrieval toward more accurate and relevant knowledge for MLLM inference. Building on this idea, we propose a new framework, CoVQD-guided RAG (CgRAG), which enables MLLMs to access more comprehensive and coherent external knowledge while benefiting from structured visual-text reasoning guidance, thereby improving generalization and reliability in complex cross-domain VQA scenarios. Extensive experiments on E-VQA, InfoSeek, and OKVQA benchmarks demonstrate the effectiveness of the proposed method.

Explore similar work

CardsList