cs.IRSep 2, 2026

ViSAR: Training-Free Adaptive-kk Retrieval for Visual Document Question Answering

Authors: Adrien Mialland, Marc Plantevit, Julien Gallois, Céline Robardet

Organizations: INSA Lyon, CNRS, LIRIS UMR 5205, F-69621 Villeurbanne, France · EPITA Research Laboratory (LRE), FR-94276, Le Kremlin-Bicêtre, France · Lowit, FR-69003 Lyon, France

Abstract

Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top-kk number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-kk retrieval method for late-interaction visual document retrieval. ViSAR operates directly in the embedding space to construct a query-conditioned page-level similarity matrix that highlights query-relevant semantics and dynamically determines the number of pages to retrieve. Across multiple encoders and LVLMs, ViSAR retrieves compact, query-adapted page sets that reduce RAG latency by up to 58.7%, while maintaining or improving answer accuracy compared with fixed top-kk and adaptive retrieval heuristics. Furthermore, we show that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.