cs.CVSep 29, 2026

TAEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG

Authors: Yalun Wu, Bingzhou Wang, Boyang Wang, Peiying Wang, Shaojie He, Yunhan Wang, Shaozu Yuan, Jiawei Wang

Organizations: NExT++ Lab, National University of Singapore

Abstract

Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term this loss of usable evidence over a reasoning trajectory trajectory-level evidence utilization degradation. To address it, we propose Trajectory-Aware Evidence Coordination (TAEC), a training-free framework that coordinates evidence use around unresolved answer requirements. TAEC tracks these requirements in a shared trajectory state to guide which evidence enters the context, how accumulated memory is retained, and at what level of detail visual evidence is examined. Under a unified evaluation protocol on ViDoSeek, SlideVQA, and MMLongBench-Doc, TAEC achieves the best overall performance against leading training-free visual RAG baselines, with the highest average accuracy across multiple proprietary vision-language models. These results demonstrate that aligning evidence with evolving reasoning needs improves evidence use throughout multi-step visual RAG.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation

    Sep 14, 2026Yucheng Shen, Lingyong Yan, Jiulong Wu +4Visual EvidenceKnowledge-Based Visual Question Answering

  2. Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation

    May 13, 2026Weiqing Luo, Zongye Hu, Xiao Wang +3Multimodal Retrieval Augmented GenerationVisual Evidence

  3. Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation

    May 2, 2026Peiyang Liu, Ziqiang Cui, Xi Wang +2Multimodal Retrieval Augmented GenerationKnowledge-Based Visual Question Answering