Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term this loss of usable evidence over a reasoning trajectory trajectory-level evidence utilization degradation. To address it, we propose Trajectory-Aware Evidence Coordination (TAEC), a training-free framework that coordinates evidence use around unresolved answer requirements. TAEC tracks these requirements in a shared trajectory state to guide which evidence enters the context, how accumulated memory is retained, and at what level of detail visual evidence is examined. Under a unified evaluation protocol on ViDoSeek, SlideVQA, and MMLongBench-Doc, TAEC achieves the best overall performance against leading training-free visual RAG baselines, with the highest average accuracy across multiple proprietary vision-language models. These results demonstrate that aligning evidence with evolving reasoning needs improves evidence use throughout multi-step visual RAG.
Figures & tables
Figure 1: Trajectory-level evidence utilization degradation in multi-step visual RAG.
Figure 2: Overview of Trajectory-Aware Evidence Coordination. At each step, the shared requirement state guides evidence admission, memory exposure, and visual detail allocation.
Figure 3: ReAct on ViDoSeek with Gemini-3.5-Flash, grouped by the number of searches. Marker area is proportional to the number of questions in each group.
ViDoSeek
SlideVQA
MMLongBench-Doc
System
Gemini
Kimi
GPT-5.6
4o-mini
Gemini
Kimi
GPT-5.6
4o-mini
Gemini
Kimi
GPT-5.6
4o-mini
Avg.
Vanilla
76.7
81.8
82.0
59.2
79.6
78.0
79.3
61.9
40.1
31.8
38.1
15.6
60.3
ReAct
79.5
82.1
82.1
56.1
77.2
82.0
81.0
59.3
42.4
36.5
39.1
15.2
61.0
ViDoRAG
79.7
80.8
81.3
65.8
72.5
74.5
78.3
58.7
31.5
37.2
37.8
14.1
59.4
DAG agent
80.1
84.8
84.1
65.7
76.7
80.9
77.7
56.6
44.5
43.9
38.3
14.3
62.3
TAEC
87.5
88.1
87.0
71.0
83.6
83.8
82.2
62.0
50.1
45.8
42.6
17.7
66.8
Table 1: Accuracy (%). Bold and underlining mark the best and second-best results per column; Avg. is the mean across twelve settings. MMLongBench-Doc uses 847 answerable questions; full-set results appear in Appendix E.1 . Supplementary results for M3RAG, RL-trained agents, and open-weight backbones appear in Appendices D , J.3 , and J , respectively.
Figure 4: Component ablation with Gemini-3.5-Flash. Values show accuracy (%); colour indicates gains over the DAG agent, normalised by TAEC’s gain on each benchmark.
Figure 5: Analysis with Gemini-3.5-Flash (Table 1 ). (a) Accuracy by ReAct search count (Table 3 ). (b) TAEC’s accuracy gain over the DAG agent, decomposed by coverage group on questions with multiple ReAct searches. Centres show net gains (percentage points); hollow sectors indicate negative contributions. (c) Mean admitted and distinct pages per question. MMLongBench-Doc coverage uses answerable questions with annotated evidence pages.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Two appendix studies whose numbers are in the tables named below. (a) Closed-source backbones, ReAct → TAEC with DAG agent marked, sorted by ReAct accuracy (Table 1 ); the acting models are GPT-4o-mini, Gemini-3.5-Flash, Kimi-K3 and GPT-5.6-Sol. (b) Matched subsets as a flow: the same 1,142 ViDoSeek questions in the centre, grouped by the behaviour of the two systems (A: both stopped after one search; B: ReAct kept searching while TAEC stopped; C: the remainder), with ribbons carrying each group into “gold page found” or “not found” under ReAct on the left and TAEC on the right.
All questions
Gold page found at step 1
Searches
n
Hit
Post-hit
Acc.
Found, unused
n
Acc.
1
642
93.8
96.0
93.6
3.7
602
96.0
2
126
95.2
91.7
89.7
7.9
117
91.5
3
71
80.3
87.7
77.5
9.9
37
83.8
4–5
84
75.0
87.3
67.9
9.5
24
91.7
6–8
92
63.0
63.8
41.3
22.8
33
48.5
Appendix
Table 2: ReAct, unconstrained, ViDoSeek, Gemini-3.5-Flash, grouped by the number of searches the agent issued. “Found, unused” is the share of questions whose gold page reached the context and whose answer is still wrong. Right: the subset whose gold page was already among the first five retrieved pages.
Searches needed
n
Vanilla
ReAct
DAG agent
TAEC
1
642
86.7
93.6
91.9
94.1
2–3
197
78.7
85.3
79.7
86.3
4–5
84
69.0
67.9
63.1
81.0
6+
219
48.4
37.9
52.5
71.7
Appendix
Table 3: ViDoSeek accuracy by question difficulty, Gemini-3.5-Flash. Difficulty is the number of searches unconstrained ReAct issued on that question, so all four systems face the same questions in each row.
What became of the gold page (%)
Benchmark
System
Searches
Coverage
Found, used
Found, unused
Missed, right
Missed, wrong
ViDoSeek
Vanilla
1.00
75.0
58.8
16.2
5.0
20.0
ReAct
6.00
74.6
58.2
16.4
3.4
22.0
ViDoRAG
—
76.0
61.0
15.0
4.4
19.6
Pure DAG
1.27
68.4
57.6
10.8
7.4
24.2
TAEC
1.28
90.8
76.6
14.2
2.4
6.8
Appendix
Table 4: The questions on which ReAct kept searching, Gemini-3.5-Flash. The subset is defined by that agent alone, so all five systems answer the same questions: 500 on ViDoSeek, 765 on SlideVQA and 541 on MMLongBench-Doc. Coverage and the four shares are taken over the questions in the subset that carry a reference page, which on MMLongBench-Doc is 532 of the 541. Coverage is the sum of the first two shares. ViDoRAG records no search count, and its coverage is taken over the pages its pipeline records for the question.
System
Hit rate
Post-hit acc.
Acc. on misses
Acc.
Vanilla
86.6
84.9
23.5
76.7
ReAct
85.4
89.1
23.4
79.5
DAG agent
84.4
90.1
25.8
80.1
TAEC
94.7
90.6
31.7
87.5
Appendix
Table 5: Accuracy decomposed into coverage and evidence use per system, ViDoSeek, Gemini-3.5-Flash. Post-hit accuracy conditions on a quantity each system changes, so it is reported here for completeness and is not compared across rows; the like-for-like comparison is Figure 9 .
System
Model calls
Images
Context (MB)
Peak call (MB)
Vanilla
1.00
4.9
0.22
0.22
ReAct
5.34
92.8
3.73
0.57
DAG agent
4.94
20.0
0.87
0.23
TAEC
2.91
14.7
0.61
0.22
Appendix
Table 6: Cost per question, ViDoSeek, Gemini-3.5-Flash, same protocol as Table 1 . Cost is what a run actually transmits: images and context are summed over every model call of a question, so a page re-sent on a later call is counted again, and the last column is the largest single call. Wall-clock time is not reported; it is set by concurrency and rate limits rather than by the method.
Figure 7: What a trajectory spends and what the spending buys, ViDoSeek, Gemini-3.5-Flash. (a) The accuracy a system reaches when a question is allowed at most k searches, the numbers of Table 19 ; a ring marks the smallest budget at which a system already holds all the accuracy it will reach. The substrate and TAEC are saturated by eight searches, while the unmanaged agent is still gaining at twenty and reaches neither. (b) Accuracy against the context a question transmits, the numbers of Table 6 , with point area proportional to the images sent. TAEC answers eight points above the unmanaged agent while sending a sixth of the context and a sixth of the images.
Components
ViDoSeek
SlideVQA
MMLongBench-Doc
Configuration
A
M
V
Acc.
Δ
Acc.
Δ
Acc.
Δ
DAG agent (substrate)
80.1
—
76.7
—
44.5
—
+ Admission
✓
86.3
+6.2
82.3
+5.6
48.2
+3.7
+ Exposure
✓
84.1
+4.0
82.8
+6.1
47.2
+2.7
+ Allocation
✓
83.5
+3.4
82.4
+5.7
46.8
+2.3
TAEC (A+M+V)
✓
✓
✓
87.5
+7.4
83.6
+6.9
50.1
+5.6
Appendix
Table 7: Component ablation behind Figure 4 , Gemini-3.5-Flash; the substrate and TAEC rows are those of Table 1 , and Δ is against the substrate.
System
ViDoSeek
SlideVQA
MMLongBench-Doc
Vanilla
76.7
79.6
40.1
ReAct
79.5
77.2
42.4
ViDoRAG
79.7
72.5
31.5
M3RAG (reimplemented)
82.5
81.5
47.6
DAG agent
80.1
76.7
44.5
TAEC
87.5
83.6
50.1
Appendix
Table 8: Gemini-3.5-Flash. The other five rows are those of Table 1 ; best and second best per column.
ViDoSeek
SlideVQA
MMLongBench-Doc
Backbone
M3RAG
TAEC
M3RAG
TAEC
M3RAG
TAEC
Gemini-3.5-Flash
82.5
87.5
81.5
83.6
47.6
50.1
Kimi-K3
80.7
88.1
81.5
83.8
44.5
45.8
GPT-5.6-Sol
78.4
87.0
77.3
82.2
44.3
42.6
GPT-4o-mini
49.6
71.0
44.9
62.0
18.7
17.7
Appendix
Table 9: M3RAG against TAEC on the four acting models of Table 1 , under the same index, budgets and judge, and on the full benchmark in every cell. MMLongBench-Doc is scored on the 847 answerable questions, as in Table 1 .
ViDoSeek
SlideVQA
MMLongBench-Doc
Backbone
System
Single
Multi
All
All
Text
Table
Chart
Figure
All
Gemini-3.5-Flash
Vanilla
71.8
83.1
76.7
79.6
41.0
49.7
55.1
35.2
40.1
ReAct
76.3
83.7
79.5
77.2
42.5
60.1
60.7
29.1
42.4
ViDoRAG
77.2
82.9
79.7
72.5
36.6
34.3
40.2
29.1
31.5
M3RAG †
81.4
83.9
82.5
81.5
50.0
67.8
61.7
33.5
47.6
DAG agent
77.8
83.1
80.1
76.7
41.8
69.2
55.1
34.6
44.5
Appendix
Table 10: Accuracy by question type, same runs and judging batch as Table 1 . ViDoSeek is split by hop count; MMLongBench-Doc by the modality of the evidence, counting the questions annotated with exactly one evidence source, so questions drawing on several modalities and the layout-only ones enter only “All”, which is taken over the 847 answerable questions; SlideVQA carries no question-type annotation. Best per column within each backbone block in bold. † our reimplementation, for the reasons given in Appendix D .
System
Gemini-3.5-Flash
Kimi-K3
GPT-5.6-Sol
GPT-4o-mini
Vanilla
36.8
32.3
31.5
22.5
ReAct
33.3
32.4
33.5
16.9
ViDoRAG
28.0
33.8
34.7
15.6
M3RAG †
40.7
42.0
44.3
24.1
DAG agent
45.7
41.2
39.5
23.8
TAEC
46.9
42.9
40.1
25.2
Appendix
Table 11: MMLongBench-Doc over all 1,091 questions, including the 244 annotated unanswerable, same runs and judging batch as Table 1 . Best per column in bold. † our reimplementation, for the reasons given in Appendix D .
Figure 8: One ViDoSeek question, Kimi-K3, from the runs behind Table 1 . The answer is on page 7 of a national renal registry report, and TAEC is the only system that retrieves it. (a) the three baselines: none reaches the page, and two answer anyway. (b) the TAEC trajectory, the same five searches ReAct issued: admission keeps 25 of the 48 candidates it sees across four calls scored on requirement coverage; exposure ranks the five memories it carries by an energy that decays with their age and compresses the one it has verified as resolved, so the carried graph is sent at 1119 characters rather than 1160; allocation spends 687K pixels over the five images exposure keeps visible; and the requirement state carried across turns provides the basis for the next query.
Neutral tokens added
0
16K
32K
Accuracy
78.8
74.8
73.6
Appendix
Table 12: The substrate on ViDoSeek, Gemini-3.5-Flash, with retrieval frozen and the gold page in the context in every column.
ReAct stopped after one search
ReAct searched again
Benchmark
Backbone
n
ReAct
Vanilla
Δ
n
ReAct
Vanilla
Δ
ViDoSeek
Gemini-3.5-Flash
642
93.5
86.8
+6.7
500
61.6
63.8
−2.2
GPT-4o-mini
910
62.2
63.8
−1.6
232
32.3
40.9
−8.6
SlideVQA
Gemini-3.5-Flash
1450
91.8
91.2
+0.6
765
49.4
57.6
−8.2
Kimi-K3
1762
88.6
86.2
+2.4
453
56.3
46.1
+10.2
GPT-4o-mini
1827
62.6
65.7
−3.1
388
44.1
44.3
−0.3
Appendix
Table 13: ReAct against single-pass retrieval on the same questions, split by the depth ReAct chose. Δ is ReAct minus Vanilla. The rows are the four cells of Table 1 where ReAct falls below single-pass retrieval, together with three where it does not: Gemini-3.5-Flash on ViDoSeek and MMLongBench-Doc, and Kimi-K3 on SlideVQA.
Gemini-3.5-Flash
Kimi-K3
GPT-5.6-Sol
GPT-4o-mini
Perceptual ability, on the substrate
Reads a retrieved gold page
85.0
85.3
82.5
64.2
Agentic ability, on the substrate
Commits without a gold page
22.6
12.2
14.8
9.5
Over-searches an answered question
6.0
15.5
12.8
47.9
Searches per question
1.33
1.72
1.62
3.10
Appendix
Table 14: Each column describes one acting model, averaged over the three benchmarks of Table 1 . The first block is measured on the substrate, where the four columns differ in nothing but the model. Reads a retrieved gold page is accuracy on the single-search questions whose gold page reached the context. The two agentic rows are the two directions in which a sufficiency judgement fails: commits without a gold page is the share of all questions answered after one search with no gold page in the context, and over-searches an answered question is the share of the questions whose gold page a single retrieval already reached on which the system issued a further search.
Backbone
Benchmark
Vanilla
ReAct
ViDoRAG
M3RAG †
DAG agent
TAEC
Qwen2.5-VL-7B
ViDoSeek
67.9
39.1
—
22.1
63.9
66.3
SlideVQA
57.0
33.8
—
—
55.3
55.4
Qwen3-VL-4B
ViDoSeek
68.1
63.4
70.2
58.1
68.0
71.5
SlideVQA
58.3
59.0
62.8
51.5
60.5
60.4
MMLongBench-Doc
15.0
15.8
4.7
14.5
14.8
16.1
VRAG-RL 7B (RL)
ViDoSeek
—
12.7
—
—
58.0
58.7
Appendix
Table 15: Accuracy on open-weight backbones, all rows judged with the prompt of Table 1 and run under the same harness; as in Table 1 , MMLongBench-Doc is scored on its 847 answerable questions. ViDoRAG was run on Qwen3-VL-4B only. The VRAG-RL row uses the released checkpoint as the acting model in our harness; its own loop does not transfer to our tool interface, which accounts for the value in the ReAct column. † our reimplementation, for the reasons given in Appendix D ; its planner and verifier run on the acting model, which is what the Qwen2.5-VL-7B row reflects.
Searches TAEC issued
n
Vanilla
TAEC
Δ
1
442
80.3
78.7
−1.6
2
139
76.3
71.2
−5.1
3–4
178
76.4
65.7
−10.7
5+
230
69.6
62.2
−7.4
Appendix
Table 16: Qwen2.5-VL-7B on the 989 ViDoSeek questions whose annotated page single-pass retrieval reached, grouped by the number of searches TAEC issued on them. The evidence was present after one search in every row.
System
Training
ViDoSeek
Source
Vanilla
none
32.9
Shen et al. (2026)
ReAct
none
33.9
Shen et al. (2026)
ReAct
none
39.1
ours
ViDoRAG
none
69.0
Shen et al. (2026)
M3RAG
none
69.4
Shen et al. (2026)
TAEC
none
66.3
ours
Appendix
Table 17: Qwen2.5-VL-7B on ViDoSeek. Published rows are taken from the VISOR table and use its judge, retriever and page collection; ours use the protocol of Table 1 . The two blocks are indicative of what each line of work reports on this backbone and are not comparable cell by cell. The bottom block uses the released VRAG-RL checkpoint as the acting model in our harness; its own ReAct-style loop does not transfer to our tool interface, which accounts for the low value in the first row.
Retriever
Vanilla
DAG agent
TAEC
Δ
Dense (main)
76.7
80.1
87.5
+7.4
BM25 over page text
72.8
80.2
80.7
+0.5
Hybrid (dense + link structure)
82.7
78.5
87.1
+4.4
Appendix
Table 18: The same three systems on retrievers of different strength, ViDoSeek, Gemini-3.5-Flash, same judging batch as the main tables. The last column is TAEC minus the stronger baseline in the row.
Searches used ≤k
Vanilla
ReAct
DAG agent
TAEC
2
76.7
62.4
79.0
86.3
3
76.7
67.3
79.4
86.8
5
76.7
72.2
79.9
87.3
8
76.7
75.6
80.1
87.5
10
76.7
76.9
80.1
87.5
20
76.7
79.5
80.1
87.5
Appendix
Table 19: Accuracy attained within k searches, ViDoSeek, Gemini-3.5-Flash, same runs and judging batch as Table 1 . A question counts when the system answered it using at most k searches and answered it correctly. The table reads the multi-step range, from two searches upward, where the components are active. No TAEC trajectory exceeds eight searches and no substrate trajectory exceeds seven, so both are saturated from eight on, while ReAct continues to gain up to nineteen and is still two and a half points short at ten.
Figure 9: The four-group decomposition, all three benchmarks, Gemini-3.5-Flash. (a) One cell per group and benchmark: the substrate’s accuracy and TAEC’s on that group, and its size; fill is the group’s contribution to the accuracy difference. (b) The same twelve groups against the identity line, point area ∝ group size.
Difference
Effect on the reported numbers
Retrieval index and pages placed per search held common to every system
Published numbers come from per-system indices of different strength; ours are read against one pooled index of 36,233 pages
One judge and one prompt for every row
ViDoRAG grades on a five-point scale with GPT-4o and counts four or above; an in-run 7B judge scores the same predictions 37.9 where GPT-4.1 scores 80.1
Answers scored as returned
A trajectory that terminates without an answer is scored on a plain-text fallback rather than as zero, which is worth 8.5 points to ReAct
Appendix
Table 20: Protocol differences between this paper and the works we compare against, and how each affects the numbers. Rows from different retriever fingerprints are never placed in one table.
Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually rich documents by retrieving relevant page images as visual evidence and reasoning over their content. However, effectively utilizing this visual evidence is usually impeded by two main challenges. First, answer-relevant evidence is sparse and may be concentrated in a small region of one page or dispersed across multiple pages. Second, existing agentic methods often generate answers based on raw exploration trajectories or compressed textual memories rather than an explicitly organized set of supporting images, making answers susceptible to exploration noise and obscuring the evidence-backed reasoning trace. We argue that the bottleneck lies not only in evidence discovery but also in its preservation and organization before answer generation. We propose SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation. During exploration, SCoRE retains only query-relevant observations and their source pointers in a maintained textual ledger, preserving earlier evidence while keeping the visual context bounded. At termination, it reloads the referenced original images and consolidates the visual evidence for answering, arranging it into a logical sequence. This decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages. To enable end-to-end optimization of this unified rollout, our training paradigm combines filtered cold-start trajectory distillation with evidence-aware reinforcement learning, whose reward promotes evidence coverage, consolidation compactness, and answer correctness.
Yucheng Shen, Lingyong Yan, Jiulong Wu +4
School of Computer Science and Technology, Soochow University · Baidu Inc.
Visual evidence selection is a critical component of multimodal retrieval-augmented generation (RAG), yet existing methods typically rely on semantic relevance or surface-level similarity, which are often misaligned with the actual utility of visual evidence for downstream reasoning. We reformulate multimodal evidence selection from an information-theoretic perspective by defining evidence utility as the information gain induced on a model's output distribution. To overcome the intractability of answer-space optimization, we introduce a latent notion of evidence helpfulness and theoretically show that, under mild assumptions, ranking evidence by information gain on this latent variable is equivalent to answer-space utility. We further propose a training-free, surrogate-accelerated framework that efficiently estimates evidence utility using lightweight multimodal models. Experiments on MRAG-Bench and Visual-RAG across multiple model families demonstrate that our method consistently outperforms state-of-the-art RAG baselines while achieving substantial reductions in computational cost.
Weiqing Luo, Zongye Hu, Xiao Wang +3
♦Arizona State University · ♥Texas A&M University · ♠Morgan Stanley
Iterative Retrieval-Augmented Generation (iRAG) has emerged as a powerful paradigm for answering complex multi-hop questions by progressively retrieving and reasoning over external documents. However, current systems predominantly operate on parsed text, which creates two critical bottlenecks: (1) \textit{Coarse-grained attribution}, where users are burdened with manually locating evidence within lengthy documents based on vague text-level citations; and (2) \textit{Visual semantic loss}, where the conversion of visually rich documents (e.g., slides, PDFs with charts) into text discards spatial logic and layout cues essential for reasoning. To bridge this gap, we present \textbf{Chain of Evidence (CoE)}, a retriever-agnostic visual attribution framework that leverages Vision-Language Models to reason directly over screenshots of retrieved document candidates. CoE eliminates format-specific parsing and outputs precise bounding boxes, visualizing the complete reasoning chain within the retrieved candidate set. We evaluate CoE on two distinct benchmarks: \textbf{Wiki-CoE}, a large-scale dataset of structured web pages derived from 2WikiMultiHopQA, and \textbf{SlideVQA}, a challenging dataset of presentation slides featuring complex diagrams and free-form layouts. Experiments demonstrate that fine-tuned Qwen3-VL-8B-Instruct achieves robust performance, significantly outperforming text-based baselines in scenarios requiring visual layout understanding, while establishing a retriever-agnostic solution for pixel-level interpretable iRAG. Our code is available at https://github.com/PeiYangLiu/CoE.git.
Peiyang Liu, Ziqiang Cui, Xi Wang +2
National Engineering Research Center for Software Engineering, Peking University Beijing, China · City University of Hong Kong Hong Kong SAR, China · Tencent Technology Beijing, China