Modern multi-hop LLM agents are equipped with built-in mechanisms to detect errors in intermediate reasoning steps. Such errors trigger corrective actions from these agents, which mostly follow the paradigm of retrying the steps or the reasoning trajectories. Not only are these retries expensive, we present in this paper that they are also potentially unnecessary. To this end, we introduce Evidence-Inference Reconstruction (EIR), which uses structured state to guide one retrieval trajectory, accumulating source evidence in the process. We show that as long as the relevant evidence has been collected, EIR is capable of generating the correct answer in a single final model call even if erroneous evidence has been mixed in due to incorrect intermediate reasoning steps. In one evaluation, using Haiku 4.5 and GPT-4.1 Mini, we evaluate EIR on matched 1,000-question subsets of HotpotQA, 2WikiMultiHopQA, and MuSiQue, showing that EIR improves Answer F1, the overlap between the model's and the correct answer, over the baseline by 8.3--32.8 points, Agentic SSR by 10.6--29.1 points, and Reflexion by 1.1--15.9 points. Additionally, we show that EIR averages 4.85 total model calls per question, compared with 35.29 for Agentic SSR and 12.41 for Reflexion. Together, these results corroborate EIR's central premise: separating evidence retrieval from the final answer model call can improve answer accuracy while utilizing substantially less computation.
Figures & tables
Figure 1: Evidence-Inference Reconstruction. (a) Requirements and evidence-citing conclusions guide document selection while retrieved source evidence is accumulated separately. (b) Retrieval ends with one final reconstruction from only the original question q and accumulated evidence ET ; the retrieval state and provisional answer are withheld.
Figure 2: EIR trace showing an incorrect intermediate candidate followed by a correct final answer. This instrumented reproduction, outside the confirmatory protocol, produces an incorrect candidate despite retrieving the necessary evidence. A fresh call over only the question and source text derives the correct answer. Appendix A documents its eight-call procedure.
Panel A. EIR and Direct
Answer F1
EIR minus Direct
Calls/question
Model
Benchmark
Direct
EIR
Difference
95% interval
Direct
EIR
Haiku 4.5
HotpotQA
64.37
72.64
+8.27
[ +6.39 , +10.18 ]
2.85
4.52
2WikiMultiHopQA
49.69
70.22
+20.53
[ +17.95 , +23.10 ]
3.68
5.11
MuSiQue
40.82
51.98
+11.16
[ +8.73 , +13.66 ]
3.49
5.32
GPT-4.1 Mini
HotpotQA
56.19
74.05
+17.85
[ +15.53 , +20.13 ]
2.78
4.08
Table 1: Main evaluation using the paragraphs provided with each question. Each row represents 1,000 questions. Bold marks the highest Answer F1 and fewest calls within each displayed group. Reflexion’s first trial includes no retry. Oracle stopping permits five trials and stops at the first exact match to the reference, with all trials and retries counted. Superscripts mark comparison scores significantly below EIR after Holm correction: Agentic SSR ( ∗ ), Reflexion first trial ( † ), and Reflexion with oracle stopping ( ‡ ).
Model
Benchmark
Full Pool
EIR
Haiku 4.5
HotpotQA
72.11
72.64
2WikiMultiHopQA
64.62
70.22
MuSiQue
53.51
51.98
GPT-4.1 Mini
HotpotQA
77.12
74.05
2WikiMultiHopQA
68.66
74.89
MuSiQue
57.84
57.54
Table 2: Answer F1 with every provided paragraph versus EIR’s selected evidence. Full Pool reads all paragraph text together and answers in one call, without model calls to select evidence. Bold marks the higher observed score, not statistical significance. Paired tests and intervals appear in Appendix G , Table 18 .
GPT-4.1 Mini
Haiku 4.5
Method
Answer F1
Calls/question
Answer F1
Calls/question
Direct
48.30
6.19
57.07
6.33
EIR
61.67
7.61
61.35
12.03
Agentic SSR
49.14
70.29
52.02
88.04
Reflexion, first trial
51.51
6.17
50.38
8.30
Reflexion, oracle stopping
56.99
27.92
62.21
32.56
Table 3: Extended HotpotQA evaluation requiring corpus search. The same 1,000 questions are answered using each model, with evidence found among 5.23 million paragraphs. Bold marks the highest Answer F1 and fewest calls among displayed methods. Oracle-stopped Reflexion includes all preceding trials and reflections, while first-trial scores include just the first trial.
Figure 3: EIR answer outcomes across benchmarks and models. Each bar represents 1,000 questions. Teal shows answers that exactly match the reference after normalization. The remaining answers are divided into those for which EIR read every annotated supporting paragraph (tan) and those missing at least one supporting paragraph (gray). Each percentage uses all 1,000 questions as its denominator, and the three portions sum to 100%. The teal portion includes all exact matches without distinguishing their evidence coverage. Counts appear in Appendix F , Table 15 , and Appendix C , Table 5 .
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Agentic SSR on the Humphrey question. The first two reads establish the relationships needed to identify his aunt. Subsequent actions pursue his father’s relatives and unrelated family branches. On the final turn, all three judgments reject the factual reasoning step, but their proposed corrections do not meet the agreement rule for refinement, and the agent submits Lady Anne Neville. The diagram shows all verification calls for the first turn and condenses later turns.
Figure 5: Reflexion on the Humphrey question. All five trials read the supporting relationships but identify Humphrey’s mother, Eleanor de Bohun, as his aunt. The intervening reflections call for further verification of his family relationships. None of the answers matches the reference, so all five trials run, using 18 action calls and four reflection calls.
Setting
Value
Questions and models
Model snapshots
claude-haiku-4-5-20251001 and gpt-4.1-mini-2025-04-14
Development split: 12,576 eligible questions, 11,506 unused earlier, 1,000 selected
MuSiQue
Answerable development questions: 2,417 eligible, 1,347 unused earlier, 1,000 selected
Selection
Questions drawn in sorted order from those unused in earlier evaluations, using fixed question lists
Appendix
Table 4: Question selection, model, statistical, and cost settings for the main evaluation. Output limits count generated tokens, and dollar costs are reconstructed at the fixed prices below rather than taken from invoices. Full Pool and full-corpus settings appear in Appendices G and H .
Answer EM
Support F1
Model
Benchmark
Direct
EIR
Direct
EIR
95% CI (diff.)
Haiku 4.5
HotpotQA
45.0
54.3
68.80
74.50
[ +4.13 , +7.18 ]
2Wiki
33.0
60.2
66.15
85.69
[ +17.18 , +21.83 ]
MuSiQue
31.5
40.7
48.83
67.64
[ +16.35 , +21.19 ]
GPT-4.1 Mini
HotpotQA
34.5
58.8
65.45
68.67
[ +1.67 , +4.81 ]
2Wiki
21.7
65.6
62.93
72.67
[ +7.71 , +11.83 ]
Appendix
Table 5: Exact Match and Support F1 for Direct and EIR. Each row includes all 1,000 questions, and bold marks the higher score in each pair. Intervals are paired bootstrap 95% intervals for EIR minus Direct Support F1, and each excludes zero.
Model
Benchmark
Δ F1 (pp)
95% CI (pp)
raw p
Holm p
sig.
EIR vs. Agentic SSR
Haiku 4.5
HotpotQA
+12.59
[ +10.46 , +14.73 ]
<10−4
<10−4
yes
Haiku 4.5
2Wiki
+29.12
[ +26.34 , +31.82 ]
<10−4
<10−4
yes
Haiku 4.5
MuSiQue
+10.56
[ +8.22 , +12.86 ]
<10−4
<10−4
yes
GPT-4.1 Mini
HotpotQA
+14.54
[ +12.40 , +16.70 ]
<10−4
<10−4
yes
GPT-4.1 Mini
2Wiki
+26.66
[ +24.08 , +29.32 ]
<10−4
<10−4
yes
Appendix
Table 6: Answer F1 differences, EIR minus each baseline, in percentage points. Each row compares both methods on the same 1,000 questions, with paired bootstrap 95% intervals. Holm adjustment covers the six rows for each baseline, and values below the reporting resolution appear as <10−4 .
Model
Benchmark
Δ EM (pp)
Only one exact ( b / c )
McNemar p
EIR vs. Agentic SSR
Haiku 4.5
HotpotQA
+12.90
28/157
<10−4
Haiku 4.5
2Wiki
+34.90
15/364
<10−4
Haiku 4.5
MuSiQue
+9.60
38/134
<10−4
GPT-4.1 Mini
HotpotQA
+21.40
28/242
<10−4
GPT-4.1 Mini
2Wiki
+36.70
19/386
<10−4
Appendix
Table 7: Exact Match differences, EIR minus each baseline, in percentage points. Each row compares both methods on the same 1,000 questions, and in b/c , b counts questions answered exactly only by the baseline and c those answered exactly only by EIR. p -values are unadjusted exact two-sided McNemar tests.
Model
Benchmark
EIR
SSR
Refl. t1
Refl. oracle
Δ (EIR − oracle)
Holm p , sig.
Haiku 4.5
HotpotQA
74.50
68.73
71.00
70.34
+4.16
<10−4 , yes
2Wiki
85.69
67.30
77.17
79.26
+6.44
<10−4 , yes
MuSiQue
67.64
51.22
61.23
58.44
+9.21
<10−4 , yes
GPT-4.1 Mini
HotpotQA
68.67
64.07
67.08
69.36
−0.68
0.777 , no
2Wiki
72.67
62.16
71.52
74.97
−2.31
0.009 , yes
MuSiQue
62.50
53.28
59.91
63.38
−0.88
0.777 , no
Appendix
Table 8: Support F1 on a 0–100 scale. Each row includes all 1,000 questions, where SSR is Agentic SSR, Refl. t1 is Reflexion’s first trial, and Refl. oracle is oracle-stopped Reflexion. The last two columns give EIR-minus-oracle differences with Holm adjustment across those six comparisons.
Method
Answer F1
Calls/q
Calls/EIR
$/q
Cost/EIR
EIR
66.89
4.85
1.00 ×
0.00738
1.00 ×
Agentic SSR
47.52
35.29
7.28 ×
0.02681
3.63 ×
Reflexion (oracle)
58.56
12.41
2.56 ×
0.01149–0.01274
1.56–1.73 ×
Direct
47.93
3.24
0.67 ×
0.00311
0.42 ×
Appendix
Table 9: Average accuracy and computation in the main evaluation. Rows are methods averaged over the three benchmarks and two models, with ratios relative to the EIR mean and dollar costs at the fixed prices in Table 4 . Bold marks the highest Answer F1, fewest calls, and lowest cost, and the Reflexion cost range reflects missing token breakdowns rather than statistical uncertainty.
Figure 6: EIR accuracy across read budgets. Each point averages the same 1,000 questions per benchmark and model. The dotted line marks the budget used in the main evaluation, H=6 .
Model
Benchmark
H=4
H=6
H=8
H=10
H=10 vs. H=6
Haiku 4.5
HotpotQA
4.24
4.52
4.83
5.09
+12.6%
2WikiMultiHopQA
4.74
5.11
5.42
5.78
+13.1%
MuSiQue
4.76
5.32
5.79
6.14
+15.4%
GPT-4.1 Mini
HotpotQA
3.91
4.08
4.14
4.20
+3.1%
2WikiMultiHopQA
4.54
4.86
5.04
5.27
+8.4%
MuSiQue
4.54
5.19
5.73
6.00
+15.6%
Appendix
Table 10: Mean model calls per question across read budgets. Each value averages 1,000 questions and includes the final answer call. The last column gives the relative increase from H=6 to H=10 .
Model
Benchmark
H=4
H=8
H=10
Haiku 4.5
HotpotQA
−0.65
+0.68
+0.87
2WikiMultiHopQA
−0.48
−0.21
+0.08
MuSiQue
−3.18∗
+1.46
+3.37∗
GPT-4.1 Mini
HotpotQA
−1.07
−0.20
−0.20
2WikiMultiHopQA
−0.29
−1.68
−0.03
MuSiQue
−1.90
+0.25
+2.11
Appendix
Table 11: Answer F1 change relative to H=6 . Each value is the named budget minus H=6 on the same 1,000 questions, in percentage points, calculated before rounding. Differences use paired bootstrap tests with 10,000 resamples, and Holm adjustment covers all 18 comparisons. ∗ marks significance after adjustment.
Model
Benchmark
Corrected control
EIR
Difference
95% CI
Holm p
Haiku 4.5
HotpotQA
73.91
72.64
−1.27
[ −2.71 , +0.12 ]
0.2256
2Wiki
69.41
70.22
+0.80
[ −0.48 , +2.13 ]
0.4120
MuSiQue
47.15
51.98
+4.83∗
[ +2.79 , +6.86 ]
<0.001
GPT-4.1 Mini
HotpotQA
75.49
74.05
−1.44
[ −2.83 , −0.06 ]
0.1624
2Wiki
72.45
74.89
+2.45∗
[ +1.08 , +3.89 ]
0.0060
MuSiQue
56.25
57.54
+1.30
[ −0.74 , +3.28 ]
0.4120
Appendix
Table 12: Answer F1 with and without the lists in later retrieval prompts, after correcting the control. Each row compares both methods on the same 1,000 questions, and differences are EIR minus the corrected control, calculated before rounding. Intervals are paired bootstrap 95% intervals with 10,000 resamples and seed 20260806. Holm adjustment covers the six comparisons, and ∗ marks significance after adjustment.
Model
Benchmark
Repair
EIR
Difference
95% CI
Holm p
Δ EM
Haiku 4.5
HotpotQA
68.47
72.64
−4.17
[ −6.02 , −2.37 ]
<10−4
−1.80
Haiku 4.5
2Wiki
65.64
70.22
−4.57
[ −6.80 , −2.30 ]
<10−4
−3.40
Haiku 4.5
MuSiQue
53.32
51.98
+1.34
[ −0.67 , +3.36 ]
0.1968
+1.20
GPT-4.1 Mini
HotpotQA
71.08
74.05
−2.97
[ −4.76 , −1.21 ]
0.0030
−4.40
GPT-4.1 Mini
2Wiki
69.83
74.89
−5.06
[ −6.98 , −3.17 ]
<10−4
−4.60
GPT-4.1 Mini
MuSiQue
56.21
57.54
−1.33
[ −3.70 , +1.03 ]
0.3476
−2.10
Appendix
Table 13: Answer F1 for dependency repair and default EIR. Each row compares both methods on the same 1,000 questions, and differences are Repair minus EIR, calculated before rounding. Holm p -values keep the original groups of four comparisons within each benchmark and model, including comparisons not displayed here.
Model
Benchmark
Repairs applied
Detected, not applied
Difference
95% CI
Holm p
Haiku 4.5
HotpotQA
68.47
68.81
−0.34
[ −2.02 , +1.26 ]
1.000
2Wiki
65.64
65.37
+0.28
[ −1.63 , +2.21 ]
1.000
MuSiQue
53.32
51.06
+2.26
[ +0.67 , +3.92 ]
0.031
GPT-4.1 Mini
HotpotQA
71.08
71.27
−0.20
[ −1.86 , +1.45 ]
1.000
2Wiki
69.83
68.75
+1.08
[ −0.87 , +3.01 ]
1.000
MuSiQue
56.21
54.34
+1.87
[ −0.19 , +3.88 ]
0.379
Appendix
Table 14: Answer F1 with proposed repairs applied or left unapplied. Each row compares both methods on the same 1,000 questions, and differences are applied minus unapplied. Holm adjustment covers the six comparisons.
Model
Benchmark
EM errors
Complete
Incomplete
Incomplete (%)
Haiku 4.5
HotpotQA
457
286
171
37.42
2Wiki
398
368
30
7.54
MuSiQue
593
220
373
62.90
GPT-4.1 Mini
HotpotQA
412
192
220
53.40
2Wiki
344
265
79
22.97
MuSiQue
554
179
375
67.69
Appendix
Table 15: Evidence read for answers that fail Exact Match. Rows count EM errors per benchmark and model, divided into those with every annotated supporting paragraph read (Complete) and those missing at least one (Incomplete). Incomplete (%) divides by EM errors, not by all 1,000 questions.
Model
Benchmark
N
Added support
Model-selected
Exact (%)
Haiku 4.5
HotpotQA
171
46
20
11.70
2Wiki
30
3
3
10.00
MuSiQue
373
89
35
9.38
GPT-4.1 Mini
HotpotQA
220
68
39
17.73
2Wiki
79
18
15
18.99
MuSiQue
375
111
62
16.53
Appendix
Table 16: Exact answers after adding evidence to the same known failures. Rows are the EM errors missing support ( N ) per benchmark and model. Added support supplies the missing annotated paragraphs, Model-selected lets the model choose at most two more, and Exact (%) divides Model-selected by N .
Model
Benchmark
EIR Answer F1
Extended Answer F1
Δ F1
95% CI
Δ EM
More requested (%)
Haiku 4.5
HotpotQA
72.64
74.54
+1.90
[ +1.06,+2.82 ]
+1.90
17.60
2Wiki
70.22
70.07
−0.15
[ −0.52,+0.17 ]
−0.20
21.00
MuSiQue
51.98
54.72
+2.74
[ +1.50,+4.04 ]
+2.30
53.00
GPT-4.1 Mini
HotpotQA
74.05
78.25
+4.20
[ +2.72,+5.72 ]
+3.90
57.50
2Wiki
74.89
75.63
+0.74
[ −0.40,+1.87 ]
+1.00
56.90
MuSiQue
57.54
62.43
+4.89
[ +3.01,+6.74 ]
+3.90
85.50
Appendix
Table 17: Additional retrieval applied to every question. Each row compares both methods on the same 1,000 questions, and More requested (%) is the percentage for which the model asks to read more. Differences are computed before rounding, and both 2WikiMultiHopQA intervals include zero.
Model
Benchmark
EIR
Full Pool
Difference
Paired 95% CI
Haiku 4.5
HotpotQA
72.64
72.11
+0.52
[ −1.30 , +2.30 ]
2Wiki
70.22
64.62
+5.60∗
[ +3.44 , +7.74 ]
MuSiQue
51.98
53.51
−1.53
[ −4.04 , +0.98 ]
GPT-4.1 Mini
HotpotQA
74.05
77.12
−3.08∗
[ −4.88 , −1.24 ]
2Wiki
74.89
68.66
+6.23∗
[ +4.30 , +8.23 ]
MuSiQue
57.54
57.84
−0.29
[ −2.79 , +2.18 ]
Appendix
Table 18: Answer F1 for EIR and Full Pool. Each row compares both methods on the same 1,000 questions, and differences are EIR minus Full Pool, calculated before rounding. ∗ marks significance after Holm adjustment across the six tests.
Model
Method
Answer F1
EM
Support F1
Calls
Cost ($)
GPT-4.1 Mini
Direct
48.30
27.30
50.99
6,194
3.4639
EIR
61.67
46.30
55.31
7,606
5.2582
Agentic SSR
49.14
28.70
49.25
70,292
34.4042
Reflexion trial 1
51.51
31.00
49.14
6,169
2.3206
Reflexion (oracle)
56.99
38.60
54.33
27,920
15.6130
Haiku 4.5
Direct
57.07
37.70
59.19
6,329
11.3336
Appendix
Table 19: Full-corpus HotpotQA results. Rows are methods on 1,000 questions per model, with accuracy on a 0–100 scale and calls and costs totaled over all scored questions. Oracle-stopped Reflexion includes every trial and reflection, and costs use the fixed prices in Table 4 .
Model
EIR versus
Answer F1 difference
Paired 95% CI
Answer F1 Holm p
EM Holm p
GPT-4.1 Mini
Direct
+13.38
[ +10.91 , +15.83 ]
0.0010
9.938×10−31
Agentic SSR
+12.53
[ +10.11 , +14.96 ]
0.0010
9.454×10−28
Reflexion trial 1
+10.16
[ +7.73 , +12.60 ]
0.0010
5.997×10−21
Reflexion (oracle)
+4.68
[ +2.01 , +7.33 ]
0.0024
1.019×10−5
Haiku 4.5
Direct
+4.28
[ +2.09 , +6.43 ]
0.0010
9.293×10−9
Agentic SSR
+9.33
[ +6.91 , +11.68 ]
0.0010
4.217×10−16
Appendix
Table 20: Full-corpus paired comparisons, EIR minus the named baseline. Each row compares both methods on the same 1,000 questions, with Answer F1 differences calculated before rounding and a separate Exact Match test in the last column. Holm p -values keep the original groups of five comparisons per metric and model, four of which are shown, and 0.0010 rounds 0.0009999 rather than denoting p<10−4 .
Model
Method
Questions
Later actions in same trial
Actions in later trials
Reflections
Final answer calls
GPT-4.1 Mini
EIR
420
1,004
—
—
420
Reflexion (oracle)
547
846
8,442
1,299
—
Haiku 4.5
EIR
589
3,581
—
—
589
Reflexion (oracle)
542
602
9,235
1,106
—
Appendix
Table 21: Calls after annotated support is first acquired. Questions counts runs in which EIR, or at least one Reflexion trial, reads every paragraph labeled as necessary. Counts exclude the read that completes that support, dashes mark call types absent from a method, and Agentic SSR is not covered.
Model
Method
Rejected actions
Recorded actions
Rejected (%)
GPT-4.1 Mini
EIR
216
6,606
3.27
Direct
41
6,194
0.66
Reflexion, all trials
619
25,339
2.44
Haiku 4.5
EIR
2,377
11,024
21.56
Direct
70
6,328
1.11
Reflexion, all trials
14,045
30,154
46.58
Appendix
Table 22: Rejected actions in the full-corpus evaluation. Rows are methods per model, with denominators counting recorded actions and excluding unfinished requests, reflection, and separate final answering. Reflexion totals include actions from all trials for the same questions.
Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with retrieval-augmented generation (RAG) remains fundamentally misaligned. Current RAG systems optimize for providing context before reasoning begins, while reasoning models require evidence injection during multi-step inference chains. We introduce ReaLM-Retrieve, a reasoning-aware retrieval framework that addresses this mismatch through three key innovations: (1) a step-level uncertainty detector that identifies knowledge gaps at reasoning-step granularity rather than token or sentence level; (2) a retrieval intervention policy that learns when external evidence maximally benefits ongoing reasoning; and (3) an efficiency-optimized integration mechanism that reduces per-retrieval overhead by 3.2x compared to naive integration. Experiments on MuSiQue, HotpotQA, and 2WikiMultiHopQA demonstrate that ReaLM-Retrieve achieves on average 10.1% absolute improvement in answer F1 over standard RAG (range: 9.0-11.8% across the three benchmarks) while reducing retrieval calls by 47% compared to fixed-interval approaches like IRCoT (all improvements significant at p<0.01, paired bootstrap). On the challenging MuSiQue benchmark requiring 2-4 hop reasoning, our method achieves 71.2% F1 with an average of only 1.8 retrieval calls per question. Analysis shows that ReaLM-Retrieve also improves retrieval quality itself, achieving 81.3% Recall@5 with consistently higher precision and MRR than fixed-interval baselines on supporting evidence, establishing new state-of-the-art efficiency-accuracy trade-offs for reasoning-intensive retrieval tasks.
Dongxin Guo, Jikun Wu, Siu Ming Yiu
The University of Hong Kong Hong Kong, China · Brain Investing Limited Hong Kong, China · Stellaris AI Limited Hong Kong, China
Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a procedural property of the agent trajectory, decomposing wrong answers into pre-evidence discipline failures and post-gold-read failures using saved tool-call traces, retrieved evidence, read passages, and final answers. Across 12,000 paired trajectories on HotpotQA, 2WikiMultiHopQA, and MuSiQue, the two failure types are largely non-redundant: the both-trigger rate is in [11.2%, 13.1%] across regex and spaCy entity extractors. We then evaluate Read-Gate, a minimal runtime invariant requiring an agent to read after search and before finalization. Forced reading improves LLM-Acc by 14.9-19.9 points on trajectories that would otherwise skip reading and by 3.2-9.4 points on full minimal-reasoning cells. Additional diagnostics show that larger hidden thinking budgets do not necessarily increase evidence inspection. Together, these results indicate that evidence-gathering should be evaluated as a trajectory-level control problem, separately from answer-side reasoning.
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.
Enjun Du, Hange Zhou, Chenxu Du +4
The Hong Kong University of Science and Technology (Guangzhou) · The University of Hong Kong · Tsinghua University +1