Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.
Figures & tables
Benchmark
Full
Drop-one
Control
BNR
MuSiQue [ 25 ]
87.4
65.4
87.3
27.4
HotpotQA [ 31 ]
81.2
69.7
81.5
16.6
2WikiMultiHopQA [ 7 ]
88.6
74.5
89.1
17.0
FalseCoTQA [ 21 ]
63.8
36.3
59.9
43.5
Plausible Distractors [ 3 ]
87.8
44.8
86.4
48.9
Table 1: Panel-mean Full, Drop-one, Control, and BNR on existing multi-hop benchmarks (percent). Means give equal weight to DeepSeek V4 Flash, Gemini 3.6 Flash, and Qwen 3.7 Max. Drop-one and Control are accuracies over all audited questions, whereas BNR conditions on Full-correct ones.
LongBench v2
RealHop-LongBench
Model
Eff.
Acc.
Rank
Trial 1
Trial 2
Trial 3
Avg@3
Rank
Change
MCQ
Claude Opus 5
max
70.4
2
80.6
85.2
81.9
82.6
1
↑1
94.4
GPT-5.6 Sol
xhigh
68.1
3
63.9
63.9
65.7
64.5
2
↑1
75.5
Kimi K3
max
65.7
6
52.8
59.7
50.9
54.5
3
↑3
64.8
Claude Opus 4.8
max
64.8
8
53.7
49.1
55.1
52.6
4
↑4
79.6
Grok 4.6
xhigh
65.3
7
43.1
40.3
44.0
42.4
5
↑2
60.2
Table 2: Source ranks use accuracy on the corresponding source questions under the official LongBench v2 multiple-choice template. Trials 1–3 and Avg@3 are open-ended scores on RealHop-LongBench ; the rightmost MCQ column evaluates the same constructed items under that template, run once. Eff. denotes the provider-specific reasoning-effort setting. Arrows give the rank change from the source leaderboard to Avg@3. Spread is the range of displayed scores across models.
MuSiQue
RealHop-MuSiQue
Metric
Panel
Model
Original
Gold-only
+Flexible
+Strict
Full
Accuracy
Earlier
DeepSeek R1
84.9
72.6
57.5
42.5
37.0
Doubao 1.5 Pro
74.0
67.1
46.6
30.1
24.7
GPT-4o
71.2
68.5
31.5
24.7
16.4
Qwen Max
64.4
58.9
28.8
24.7
21.9
GLM-4-Plus
58.6
50.0
21.4
20.0
11.4
Table 3: Component ablation on a 73-item MuSiQue cohort (percent). Original denotes source questions. Accuracy uses model-specific complete-case subsets; BNR averages hops within each question and is reported only for current models. Mean weights models equally.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Model
Full
Drop-one
Control
BNR
MuSiQue
DeepSeek V4 Flash
87.3
64.1
86.9
28.9
Gemini 3.6 Flash
88.7
68.2
87.8
25.2
Qwen 3.7 Max
86.1
64.0
87.3
28.2
HotpotQA
DeepSeek V4 Flash
80.7
68.8
81.3
17.2
Gemini 3.6 Flash
80.7
69.0
81.5
16.9
Qwen 3.7 Max
82.3
71.2
81.8
15.6
Appendix
Table 4: Per-model Full, Drop-one, Control, and BNR on the external multi-hop audit (percent).
Drop survival
Control survival
Model
Original
RealHop-MuSiQue
Original
RealHop-MuSiQue
DeepSeek V4 Flash
73.4
5.7
96.6
93.0
Qwen 3.7 Max
74.4
5.8
97.8
89.7
Gemini 3.6 Flash
76.3
6.2
97.2
92.8
Appendix
Table 5: Matched-background specificity on aligned MuSiQue support deletions (percent), computed on Full-correct items with valid Drop and Control results.
Model
∣I∩∣
BNR Original
BNR RealHop-MuSiQue
Increase [95% CI]
DeepSeek V4 Flash
559
28.04
94.66
66.6 [64.0, 69.1]
Gemini 3.6 Flash
599
24.36
94.05
69.7 [67.2, 72.2]
Qwen 3.7 Max
482
27.35
94.90
67.5 [64.8, 70.3]
Appendix
Table 6: MuSiQue BNR on questions with Full correctness on both sides. BNR is percent; the paired increase and its interval are percentage points.
Model
Wrong hit
Share of errors
DeepSeek V4 Flash
14.4
69.5
Gemini 3.6 Flash
11.1
67.7
Qwen 3.7 Max
20.3
64.3
Appendix
Table 7: Planned competitor-endpoint matches among Full answers on the 790 RealHop-MuSiQue questions (percent).
Model
Paired n
Original
RealHop-MuSiQue
Δ
GPT-4o
353
62.6
17.8
−44.8
Qwen Turbo
352
33.8
8.5
−25.3
Qwen Max
352
55.1
15.6
−39.5
GLM-4-Plus
337
51.9
15.1
−36.8
Doubao 1.5 Pro
340
64.7
20.9
−43.8
DeepSeek R1
305
79.3
31.5
−47.9
Appendix
Table 8: Paired accuracy on source MuSiQue and the corresponding RealHop-MuSiQue questions for six earlier models (percent).
Question
Reference key
Defect
Who is the child of Mahmoud Mirza’s father?
Ahmad Shah Qajar
Self-referential: Mahmoud Mirza is himself a child of that father, and all nine models answer with his name.
What team does the winner of the 2017 BBC African Footballer of the Year play for?
Egypt national football team
Club and national team are both valid; models answer with the club.
What mountain can you see from Portland, in the state that Raven Creek is located in?
Tualatin Mountains
Several mountains are visible from Portland; the key records one.
Bancroft’s county borders what county?
Haliburton County
Multiple bordering counties; answers that name the key alongside another county are scored incorrect.
When did Nissan, the Acura Legend maker and the Scion owner open US assembly plants?
1981
Models answer “early 1980s”, which the key cannot accept.
When did the luxury division of the employer of Katsuaki Watanabe change the body style of the rx 350?
Sales began worldwide in April 2012
Key is a sentence rather than a date; models answer “March 2012”.
Appendix
Table 9: The seven source questions removed from the component ablation.
Model
Full
BNR [95% CI]
DeepSeek V4 Flash
32.4
97.6 [94.3, 100.0]
Gemini 3.6 Flash
41.2
98.1 [95.5, 100.0]
Qwen 3.7 Max
26.9
96.5 [93.1, 99.4]
Appendix
Table 10: Targeted-hop intervention results on RealHop-LongBench (percent).
Model
Full
Mask 1
Mask 2
Mask 3
Mask avg.
Accuracy decrease [95% CI]
Full-correct survival
DeepSeek V4 Flash
32.4
18.5
16.7
16.7
17.3
15.1 [8.8, 21.6]
24.3
Gemini 3.6 Flash
41.2
19.4
17.1
15.7
17.4
23.8 [16.8, 30.6]
21.0
Qwen 3.7 Max
26.9
17.6
11.6
17.1
15.4
11.4 [5.2, 17.7]
20.1
Appendix
Table 11: Random 20% mask results on RealHop-LongBench (percent).
Benchmark
Model
Complete n
Full fixed
Mask fixed
Δ fixed
Δ paired [95% CI]
LongBench v2
DeepSeek
100
73.0
68.7
−4.3
−4.3 [ −10.0,1.7 ]
Gemini
100
71.0
69.0
−2.0
−2.0 [ −6.3,2.7 ]
Qwen
94
69.0
65.3
−3.7
−1.8 [ −7.8,4.3 ]
NarrativeQA
DeepSeek
100
57.0
57.0
−0.0
−0.0 [ −6.7,6.7 ]
Gemini
100
54.0
54.3
+0.3
+0.3 [ −5.3,6.3 ]
Qwen
90
47.0
44.7
−2.3
−5.9 [ −13.0,1.1 ]
Appendix
Table 12: Full and random-mask accuracy on four long-context audits (percent).
Model
Trial
n+
BNR
O
R
O
R
DeepSeek V4 Flash
1
79
59
7.81
89.83
2
79
65
6.75
88.03
3
79
62
8.44
88.87
Gemini 3.6 Flash
1
80
71
5.21
89.18
2
79
69
7.59
87.78
Appendix
Table 13: FRAMES BNR recomputed independently for each recorded trial (percent). O/R denote source/constructed items; n+ is the corresponding Full-correct question count. Complete-panel counts are 80, 81, and 79 throughout. Avg@3 Full and BNR in Section 6.2 are the arithmetic mean of these trial-level metrics.
Model
Counts
BNR O
BNR R
Increase [95% CI]
DeepSeek V4 Flash
59/65/61
6.80
88.85
82.1 [75.2, 88.2]
Gemini 3.6 Flash
70/67/68
5.75
89.30
83.5 [77.4, 89.1]
Qwen 3.7 Max
49/56/52
7.79
89.57
81.8 [74.7, 88.2]
Appendix
Table 14: FRAMES BNR on trials with Full correctness on both sides. Counts give ∣It∩∣ for trials 1/2/3; O/R are source/constructed. BNR is percent; the paired increase and its interval are percentage points.
Model
Filter
n/K
Full
Control
Drop
C−D [95% CI]
DeepSeek V4 Flash
All feasible
80/204
77.50
78.09
9.81
68.3 [60.9, 75.6]
±25%
80/191
77.50
77.74
7.92
69.8 [62.7, 76.8]
±10%
56/78
74.40
77.38
9.33
68.1 [57.9, 77.7]
±5%
30/38
74.44
80.56
12.78
67.8 [53.3, 81.1]
Gemini 3.6 Flash
All feasible
81/207
86.01
83.82
9.37
74.5 [66.9, 81.6]
±25%
81/194
86.01
84.40
7.72
76.7 [69.2, 83.8]
Appendix
Table 15: Constructed FRAMES background-deletion sensitivity. n/K counts questions/paired hop conditions, not model calls. All score columns are item-macro condition means multiplied by 100; C−D is Control minus Drop, with a paired question-bootstrap interval. Length bands additionally require sentence-aligned deletion.
Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks two critical vulnerabilities: (1) reliance on internal parametric knowledge rather than adherence to the provided context, and (2) exploitation of dataset shortcuts, such as single-document cues or type-matching, that diminish the need for genuine evidence aggregation across multiple documents. We introduce CRiT-QA (Counterfactual Reasoning with Traps), a dataset explicitly designed to address both limitations. To neutralize reliance on memorized knowledge and enforce strict context dependency, CRiT-QA transforms factual reasoning chains with counterfactual entities. Furthermore, it injects multi-anchor distractor chains, plausible but incorrect reasoning paths that diverge at different hops. These traps require models to follow the entire reasoning process rather than exploiting shallow heuristics. Our experiments show that LLMs exhibit substantial performance degradation on CRiT-QA compared to standard datasets, exposing their vulnerability to counterfactual conditions and distractor traps. CRiT-QA thus serves as a rigorous diagnostic tool for evaluating genuine multi-hop reasoning and provides a foundation for developing more reliable, evidence-grounded LLMs.
JungMin Yun, JuneHyoung Kwon, YoungBin Kim
Department of Artificial Intelligence, Chung-Ang University · Graduate School of Advanced Imaging Sciences, Multimedia and Film, Chung-Ang University
Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park +5
University of Wisconsin-Madison · Massachusetts Institute of Technology · Stanford University +2
With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop reasoning, which requires models to engage in deliberate thinking and multi-step interaction, serves as a critical testbed for assessing such capabilities. However, existing benchmarks typically provide only final questions and answers, while lacking the intermediate hop-level questions that gradually connect atomic questions to the final multi-hop query. This limitation prevents researchers from analyzing at which step an agent fails and restricts more fine-grained evaluation of model capabilities. Moreover, most current benchmarks are manually constructed, which is both time-consuming and labor-intensive, while also limiting scalability and generalization. To address these challenges, we introduce AgenticRAGTracer, the first Agentic RAG benchmark that is primarily constructed automatically by large language models and designed to support step-by-step validation. Our benchmark spans multiple domains, contains 1,305 data points, and has no overlap with existing mainstream benchmarks. Extensive experiments demonstrate that even the best large language models perform poorly on our dataset. For instance, GPT-5 attains merely 22.6% EM accuracy on the hardest portion of our dataset. Hop-aware diagnosis reveals that failures are primarily driven by distorted reasoning chains -- either collapsing prematurely or wandering into over-extension. This highlights a critical inability to allocate steps consistent with the task's logical structure, providing a diagnostic dimension missing in traditional evaluations. We believe our work will facilitate research in Agentic RAG and inspire further meaningful progress in this area. Our code and data are available at https://github.com/YqjMartin/AgenticRAGTracer.
Qijie You, Wenkai Yu, Wentao Zhang
University of Science and Technology Beijing · 2Peking University · 3Zhongguancun Academy +1