AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
Organizations: Korea University · Sookmyung Women’s University
Abstract
Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.
Figures & tables
| Benchmark | N | Search | Synth. | Tool-use | Resource | Diagnostic |
| HotpotQA [ 16 ] | 113K | ✓ | ✓ | |||
| 2WikiMultiHopQA [ 17 ] | 167K | ✓ | ||||
| OpenScholar [ 7 ] | 2,967 | ✓ | ✓ | |||
| WebArena [ 18 ] | 812 | ✓ | ✓ | ✓ | ||
| GAIA [ 19 ] | 466 | ✓ | ✓ | ✓ | ||
| BrowseComp [ 20 ] | 1,266 | ✓ | ✓ | ✓ |
| Model | Search | Synthesis | Tool-use | Resource | Failure Ratio | Outcome | ||||||||
| Recall (ST) | Recall (MT) | SecRec (ST) | SecRec (MT) | Conversion (ST) | Conversion (MT) | Calls /Turn | Avg. Tokens | Avg. Turns | Avg. Budgets | Max Tokens | Max Turns | Max Budgets | Accuracy | |
| Closed Frontier | ||||||||||||||
| Gemini-3 Pro | 0.868 | 0.752 | 0.859 | 0.655 | 0.957 | 0.970 | 1.05 | 92.4 | 9.4 | 19.8 | 4.4 | 0.1 | 0.0 | 0.891 |
| Gemini-3 Flash | 0.724 | 0.501 | 0.680 | 0.386 | 0.625 | 0.644 | 1.01 | 180.3 | 14.1 | 27.9 | 27.7 | 2.1 | 0.2 | 0.547 |
| GPT-5.3 codex | 0.812 | 0.623 | 0.775 | 0.470 | 0.896 | 0.902 | 1.44 | 69.5 | 8.5 | 25.2 | 0.3 | 0.0 | 0.7 | 0.825 |
| GPT-5.4 | 0.808 | 0.643 | 0.752 | 0.454 | 0.878 | 0.846 | 1.74 | 67.5 | 7.7 | 26.5 | 0.3 | 0.0 | 4.9 | 0.790 |
| Model | n | Recover | Stay | Think | Re-nav | Submit |
| Closed Frontier | ||||||
| Gemini-3 Pro | 882 | 3.2% | 33.9% | 5.4% | 50.2% | 7.3% |
| Gemini-3 Flash | 2165 | 1.2% | 39.7% | 1.5% | 50.2% | 7.0% |
| GPT-5.3 codex | 1559 | 1.9% | 39.8% | 0.0% | 53.4% | 4.9% |
| GPT-5.4 | 1652 | 3.1% | 39.6% | 1.8% | 52.1% | 3.0% |
| Claude Opus 4.6 | 904 | 1.4% | 12.7% | 82.3% | 3.4% | 0.1% |
| Model | Zero-tool | No-search | Full-tool | ||
| (%) | (pp) | (%) | (pp) | (%) | |
| DeepSeek-chat | 18.3 | 53.3 | 56.4 | ||
| Gemma-4-31B | 31.7 | 54.8 | 64.8 | ||
| GPT-5.3 codex | 48.6 | 79.3 | 82.5 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Description | Drops | Cumulative |
| Generation | |||
| 1 | Seed selection | — | 953 |
| 2 | Citation-chain expansion | — | 17,983 |
| 3 | Q/A generation (BFS + DFS) | — | 1,651 |
| 4 | Distractor generation (3-trap) | — | 1,651 |
| Filtering | |||
| Theme | Description | |
| Paper-content extraction broken | 12 | The gold paper’s HTML or section-extraction pipeline produced empty, mangled, or partially captured sections; the auditor saw that the cited paper content needed to ground the answer was missing. |
| Weak chain / target paper under-utilised | 6 | The synthesis relies on only one of the two cited papers, or the bridge terminal step does not actually carry the information the question demands. |
| Trivial or low-quality question | 6 | The question or its gold answer is too simple to be diagnostic of multi-hop competence (e.g., the answer is a single-word fact derivable without genuine multi-paper reasoning). |
| Answer-question alignment failure | 6 | The gold answer is on-topic but does not satisfy what the question explicitly asks ( how vs what , missing tradeoff or comparison, synthesis claim that doesn’t follow from the cited evidence). |
| Other (sundry, 1 each) | 5 | Option text misquotes the paper; a numerical detail is incorrect; a non-gold option is also defensible; the answer compares the wrong pair of methods; or the asserted target-paper claim actually originates in the seed paper. |
| Total retired at Stage 7 | 35 |
| Display name | Identifier | Display name | Identifier |
| Gemini-3 Pro | gemini-3-pro-preview | MiniMax-M2.7 | MiniMaxAI/MiniMax-M2.7 |
| Claude Opus 4.6 | claude-opus-4-6 | Qwen3.5-27B | Qwen/Qwen3.5-27B |
| GPT-5.3-codex | gpt-5.3-codex | Qwen3.5-35B-A3B | Qwen/Qwen3.5-35B-A3B |
| Claude Sonnet 4.6 | claude-sonnet-4-6 | Qwen3.6-35B-A3B | Qwen/Qwen3.6-35B-A3B |
| GPT-5.4 | gpt-5.4 | Qwen3-32B | Qwen/Qwen3-32B |
| Gemini-3 Flash | gemini-3-flash-preview | Qwen3-30B-A3B | Qwen/Qwen3-30B-A3B-Instruct-2507 |
| Model | Full-tool ($) | Zero-tool ($) | $/sample (full) | Input/sample (K) | Output/sample |
| Gemini 3 Pro | 159.27 | 37.14 | 0.158 | 87 | 5,003 |
| Gemini 3 Flash | 77.49 | 25.17 | 0.077 | 169 | 8,517 |
| Claude Opus 4.6 | 405.54 | 79.47 | 0.401 | 107 | 3,890 |
| Claude Sonnet 4.6 | 205.23 | 43.86 | 0.203 | 92 | 4,165 |
| GPT-5.3-codex | 80.32 | 14.80 | 0.079 | 67 | 2,208 |
| GPT-5.4 | 78.99 | 28.14 | 0.078 | 65 | 2,444 |
| MT-fail | ST (after gold) | MT-both (after both) | |||||||||||
| Model | 0-Gold | Prem | Explored | Stayed | Imm | Thnk | Read | Imm | Thnk | Read | |||
| Closed Frontier | |||||||||||||
| Gemini-3 Pro | 110 | 20.9 | 17.3 | 40.0 | 21.8 | 493 | 35.7 | 11.6 | 48.7 | 333 | 39.6 | 10.2 | 47.4 |
| Gemini-3 Flash | 221 | 24.0 | 4.1 | 56.6 | 15.4 | 411 | 5.1 | 1.5 | 89.1 | 222 | 13.1 | 3.6 | 74.3 |
| GPT-5.3 codex | 167 | 16.2 | 34.7 | 31.1 | 18.0 | 461 | 47.3 | 0.2 | 47.1 | 276 | 66.3 | 1.1 | 22.5 |
| GPT-5.4 | 158 | 17.1 | 53.2 | 27.2 | 2.5 | 459 | 31.4 | 14.4 | 45.8 | 285 | 38.9 | 29.5 | 21.1 |
| Model | st_d1 | st_d2 | mt_d1 | mt_d2 | |
| ( ) | ( ) | ( ) | ( ) | (pp) | |
| GPT-5.3 codex | |||||
| DeepSeek-chat | |||||
| Gemma-4-31B |
| Group | Share | Conversion | Zero-tool | |
| Read a direct -labelled section (header match) | 4,725 | 42.7% | 74.0% | 35.2% |
| Read the labelled content inside its parent section | 4,765 | 43.0% | 78.1% | 40.1% |
| Read no direct -labelled content | 1,582 | 14.3% | 49.3% | 31.5% |
| Pair | Discordant | |||
| MT | Gemini-3 Pro vs Claude Sonnet 4.6 | 226 | 20/3 | 0.0005 |
| MT | Claude Sonnet 4.6 vs GPT-5.4 | 193 | 24/16 | 0.27 |
| MT | GPT-5.4 vs GLM-5.1 | 227 | 36/22 | 0.09 |
| MT | GLM-5.1 vs DeepSeek-reasoner | 203 | 35/21 | 0.08 |
| MT | DeepSeek-reasoner vs MiniMax-M2.7 | 168 | 42/12 | 0.0001 |
| Closed, ST | Gemini-3 Pro vs Claude Sonnet 4.6 | 437 | 25/13 | 0.073 |
| Bigram | (pp) | Sig. count | Models |
| Productive (positive ) | |||
| read_section think | |||
| think submit_answer | |||
| list_sections read_section | |||
| get_references get_paper_info | |||
| get_paper_info get_references | |||
| Model | Top-1 trigram | Concentration |
| Closed Frontier | ||
| Gemini-3 Flash | gpi ls rs | |
| Gemini-3 Pro | ls rs rs | |
| Claude Sonnet 4.6 | gpi gr th | |
| GPT-5.4 | ls rs rs | |
| GPT-5.3 codex | gpi ls rs | |
| Model | Acc | NA | MaxT | Bud | Tok | Model | Acc | NA | MaxT | Bud | Tok |
| Closed Frontier | Open Frontier | ||||||||||
| Gemini-3 Pro | 89.1 | 4.5 | 0.0 | 0.0 | 4.5 | GLM-5.1 | 69.6 | 7.9 | 0.4 | 3.2 | 4.4 |
| GPT-5.3 codex | 82.5 | 1.0 | 0.0 | 0.7 | 0.3 | DeepSeek-R1 | 59.6 | 6.4 | 0.3 | 0.0 | 6.1 |
| GPT-5.4 | 79.0 | 5.2 | 0.0 | 4.9 | 0.3 | DeepSeek-V3 | 56.4 | 14.3 | 0.3 | 0.0 | 14.0 |
| Claude Opus 4.6 | 77.5 | 6.7 | 0.0 | 0.0 | 6.7 | Kimi-K2.5 | 44.1 | 9.9 | 2.0 | 2.4 | 5.5 |
| Claude Sonnet 4.6 | 76.6 | 4.4 | 0.6 | 0.0 | 3.8 | MiniMax-M2.7 | 41.3 | 7.2 | 1.3 | 4.8 | 1.1 |
| Model | GPQA-D | HLE | SWE-B Pro | T-Bench 2.0 | NL2Repo | BrowseComp | HLE+tools | AgentHop |
| Closed Frontier | ||||||||
| Gemini-3 Pro | 91.9 | 37.5 | 43.3 | 56.9 | — | 59.2 | — | 89.1 |
| GPT-5.3 codex | 92.6 | — | 56.8 | 77.3 | — | 77.3 | — | 82.5 |
| GPT-5.4 | 92.8 | 39.8 | 57.7 | 75.1 | 41.3 ∗ | 82.7 | 52.1 | 79.0 |
| Claude Opus 4.6 | 91.3 | 40.0 | 57.3 ∗ | 65.4 | 49.8 ∗ | 84.0 | 53.1 | 77.5 |
| Claude Sonnet 4.6 | 89.9 | 33.2 | — | 59.1 | — | 74.0 | 49.0 | 76.6 |
| Model | Metric | Run 0 | Run 1 | Run 2 | Mean | Std | Range |
| DeepSeek-chat | Accuracy | 0.564 | 0.553 | 0.542 | 0.553 | 0.009 | 0.022 |
| Section recall | 0.439 | 0.452 | 0.433 | 0.441 | 0.008 | 0.019 | |
| Conversion | 0.669 | 0.652 | 0.642 | 0.654 | 0.011 | 0.027 | |
| Tokens (k) | 156.7 | 156.1 | 155.1 | 156.0 | 0.7 | 1.6 | |
| Gemma-4-31B | Accuracy | 0.648 | 0.585 | 0.581 | 0.605 | 0.031 | 0.067 |
| Section recall | 0.275 | 0.242 | 0.251 | 0.256 | 0.014 | 0.033 |
| Model | Zero | Full | Model | Zero | Full | ||
| Closed Frontier | Open Frontier | ||||||
| Gemini-3 Pro | 73.4 | 89.1 | GLM-5.1 | 37.6 | 69.6 | ||
| GPT-5.3 codex | 48.6 | 82.5 | DeepSeek-reasoner | 25.0 | 59.5 | ||
| GPT-5.4 | 59.3 | 79.0 | DeepSeek-chat | 18.3 | 56.4 | ||
| Claude Opus 4.6 | 55.2 | 77.5 | Kimi-K2.5 | 12.1 | 44.1 | ||
| Claude Sonnet 4.6 | 48.5 | 76.6 | MiniMax-M2.7 | 14.5 | 41.3 | ||
| Axis | Stratum | Closed | Open | Comp.-eff. | |
| Citations | Q1 (1–224) | 252 | 57.5 | 21.7 | 19.2 |
| Q2 (226–585) | 253 | 59.2 | 22.2 | 17.8 | |
| Q3 (586–1,554) | 253 | 57.1 | 19.9 | 16.5 | |
| Q4 (1,554–54,992) | 253 | 54.4 | 22.2 | 16.7 | |
| Year | 2022 | 72 | 53.5 | 18.3 | 19.6 |
| 2023 | 355 | 54.7 | 19.7 | 15.9 |
| Tool | Cost | Input | Returns |
| get_paper_info | 1 | paper_id | Title, abstract, year, and venue of the paper. |
| get_references | 1 | paper_id | List of cited papers with their IDs, titles, years, and the citation context (the sentence in which they were cited). |
| list_sections | 1 | paper_id | Section headers with alphabet aliases and per-section character counts, for example (A) Introduction (5869 chars) . Aliases may be passed to read_section in place of a full header, and the counts let an agent price a 5-point read before paying for it. |
| search_papers | 1 | query (str), top_k ( , default 5) | Top- matching papers from the available pool, each with title and abstract. |
| read_section | 5 | paper_id , section | Full text of the requested section. The section argument accepts the full header, the alphabet alias from list_sections , or a unique case-insensitive substring. |
| think | 0 | thought (str) | Acknowledgement; the thought is appended to the trajectory transcript and remains available for the agent’s later reasoning steps. |