cs.AISep 29, 2026

REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

Authors: Jiawen Tao, Xiaokun Yuan, Yaoming Li, Chenxu Liu, Mengzhou Wu, Tong Yang, Maxm Pan

Organizations: Hunyuan Team, Tencent · Peking University

Abstract

Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

    Jul 12, 2026JungMin Yun, JuneHyoung Kwon, YoungBin KimMulti-Hop ReasoningLLM Reasoning Strategies

  2. DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

    Sep 2, 2026Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park +5Multimodal DocumentsMulti-Hop Reasoning

  3. AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG

    Feb 22, 2026Qijie You, Wenkai Yu, Wentao ZhangAgentic BenchmarksAgentic Reasoning