Multi-hop retrieval-augmented generation requires a retriever to connect evidence distributed across documents while preserving a concise, faithful generation context. Existing indexes leave two complementary gaps: chunk-based RAG can break cross-passage evidence chains, whereas an unlabeled pairwise projection without generating-topic provenance cannot jointly preserve topic-level co-participation and per-occurrence entity descriptions. We propose STITCH-RAG, a hypergraph-based framework with three coupled components. First, a semi-merged topic hypergraph encodes multi-entity co-participation as topic-summary hyperedges while retaining per-chunk entity states linked by canonical-name equivalence. Second, spatio-temporal influence bridging propagation (STIBP) combines topic-space propagation with deterministic chunk-index linkage across name-equivalent states under frequency-adaptive decay. Third, continuous STIBP scores replace binary entity-match seeds in localized Personalized PageRank (PPR). We characterize the condition under which this prior assigns more PPR mass to ground-truth evidence than a binary prior. Under the reported protocol, STITCH-RAG attains the highest reported Contain-Acc and LLM-Acc point estimates among the compared methods on HotpotQA and 2WikiMultiHopQA, and higher Recall@8 than the methods included in the standardized retrieval comparison. Results on the mixed-domain benchmark remain auxiliary preference-based evidence because only LLM-judged accuracy is available.
Figures & tables
Figure 1: Representation trade-offs for an example query about the birthplace of Film X’s director. (a) Chunk-based RAG can return locally similar but disconnected passages. (b) The illustrated unlabeled pairwise/full-merge representation retains the path through Alice Smith, but its single static entity node mixes the evidence-bearing birthplace context with unrelated contexts such as awards. The panel depicts one unlabeled full-merge representation, not all pairwise graphs. (c) STITCH-RAG preserves the two local Alice Smith states and links them by name equivalence, allowing query-conditioned propagation through topic hyperedges and across chunk-specific entity states.
Figure 2: Framework of STITCH-RAG: (1) document chunking; (2) semi-merged topic hypergraph construction, with topic-summary hyperedges and entity-description nodes linked by ϕ ; (3) STIBP through spatial and temporal channels; (4) influence-guided chunk scoring as a continuous PPR teleportation prior; and (5) top- k chunk retrieval for answer generation.
Method
HotpotQA
2Wiki
Mix
Contain-Acc
LLM-Acc
EM
Contain-Acc
LLM-Acc
EM
LLM-Acc
Zero-shot
0.422
0.451
0.300
0.505
0.398
0.345
0.269
Standard-RAG
0.700
0.702
0.460
0.689
0.617
0.467
0.669
HippoRAG
0.690
0.835
0.609
0.555
0.575
0.476
0.823
Cog-RAG
0.822
0.843
0.562
0.768
0.700
0.554
0.831
Hyper-RAG
0.735
0.808
0.511
0.785
0.745
0.541
0.808
Table 1: Answer accuracy on HotpotQA, 2Wiki, and Mix. Best result per column bolded. Only STITCH-RAG was repeated end to end; its entries average three runs on the same fixed question samples, with all run-level standard deviations below 0.004. Baseline run-level variances are unavailable. The table therefore ranks reported point estimates but does not establish statistical superiority.
Method
HotpotQA
2Wiki
Standard-RAG
0.612
0.587
LightRAG
0.741
0.693
LinearRAG
0.723
0.712
STITCH-RAG
0.798
0.769
Table 2: Passage-level Recall@8 on HotpotQA and 2Wiki, defined as ∣Gq∩Rq8∣/∣Gq∣ where Gq is the ground-truth supporting chunk set and Rq8 is the top-8 retrieved set. Mix is excluded because supporting-fact annotations are unavailable. Only methods with publicly available retrieval outputs under the standardized embedding protocol are included.
Variant
HotpotQA LLM-Acc
2Wiki LLM-Acc
Mix LLM-Acc
Full STITCH-RAG
0.895
0.861
0.884
w/o STIBP
0.823
0.707
0.800
w/o PPR
0.825
0.793
0.854
w/o all (dense retrieval)
0.761
0.684
0.672
w/o spatial bridging
0.838
0.742
0.814
w/o index-proximity bridging
0.840
0.748
0.815
Table 3: Ablation of STITCH-RAG (three-run LLM-Acc average). The upper block removes STIBP/PPR or uses dense retrieval; the lower block removes one STIBP channel or replaces continuous priors with binary initialization.
Method
Time (s)
Token Consumption
LLM-Acc
Indexing
Retrieval
Idx. Prompt
Idx. Completion
Query Prompt
Query Completion
HippoRAG
1706.91
39.56
1382281
1321272
55845.91
3774.76
0.823
Cog-RAG
19109.94
126.88
3082137
7988713
36840.47
13246.30
0.831
Hyper-RAG
70414.64
52.20
4652473
7264711
18557.08
5521.35
0.808
LightRAG
89882.40
75.87
6588007
8685853
29795.00
8148.52
0.877
LinearRAG
633.37
23.00
0
0
6395.72
261.50
0.762
Table 4: Efficiency and LLM-Acc on Mix. Indexing costs are one-time; query-stage tokens include answer generation and method-specific query-time LLM calls under identical hardware and API conditions.
δ
Mix
0.3
0.861
0.4
0.815
0.5
0.884
0.6
0.838
0.7
0.861
Table 5: Post-hoc Mix LLM-Acc sensitivity to δ . The bold row denotes the value preselected on HotpotQA, not a value selected on Mix. All other hyperparameters are fixed.
λ
Mix
0.1
0.823
0.3
0.884
0.5
0.846
0.7
0.846
0.9
0.861
Table 6: Post-hoc Mix LLM-Acc sensitivity to λ . The bold row denotes the value preselected on HotpotQA, not a value selected on Mix. All other hyperparameters are fixed.
k
5
6
7
8
9
10
LLM-Acc
0.823
0.853
0.861
0.884
0.869
0.876
Table 7: LLM-Acc under varying top- k retrieval count on the 200-question HotpotQA development set. The bold column denotes the selected value. All other hyperparameters are fixed at their defaults.
Figure 3: Pairwise answer-quality win rates of STITCH-RAG against five baselines on HotpotQA, 2Wiki, and Mix, judged by Qwen-Max on comprehensiveness (coverage of relevant details), diversity (variety of useful perspectives), and empowerment (utility for informed judgment). The dashed circle marks the 50% win-rate line. Values outside it favor STITCH-RAG.
Strategy
HotpotQA
2Wiki
Mix
Full-Merge
0.873
0.839
0.807
Semi-Merge
0.895
0.861
0.884
No-Merge
0.879
0.844
0.831
Table 8: LLM-Acc point estimates for one implementation of each merge strategy under a fixed retrieval pipeline. The diagnostic does not evaluate run-level variance, retrieval recall, graph compression, statistical significance, or all possible implementations of the three strategies.
Decay Function
LLM-Acc
exp(−αΔt) (pure exponential)
0.853
σ(−αΔt+b) (sigmoid gate)
0.815
exp(−tanh(αΔt)) (ours)
0.884
Table 9: LLM-Acc on Mix under three temporal decay functions, with all other pipeline components held fixed (same hypergraph, same PPR initialization, same α=nu/nˉ ).
Multi-hop question answering (QA) requires reasoning across multiple documents, yet existing retrieval-augmented generation (RAG) approaches address this either through graph-based methods requiring additional online processing or iterative multi-step reasoning. We present IndexRAG, a novel approach that shifts cross-document reasoning from online inference to offline indexing. IndexRAG identifies bridge entities shared across documents and generates bridging facts as independently retrievable units, requiring no additional training or fine-tuning. Experiments on three widely-used multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA, MuSiQue) show that IndexRAG improves F1 over Naive RAG by 4.6 points on average, while requiring only single-pass retrieval and a single LLM call at inference time. When combined with IRCoT, IndexRAG achieves the best average performance among all evaluated methods, including graph-based baselines such as HippoRAG2 and FastGraphRAG, while relying on a flat vector index. Our code is available at https://github.com/Continuum-AI-Corp/IndexRAG .
Zhenghua Bao, Yi Shi
Continuum AI Shanghai, China · Continuum AI San Francisco, CA, USA
While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning. Graph-based methods address this by constructing knowledge graphs offline, but they often fragment semantics, incur high maintenance, and complicate incremental updates. We propose SAG (SQL-Retrieval Augmented Generation), a structured retrieval architecture that organizes documents into an event-entity index without building a global knowledge graph. SAG represents each chunk as a semantically complete event paired with its entities, forming a latent hyperedge that preserves n-ary relations without decomposing them into triples. At query time, SAG treats shared entities as join keys to connect related chunks. This dynamically yields a query-scoped neighborhood of events, and yet every piece of evidence remains the original chunk throughout. Experiments on HotpotQA, 2WikiMultiHopQA, and MuSiQue show that SAG achieves the best retrieval and end-to-end QA performance on every benchmark, with gains that widen as reasoning-chain complexity increases. On MuSiQue, where multi-hop evidence chaining is most demanding, SAG reaches 80.36% Recall@5, outperforming the strongest baseline by 11.52 points. This work paves the way for knowledge infrastructure that enables LLM agents to retrieve and reason over continually growing organizational knowledge.
Graph-based retrieval-augmented generation supports multi-hop retrieval by organizing corpus information into graphs. However, existing relation-free graph retrieval methods rely primarily on query-sentence similarity to search for evidence. This can exclude useful bridging evidence with low query similarity and activate incidental entities unrelated to the reasoning chain. In this paper, we propose a simple and effective approach called NexusRAG, which augments the relation-free Tri-Graph with a corpus-level entity neighborhood structure derived from joint entity co-occurrence and semantic similarity. NexusRAG employs this structure to guide two complementary propagation paths: neighborhood-constrained semantic propagation through sentences identifies the query-relevant entity frontier, while direct structural propagation between neighboring entities expands that frontier to structurally related entities. The propagated entity weights also inform neighborhood-aware passage initialization for Personalized PageRank. Experiments on three multi-hop QA benchmarks and a domain-specific subset of GraphRAG-Bench show that NexusRAG consistently outperforms existing approaches. On the GraphRAG-Bench subset, NexusRAG achieves the highest evidence recall in all question categories, exceeding baselines by 4.2-8.1 points. The implementation code is available at https://github.com/Jacob-biu/NexusRAG.
Baoxian Liu, Tong Wei
College of Software Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · School of Computer Science and Engineering, Southeast University, Nanjing 210096, China