Corpus-Guided Dual-Path Propagation for Graph Retrieval-Augmented Generation
Authors: Baoxian Liu, Tong Wei
Organizations: College of Software Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · School of Computer Science and Engineering, Southeast University, Nanjing 210096, China
Graph-based retrieval-augmented generation supports multi-hop retrieval by organizing corpus information into graphs. However, existing relation-free graph retrieval methods rely primarily on query-sentence similarity to search for evidence. This can exclude useful bridging evidence with low query similarity and activate incidental entities unrelated to the reasoning chain. In this paper, we propose a simple and effective approach called NexusRAG, which augments the relation-free Tri-Graph with a corpus-level entity neighborhood structure derived from joint entity co-occurrence and semantic similarity. NexusRAG employs this structure to guide two complementary propagation paths: neighborhood-constrained semantic propagation through sentences identifies the query-relevant entity frontier, while direct structural propagation between neighboring entities expands that frontier to structurally related entities. The propagated entity weights also inform neighborhood-aware passage initialization for Personalized PageRank. Experiments on three multi-hop QA benchmarks and a domain-specific subset of GraphRAG-Bench show that NexusRAG consistently outperforms existing approaches. On the GraphRAG-Bench subset, NexusRAG achieves the highest evidence recall in all question categories, exceeding baselines by 4.2-8.1 points. The implementation code is available at https://github.com/Jacob-biu/NexusRAG.
Figures & tables
Figure 1: Overview of NexusRAG. 1) Graph Construction. The corpus is indexed into a relation-free entity–sentence–passage Tri-Graph. 2) Neighbor Prior. A sparse corpus-level entity neighbor matrix W is constructed from co-occurrence and semantic similarity with rank and weight pruning. 3) Entity Propagation. Given a query q , neighbor-constrained semantic propagation identifies a query-relevant entity frontier, which structural propagation expands through W to obtain the activated entity set. 4) Passage Retrieval. Cumulative entity weights and neighbor-aware passage weights initialize Personalized PageRank for top- K passage retrieval.
Figure 2: Retrieval failure analysis of LinearRAG. (a) Bridge entities account for a substantial fraction of retrieval errors. (b) Query-dependent sentence mediation can both suppress necessary bridge evidence and over-activate incidental entities.
Method
HotpotQA
2Wiki
MuSiQue
Medical
Con.
LLM.
Avg.
Con.
LLM.
Avg.
Con.
LLM.
Avg.
LLM.
Direct Zero-shot LLM Inference
llama-8B
31.10
27.30
29.20
33.60
16.20
24.90
7.40
8.10
7.75
27.31
llama-13B
24.20
16.80
20.50
21.90
10.50
16.20
3.30
4.40
3.85
28.86
GPT-3.5-turbo
33.40
43.20
38.30
28.70
31.00
29.85
10.30
21.90
16.10
45.60
GPT-4o-mini
38.90
40.20
39.55
36.30
31.40
33.85
13.60
15.80
14.70
42.10
Table 1: Main results on four benchmarks across three generation backbones. All methods are evaluated with GPT-4o-mini; Qwen3.6-27B-FP8 and DeepSeek-V4-Flash provide additional comparisons between LinearRAG and NexusRAG. Bold and underline denote the highest and second-highest results, respectively. Columns report Contain-Acc. (%) (Con.), LLM-Acc. (%) (LLM.), and Avg. (%) (mean of Con. and LLM.). Only LLM-Acc. is used for the Medical dataset.
Method
Fact Retrieval
Complex Reasoning
Contextual
Creative Generation
Recall
Relevance
Recall
Relevance
Recall
Relevance
Recall
Relevance
Vanilla RAG (Top-5)
86.24
63.71
84.97
84.11
84.14
89.94
44.88
58.73
RAPTOR
85.40
69.38
89.70
53.20
88.86
58.73
72.70
52.71
E 2 GraphRAG
87.84
69.74
87.08
62.67
89.17
71.63
60.26
35.84
LightRAG
80.32
41.27
82.91
42.79
85.71
43.11
81.34
45.17
GFM-RAG
90.08
57.90
85.03
33.06
78.62
40.14
83.51
22.87
Table 2: Retrieval quality evaluation results (%) following the GraphRAG-Bench protocol across four question categories. Recall and relevance are reported for fact retrieval, complex reasoning, contextual understanding, and creative generation. The highest and second-highest results are shown in bold and underlined , respectively.
Method
Evaluator
Hotpot
2Wiki
MuSiQue
Med.
LinearRAG
GPT-4o-mini
69.5
65.0
38.2
65.3
DeepSeek-V4-Flash
78.1
73.8
42.8
60.5
Qwen3.6-27B-FP8
78.4
74.9
42.9
74.3
NexusRAG
GPT-4o-mini
72.9
68.7
41.0
73.91
DeepSeek-V4-Flash
80.2
74.9
44.6
63.06
Qwen3.6-27B-FP8
81.9
76.7
46.6
79.29
Table 3: LLM-Acc. (%) of the same GPT-4o-mini predictions re-judged by three evaluators. Highest results are in bold .
HotpotQA
2Wiki
MuSiQue
Med.
Variant
Con.
LLM.
Con.
LLM.
Con.
LLM.
LLM.
NexusRAG
72.2
88.7
81.4
86.0
47.1
58.4
73.96
w/o Neighbor Clamp
71.2
86.6
80.0
83.4
46.2
56.9
72.58
w/o Structural propagation
71.4
87.7
80.8
84.8
45.0
55.0
71.10
w/o Neighbor
69.8
85.6
79.8
82.8
45.7
54.6
70.81
w/o Neighbor-aware Init
71.3
87.6
80.9
84.9
46.1
56.3
72.50
Table 4: Ablation across the four benchmarks (Qwen3.6-27B-FP8).
Figure 3: Sensitivity study of the neighbor-mechanism parameters on the HotpotQA dataset.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Specification
GPU
NVIDIA RTX A6000
CPU
Intel(R) Xeon(R) Platinum 8260 CPU @ 2.30GHz
CUDA
12.8 (Driver 570.172.08)
Appendix
Table 5: Detailed machine configuration used in our experiments.
Dataset
maxT
δ
η
d
ωp
α
τ
κ
HotpotQA
3
0.4
1
0.5
0.05
0.5
0.5
5
2Wiki
3
0.4
1
0.5
0.05
0.5
0.5
5
MuSiQue
5
0.1
4
0.5
0.05
0.5
0.5
5
Medical
3
0.5
3
0.5
0.05
0.5
0.5
5
Appendix
Table 6: Hyperparameter configuration. maxT : max propagation hops; δ : pruning threshold; η : sentences per entity per iteration; d : PPR damping; ωp : passage node weight. The first group is fixed to LinearRAG’s reported values; α,τ,κ are the neighbor-mechanism parameters (the structural propagation uses the raw neighbor weight wij with no extra decay coefficient).
Dataset
τ
Avg. #neighbors
Zero-neighbor (%)
Per dataset at τ=0.5 (adopted)
HotpotQA
1.99
0.00
2Wiki
1.99
0.00
MuSiQue
2.01
0.00
Medical
2.29
0.00
Per threshold on HotpotQA
0.4
4.38
0.00
0.5
1.99
0.00
Appendix
Table 7: Neighbor graph statistics on the Qwen3.6-27B-FP8 backbone, after the top- κ=5 truncation (the effective graph used by propagation). Avg. #neighbors = average neighbors per entity in the effective graph; Zero-neighbor (%) = fraction of entities with no effective neighbors. The top half reports per-dataset statistics at the adopted threshold τ=0.5 ; the bottom half reports per-threshold statistics on HotpotQA.
Figure 4: Sensitivity to the fusion coefficient α and the neighbor threshold τ on the Qwen3.6-27B-FP8 backbone (HotpotQA; one parameter varied at a time, the other fixed at its adopted value).
Figure 5: Sensitivity to the neighbor cap κ .
Method
HotpotQA
2Wiki
K=2
K=5
K=2
K=5
Br
Ans
Br
Ans
Br
Ans
Br
Ans
LinearRAG
87.71
78.1
89.98
82.8
80.36
72.0
81.6
76.3
NexusRAG
89.98
78.7
93.49
84.8
81.05
72.7
83.13
76.7
Appendix
Table 8: Retrieval-level recall@ K (%). Bridge-Recall (Br) and Answer-Recall (Ans) at each depth K . Highest values are in bold .
Question
“What is unique about the forum an American poet, memoirist, and civil rights activist spoke at?”
Ground Truth
the oldest free public lecture series in the United States
Support Context
[“American poet, memoirist & civil-rights activist”] → “Maya Angelou” → “spoke at Ford Hall Forum”
[“Ford Hall Forum”] → “the oldest free public lecture series in the United States”
LinearRAG
Retrieved context:
1) ✕ “Cleanth Brooks (poet & critic)”: …with The Southern Review in 1935; won the Pulitzer Prize for fiction and for poetry …
2) ✕ “Grosvenor, Duke of Westminster”: …arts & charity patron; mentions Maya Angelou only in passing …
Appendix
Table 9: Case study (HotpotQA). LinearRAG retrieves topically related poet / civil-rights passages, one of which mentions Maya Angelou only in passing, but misses the bridge article Ford Hall Forum (which contains the answer sentence) and answers incorrectly, whereas NexusRAG’s structural propagation retains the Ford Hall Forum passage and answers correctly.
Question
“What is the date of birth of Archduke Karl Pius of Austria, Prince of Tuscany’s mother?”
Ground Truth
7 September 1868
Support Context
[“Archduke Karl Pius of Austria, Prince of Tuscany”] mother “Infanta Blanca of Spain”
[“Infanta Blanca of Spain”] date of birth “7 September 1868”
LinearRAG
Retrieved context:
1) ✓ hop-1 bridge (passage 349): “…archduke karl pius of austria, prince royal of hungary and bohemia, prince of tuscany (4 december 1909 – 24 december 1953) …he was the tenth and youngest child of archduke leopold salvator , prince of tuscany and infanta blanca of spain .”
2) ✕ Habsburg / film biography chunk (passage 348).
Appendix
Table 10: Case study (2WikiMultiHopQA). LinearRAG retrieves the hop-1 bridging passage (349) and correctly identifies the mother, Infanta Blanca of Spain , but its second-hop expansion follows the father Leopold Salvator , the entity that co-occurs with the mother in that very sentence, and returns his biography (630) instead; the mother’s own article (147), which states her date of birth, is never retrieved, and LinearRAG answers “not mentioned in the text”. NexusRAG returns the same hop-1 passage and retains passage 147, and answers correctly. Passage IDs are the retriever’s own indices; ✓ marks passages that carry a gold fact.
HotpotQA
2Wiki
MuSiQue
Medical
LinearRAG
Index(s)
1057.42
541.67
1109.22
179.86
Precompute(s)
0.00
0.00
0.00
0.00
Retrieval(s)
0.252
0.214
0.240
0.216
NexusRAG (ours)
Index(s)
991.31
549.20
1072.60
181.85
Appendix
Table 11: Indexing and retrieval timing on the four benchmarks (seconds).
Dataset
Method
Index(s)
Precompute(s)
Total Index (s)
Token ( ×106 )
Prompt
Completion
5M
HippoRAG
71019.97
N/A
71019.97
18.14
9.63
LinearRAG
4188.49
N/A
4188.49
0
0
NexusRAG (ours)
3949.79
106.93
4056.72
0
0
10M
HippoRAG
100245.49
N/A
100245.49
37.10
20.26
LinearRAG
8110.30
N/A
8110.30
0
0
Appendix
Table 12: Indexing cost on ATLAS-Wiki (5M and 10M tokens). Total Index (s) = base graph index ( Index(s) ) + neighbor precompute ( Precompute(s) ). LinearRAG has no neighbor precompute stage, so its total equals Index(s) ; NexusRAG adds the one-time neighbor construction. HippoRAG figures are measured in this work (Qwen3.6-27B-FP8 via vLLM).
Multi-hop question answering in retrieval-augmented gener?ation (RAG) often benefits from retrieving beyond the few candidates that will finally be read: narrow retrieval can miss an indispensable hop, while expanded retrieval introduces topical distractors. This challenge is not tied to a particu?lar knowledge-base format. Candidate pools may come from standalone retrievers, standard RAG backends, or graph-based retrieval pipelines. What is needed is a query-aware selection layer that can use relational structure to filter candidates be?fore generation. PAGE-RAG addresses this setting by using a graph as a temporary selection structure, rather than assum?ing a graph-structured knowledge base. It builds a query-local graph over retrieved candidates, records why candidates are connected, and treats each connection as a support hypothe?sis rather than support itself. We identify the resulting failure mode as a connectivity-support gap: connected candidates do not necessarily support the answer. We propose PAGE-RAG, a Provenance-Aware Graph Evidence promotion method that scores candidate paths with relevance, source-tracing meta?data, specificity, hubness, noise, and coherence signals, and applies minimal sufficient selection to promote supporting facts into a compact reader context. PAGE-RAG can serve as a complete retrieval-to-reading pipeline, and the same promo?tion stage can be inserted after existing retrieval or RAG sys?tems without replacing their upstream retrieval logic. Across three multi-hop QA benchmarks under the same final bud?get, PAGE-RAG improves support F1 and answer F1 by 10.4 and 3.3 points on a weighted average over a strong retriever. As a plug-in, PAGE-RAG further improves all reported RAG backends, including reasoning-oriented, compression-based, graph-based, and document/chunk-level systems.
Haokun Deng, Xunkai Li, Hongchao Qin +1
Department of Computer Science, Beijing Institute of Technology
GraphRAG extends retrieval-augmented generation by organizing corpora as explicit knowledge graphs, enabling graph-based retrieval for complex question answering. However, existing frameworks extract entities and relations within individual chunks, leaving cross-chunk relations -- those whose evidence spans multiple passages -- systematically absent from the index. Exhaustive LLM-based recovery of such relations is impractical due to the combinatorial explosion of chunk combinations. We present CrossAug, a GNN-guided CROSS-Chunk Graph AUGmentation method that enriches GraphRAG indices with cross-chunk relational structure as an offline step before query-time retrieval. CrossAug derives training supervision through self-supervised graph corruption, uses a topology-aware GNN to score subgraphs for missingness, and applies evidence-grounded LLM completion only to selected high-scoring regions. Experiments on three LLM-based GraphRAG frameworks across four multi-hop and long-document QA benchmarks demonstrate that CrossAug consistently improves performance, confirming the benefit of cross-chunk graph augmentation for retrieval-based question answering. Our code is available at https://github.com/DonFinliani/CrossAug.
Jiaming Zhang, Yibo Zhao, Jing Yu +2
School of Data Science and Engineering, East China Normal University
Retrieval-Augmented Generation (RAG) mitigates hallucinations in Large Language Models (LLMs) by grounding the generation process on external knowledge. However, standard RAG approaches struggle with multi-hop reasoning. While recent graph-based RAG methods improve the retrieval of interconnected chunks, they often rely on computationally expensive and error-prone LLM-based extraction pipelines. To address these issues, we propose TIGRAG (Token-Induced GraphRAG), an efficient graph-augmented RAG framework based on a token co-occurrence Knowledge Graph. TIGRAG directly models topological relationships between tokens using sliding-window co-occurrence statistics, thus enabling scalable graph construction. During inference, it combines graph-based semantic expansion and neural reranking to retrieve interconnected evidence for multi-hop reasoning. Specifically, it introduces an iterative entity-driven retrieval strategy that progressively expands the query using bridging entities extracted from previously retrieved contexts. We evaluated TIGRAG on three widely adopted multi-hop Question Answering (QA) benchmarks. Experimental results demonstrated that our framework consistently outperforms dense retrieval and graph-based RAG methods in both retrieval and downstream QA tasks, while substantially reducing indexing time, inference latency, and prompt footprint.
Gianluca Bonifazi, Christopher Buratti, Michele Marchetti +5
aUniversità Politecnica delle Marche, Ancona, Italy · bUniversità di Modena e Reggio Emilia, Modena, Italy