Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
Figures & tables
Figure 1 : Hierarchical indexing (offline). The corpus is encoded using a Matryoshka embedding model and organized from the bottom up into a hierarchical DAG via HDBSCAN with overlapping cluster assignments. The internal nodes are progressively coarser cluster centroids stored at shorter Matryoshka prefixes. The leaf layer retains full-dimensional document embeddings. Named entities are extracted from each document and stored with the embeddings. This allows for retrieval using both semantic and entity-level information.
Figure 2 : Entity-budgeted retrieval (online). Retrieval involves an iterative, top-down traversal of the hierarchy. Starting with the coarsest level, the query is scored against cluster centroids. At each level, the search space narrows to the most promising cluster until reaching the document level. The traversal score considers both similarity to the original query and similarity to the cumulative query. The latter incorporates previously retrieved documents to guide multi-hop reasoning while limiting query drift. At the leaves, candidates are re-ranked using a score that blends semantic similarity and entity overlap to prioritize documents that are both relevant and entity-consistent.
HotpotQA
2Wiki
MuSiQue
K
EM ↑
F1 ↑
Tok ↓
EM ↑
F1 ↑
Tok ↓
EM ↑
F1 ↑
Tok ↓
5
51.40
65.17
325
42.60
49.85
298
20.60
29.97
354
10
59.80
70.75
676
50.80
59.49
608
25.00
36.40
716
Table 1 : Values of Exact Match (EM), token-level F1 (F1) and LLM context size (Tok) obtained by MatRAG on the three benchmarks for K=5 and K=10 . The values represent the results computed on the validation set consisting of 500 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
p
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
1
56.60
68.66
73.80 88.70
47.20
56.74
63.45 78.35
22.80
33.64
42.60 55.87
2
59.80
70.75
75.80 91.10
50.80
59.49
66.05 81.85
25.00
36.40
43.78 57.40
3
59.60
70.67
75.60 90.30
50.00
59.19
65.35 80.25
24.00
35.66
42.92 57.22
4
59.00
70.36
74.20 89.50
49.00
58.19
65.05 79.75
23.20
33.52
42.37 56.98
Table 2 : Values of EM, F1, Recall@2 (R@2) and Recall@5 (R@5) obtained by MatRAG on the three benchmarks for p ranging from 1 to 4. The values represent the results computed on the validation set consisting of 500 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
α
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
0.00
55.80
67.52
74.70 86.80
50.40
58.25
64.70 81.15
24.00
35.07
43.80 56.70
0.25
57.20
68.48
74.30 90.30
50.20
58.79
65.80 80.70
23.40
34.79
43.03 56.22
0.50
59.80
70.75
75.80 91.10
50.80
59.49
66.05 81.85
25.00
36.40
43.78 57.40
0.75
58.80
69.81
76.40 90.20
48.40
56.74
64.70 76.60
19.80
30.20
43.70 55.75
1.00
58.00
67.62
74.30 88.90
45.80
53.31
63.45 73.80
18.60
29.83
42.61 53.37
Table 3 : Values of EM, F1, R@2 and R@5 obtained by MatRAG on the three benchmarks for α set to 0.00, 0.25, 0.50, 0.75, and 1.00. The values represent the results computed on the validation set consisting of 500 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
R@2 ↑
R@5 ↑
R@2 ↑
R@5 ↑
R@2 ↑
R@5 ↑
HippoRAG2
56.40 ± .42
82.95 ± .57
62.50 ± .86
77.80 ± .31
34.60 ± .44
50.30 ± .39
LightRAG
68.05 ± .37
84.25 ± .29
58.90 ± .82
68.70 ± .97
34.80 ± .46
48.75 ± .40
NaiveRAG
70.55 ± .54
84.10 ± .28
63.40 ± .58
69.55 ± .85
37.50 ± .43
50.35 ± .37
MatRAG
72.65 ± .89
89.30 ± .46
63.08 ± .65
79.27 ± .81
42.93 ± .40
57.08 ± .33
Table 4 : Retrieval performance on the three multi-hop benchmarks measured by Recall@2 (R@2) and Recall@5 (R@5). The results are reported only for systems that retrieve corpus documents directly. The values represent the means and standard deviations computed over 5 samples, each of which contains 1,000 different questions. For each column, the optimal value is shown in bold and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
EM ↑
F1 ↑
EM ↑
F1 ↑
EM ↑
F1 ↑
HippoRAG2
52.40 ± .58
64.10 ± .62
49.85 ± .54
55.20 ± .73
26.20 ± .80
33.55 ± .39
RAPTOR
39.50 ± .81
55.35 ± .93
32.10 ± .74
37.90 ± .41
14.15 ± .48
23.85 ± .93
KGP
23.20 ± .66
33.10 ± .44
12.05 ± .89
13.80 ± .70
11.75 ± .82
18.90 ± .61
ToG
21.30 ± .72
27.45 ± .85
15.20 ± .91
17.75 ± .76
6.40 ± .69
9.25 ± .59
GraphRAG
30.65 ± .63
40.60 ± .78
8.25 ± .56
9.10 ± .51
1.10 ± .32
2.30 ± .38
Table 5 : QA performance on the three multi-hop benchmarks, measured by EM and F1 computed against the gold answers. The values represent the means and standard deviations computed over 5 samples, each of which contains 1,000 different questions. For each column, the optimal value is shown in bold and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
HippoRAG2
312,441 ± 1,842
123,187 ± 1,253
321,905 ± 1,976
RAPTOR
28,763 ± 412
16,724 ± 318
36,204 ± 487
KGP
1,941 ± 63
1,478 ± 51
4,213 ± 74
ToG
141,820 ± 934
69,340 ± 721
148,573 ± 1,102
GraphRAG
661,392 ± 2,341
75,614 ± 843
197,840 ± 1,587
LightRAG
728,471 ± 2,813
347,605 ± 1,934
598,317 ± 2,645
Table 6 : Indexing time (Idx) in seconds across the three benchmarks. The values represent the means and standard deviations computed over 5 independent runs of the indexing procedure on the full corpus of each benchmark. For each column, the optimal value is shown in bold, and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
Res ↓
Tok ↓
Res ↓
Tok ↓
Res ↓
Tok ↓
HippoRAG2
76.84 ± 2.83
591 ± 12
80.73 ± 3.91
624 ± 14
83.12 ± 2.97
671 ± 15
RAPTOR
15.23 ± .41
854 ± 18
14.89 ± .38
843 ± 17
15.91 ± .44
856 ± 19
KGP
31.40 ± 1.57
2,478 ± 31
26.55 ± 1.49
359 ± 9
38.04 ± 1.63
2,298 ± 28
ToG
56.81 ± 1.74
124 ± 6
40.67 ± 1.62
112 ± 5
49.93 ± 1.68
121 ± 6
GraphRAG
15.34 ± .39
3,812 ± 42
16.08 ± .43
871 ± 19
16.20 ± .45
1,398 ± 24
Table 7 : Efficiency comparison across the three benchmarks in terms of response time (Res), measured in seconds, and Tok. The values represent the means and standard deviations computed over 5 samples, each composed of 1,000 different questions. For each column, the optimal value is shown in bold and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
βt
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
1.00
59.22
72.25
72.10 88.70
50.90
56.96
62.05 76.75
25.70
35.92
41.30 55.52
K∣Rt−1∣
60.50
73.66
72.65 89.30
52.70
63.42
63.08 79.27
27.40
37.89
42.93 57.08
0.00
58.70
71.24
72.05 87.72
51.24
62.31
62.58 78.81
26.42
36.17
41.17 56.31
Table 8 : Values of EM, F1, R@2 and R@5 obtained by MatRAG on the three benchmarks for βt set to 1.00, K∣Rt−1∣ and 0.00. The values represent the means computed over 5 samples, each composed of 1,000 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
l
ml
DB ↓
Sl ↑
DB ↓
Sl ↑
DB ↓
Sl ↑
1
768
0.8667
0.4484
0.7313
0.5512
0.8362
0.4492
64
0.8630
0.4628
0.7198
0.5536
0.7557
0.5406
2
768
0.8342
0.4663
0.7503
0.5085
0.7573
0.4746
128
0.7616
0.4749
0.7389
0.5059
0.7661
0.4860
3
768
0.7428
0.5206
0.6608
0.5390
0.7330
0.4876
Table 9 : Values of the Davies-Bouldin index (DB) and the Silhouette score (Sl) for the level l and embedding dimension across benchmarks. The ml column indicates the embedding dimension used for clustering with full-dimensional and Matryoshka embeddings. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
Embedding
EM ↑
F1 ↑
EM ↑
F1 ↑
EM ↑
F1 ↑
Full
59.28
72.47
49.60
60.29
26.00
36.23
Matryoshka
60.50
73.66
52.70
63.42
27.40
37.89
Table 10 : Ablation study on QA performance, as measured by EM and F1, on the three benchmarks. The Embedding column indicates whether the system uses Matryoshka truncated (Matryoshka) or full-dimensional (Full) embeddings during retrieval. Values represent the results computed over 5 samples, composed of 1,000 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
Embedding
R@2 ↑
R@5 ↑
Res ↓
R@2 ↑
R@5 ↑
Res ↓
R@2 ↑
R@5 ↑
Res ↓
Full
70.86
86.72
3.10
62.12
76.88
4.04
41.86
55.80
4.82
Matryoshka
72.65
89.30
1.85
63.08
79.27
2.49
42.93
57.08
2.99
Table 11 : Ablation study on retrieval and efficiency performance, as measured by R@2, R@5, and Res (computed as the sum of retrieval and answer generation time) on the three benchmarks. The Embedding column indicates whether the system uses Matryoshka truncated (Matryoshka) or full-dimensional (Full) embeddings during retrieval. Values represent the results computed over 5 samples composed of 1,000 different questions. For each column, the optimal value is shown in bold.
Multi-hop question answering is a fundamental challenge in retrieval-augmented generation (RAG), because deriving an answer requires integrating dispersed evidence. Iterative RAG (iRAG) is widely used for this challenge, but existing methods have two limitations. First, most methods still support each reasoning step with single-granularity evidence, making it difficult to balance information density and contextual noise. Second, existing methods often answer the original question only after aggregating evidence retrieved across intermediate steps, so redundant evidence and intermediate retrieval errors may accumulate and degrade the final answer. To address these limitations, we propose MEGRAG, an answer-aware framework that represents multi-hop reasoning as a path-structured multi-granular evidence graph. Offline, MEGRAG links passages to their sentences and extracted triples through a cross-granularity index. Online, it retrieves passages for the current query and selects aligned evidence, starting with compact triples and adding sentence or passage context as needed. MEGRAG uses the resulting intermediate answer and prior reasoning to decide whether the Initial Query has been resolved. If not, it identifies the missing information and formulates a focused next query; otherwise, it stops retrieval and returns the answer. Extensive experiments demonstrate consistent gains over a diverse set of RAG baselines.
Multi-hop question answering (QA) requires reasoning across multiple documents, yet existing retrieval-augmented generation (RAG) approaches address this either through graph-based methods requiring additional online processing or iterative multi-step reasoning. We present IndexRAG, a novel approach that shifts cross-document reasoning from online inference to offline indexing. IndexRAG identifies bridge entities shared across documents and generates bridging facts as independently retrievable units, requiring no additional training or fine-tuning. Experiments on three widely-used multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA, MuSiQue) show that IndexRAG improves F1 over Naive RAG by 4.6 points on average, while requiring only single-pass retrieval and a single LLM call at inference time. When combined with IRCoT, IndexRAG achieves the best average performance among all evaluated methods, including graph-based baselines such as HippoRAG2 and FastGraphRAG, while relying on a flat vector index. Our code is available at https://github.com/Continuum-AI-Corp/IndexRAG .
Zhenghua Bao, Yi Shi
Continuum AI Shanghai, China · Continuum AI San Francisco, CA, USA
Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA), which requires reasoning over evidence from multiple documents. Current multi-hop RAG methods generally focus on either query-side task decomposition or corpus-side knowledge graph construction. Despite their progress, these methods still struggle to achieve satisfactory performance on complex multi-hop QA tasks. To this end, we propose ConRAG, a consensus-driven multi-view RAG framework that effectively boosts LLMs on complex multi-hop QA. The core of ConRAG is to systematically optimize both the query and corpus sides and to leverage multi-view evidence (relation, entity, and text signals) for more accurate retrieval. Extensive experiments on three multi-hop QA benchmarks show that ConRAG consistently outperforms all baselines by a clear margin, e.g., up to +26.9% average performance gains over vanilla RAG, and enables Gemma-4-31B to achieve a new state-of-the-art record on the challenging MuSiQue benchmark.
Yikai Zhu, Kunfeng Chen, Qihuang Zhong +2
School of Computer Science, Wuhan University Wuhan, China