Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
Figures & tables
Figure 1 : Hierarchical indexing (offline). The corpus is encoded using a Matryoshka embedding model and organized from the bottom up into a hierarchical DAG via HDBSCAN with overlapping cluster assignments. The internal nodes are progressively coarser cluster centroids stored at shorter Matryoshka prefixes. The leaf layer retains full-dimensional document embeddings. Named entities are extracted from each document and stored with the embeddings. This allows for retrieval using both semantic and entity-level information.
Figure 2 : Entity-budgeted retrieval (online). Retrieval involves an iterative, top-down traversal of the hierarchy. Starting with the coarsest level, the query is scored against cluster centroids. At each level, the search space narrows to the most promising cluster until reaching the document level. The traversal score considers both similarity to the original query and similarity to the cumulative query. The latter incorporates previously retrieved documents to guide multi-hop reasoning while limiting query drift. At the leaves, candidates are re-ranked using a score that blends semantic similarity and entity overlap to prioritize documents that are both relevant and entity-consistent.
HotpotQA
2Wiki
MuSiQue
K
EM ↑
F1 ↑
Tok ↓
EM ↑
F1 ↑
Tok ↓
EM ↑
F1 ↑
Tok ↓
5
51.40
65.17
325
42.60
49.85
298
20.60
29.97
354
10
59.80
70.75
676
50.80
59.49
608
25.00
36.40
716
Table 1 : Values of Exact Match (EM), token-level F1 (F1) and LLM context size (Tok) obtained by MatRAG on the three benchmarks for K=5 and K=10 . The values represent the results computed on the validation set consisting of 500 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
p
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
1
56.60
68.66
73.80 88.70
47.20
56.74
63.45 78.35
22.80
33.64
42.60 55.87
2
59.80
70.75
75.80 91.10
50.80
59.49
66.05 81.85
25.00
36.40
43.78 57.40
3
59.60
70.67
75.60 90.30
50.00
59.19
65.35 80.25
24.00
35.66
42.92 57.22
4
59.00
70.36
74.20 89.50
49.00
58.19
65.05 79.75
23.20
33.52
42.37 56.98
Table 2 : Values of EM, F1, Recall@2 (R@2) and Recall@5 (R@5) obtained by MatRAG on the three benchmarks for p ranging from 1 to 4. The values represent the results computed on the validation set consisting of 500 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
α
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
0.00
55.80
67.52
74.70 86.80
50.40
58.25
64.70 81.15
24.00
35.07
43.80 56.70
0.25
57.20
68.48
74.30 90.30
50.20
58.79
65.80 80.70
23.40
34.79
43.03 56.22
0.50
59.80
70.75
75.80 91.10
50.80
59.49
66.05 81.85
25.00
36.40
43.78 57.40
0.75
58.80
69.81
76.40 90.20
48.40
56.74
64.70 76.60
19.80
30.20
43.70 55.75
1.00
58.00
67.62
74.30 88.90
45.80
53.31
63.45 73.80
18.60
29.83
42.61 53.37
Table 3 : Values of EM, F1, R@2 and R@5 obtained by MatRAG on the three benchmarks for α set to 0.00, 0.25, 0.50, 0.75, and 1.00. The values represent the results computed on the validation set consisting of 500 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
R@2 ↑
R@5 ↑
R@2 ↑
R@5 ↑
R@2 ↑
R@5 ↑
HippoRAG2
56.40 ± .42
82.95 ± .57
62.50 ± .86
77.80 ± .31
34.60 ± .44
50.30 ± .39
LightRAG
68.05 ± .37
84.25 ± .29
58.90 ± .82
68.70 ± .97
34.80 ± .46
48.75 ± .40
NaiveRAG
70.55 ± .54
84.10 ± .28
63.40 ± .58
69.55 ± .85
37.50 ± .43
50.35 ± .37
MatRAG
72.65 ± .89
89.30 ± .46
63.08 ± .65
79.27 ± .81
42.93 ± .40
57.08 ± .33
Table 4 : Retrieval performance on the three multi-hop benchmarks measured by Recall@2 (R@2) and Recall@5 (R@5). The results are reported only for systems that retrieve corpus documents directly. The values represent the means and standard deviations computed over 5 samples, each of which contains 1,000 different questions. For each column, the optimal value is shown in bold and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
EM ↑
F1 ↑
EM ↑
F1 ↑
EM ↑
F1 ↑
HippoRAG2
52.40 ± .58
64.10 ± .62
49.85 ± .54
55.20 ± .73
26.20 ± .80
33.55 ± .39
RAPTOR
39.50 ± .81
55.35 ± .93
32.10 ± .74
37.90 ± .41
14.15 ± .48
23.85 ± .93
KGP
23.20 ± .66
33.10 ± .44
12.05 ± .89
13.80 ± .70
11.75 ± .82
18.90 ± .61
ToG
21.30 ± .72
27.45 ± .85
15.20 ± .91
17.75 ± .76
6.40 ± .69
9.25 ± .59
GraphRAG
30.65 ± .63
40.60 ± .78
8.25 ± .56
9.10 ± .51
1.10 ± .32
2.30 ± .38
Table 5 : QA performance on the three multi-hop benchmarks, measured by EM and F1 computed against the gold answers. The values represent the means and standard deviations computed over 5 samples, each of which contains 1,000 different questions. For each column, the optimal value is shown in bold and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
HippoRAG2
312,441 ± 1,842
123,187 ± 1,253
321,905 ± 1,976
RAPTOR
28,763 ± 412
16,724 ± 318
36,204 ± 487
KGP
1,941 ± 63
1,478 ± 51
4,213 ± 74
ToG
141,820 ± 934
69,340 ± 721
148,573 ± 1,102
GraphRAG
661,392 ± 2,341
75,614 ± 843
197,840 ± 1,587
LightRAG
728,471 ± 2,813
347,605 ± 1,934
598,317 ± 2,645
Table 6 : Indexing time (Idx) in seconds across the three benchmarks. The values represent the means and standard deviations computed over 5 independent runs of the indexing procedure on the full corpus of each benchmark. For each column, the optimal value is shown in bold, and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
Res ↓
Tok ↓
Res ↓
Tok ↓
Res ↓
Tok ↓
HippoRAG2
76.84 ± 2.83
591 ± 12
80.73 ± 3.91
624 ± 14
83.12 ± 2.97
671 ± 15
RAPTOR
15.23 ± .41
854 ± 18
14.89 ± .38
843 ± 17
15.91 ± .44
856 ± 19
KGP
31.40 ± 1.57
2,478 ± 31
26.55 ± 1.49
359 ± 9
38.04 ± 1.63
2,298 ± 28
ToG
56.81 ± 1.74
124 ± 6
40.67 ± 1.62
112 ± 5
49.93 ± 1.68
121 ± 6
GraphRAG
15.34 ± .39
3,812 ± 42
16.08 ± .43
871 ± 19
16.20 ± .45
1,398 ± 24
Table 7 : Efficiency comparison across the three benchmarks in terms of response time (Res), measured in seconds, and Tok. The values represent the means and standard deviations computed over 5 samples, each composed of 1,000 different questions. For each column, the optimal value is shown in bold and the suboptimal value is underlined.
HotpotQA
2Wiki
MuSiQue
βt
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
EM ↑
F1 ↑
R@2 ↑ R@5 ↑
1.00
59.22
72.25
72.10 88.70
50.90
56.96
62.05 76.75
25.70
35.92
41.30 55.52
K∣Rt−1∣
60.50
73.66
72.65 89.30
52.70
63.42
63.08 79.27
27.40
37.89
42.93 57.08
0.00
58.70
71.24
72.05 87.72
51.24
62.31
62.58 78.81
26.42
36.17
41.17 56.31
Table 8 : Values of EM, F1, R@2 and R@5 obtained by MatRAG on the three benchmarks for βt set to 1.00, K∣Rt−1∣ and 0.00. The values represent the means computed over 5 samples, each composed of 1,000 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
l
ml
DB ↓
Sl ↑
DB ↓
Sl ↑
DB ↓
Sl ↑
1
768
0.8667
0.4484
0.7313
0.5512
0.8362
0.4492
64
0.8630
0.4628
0.7198
0.5536
0.7557
0.5406
2
768
0.8342
0.4663
0.7503
0.5085
0.7573
0.4746
128
0.7616
0.4749
0.7389
0.5059
0.7661
0.4860
3
768
0.7428
0.5206
0.6608
0.5390
0.7330
0.4876
Table 9 : Values of the Davies-Bouldin index (DB) and the Silhouette score (Sl) for the level l and embedding dimension across benchmarks. The ml column indicates the embedding dimension used for clustering with full-dimensional and Matryoshka embeddings. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
Embedding
EM ↑
F1 ↑
EM ↑
F1 ↑
EM ↑
F1 ↑
Full
59.28
72.47
49.60
60.29
26.00
36.23
Matryoshka
60.50
73.66
52.70
63.42
27.40
37.89
Table 10 : Ablation study on QA performance, as measured by EM and F1, on the three benchmarks. The Embedding column indicates whether the system uses Matryoshka truncated (Matryoshka) or full-dimensional (Full) embeddings during retrieval. Values represent the results computed over 5 samples, composed of 1,000 different questions. For each column, the optimal value is shown in bold.
HotpotQA
2Wiki
MuSiQue
Embedding
R@2 ↑
R@5 ↑
Res ↓
R@2 ↑
R@5 ↑
Res ↓
R@2 ↑
R@5 ↑
Res ↓
Full
70.86
86.72
3.10
62.12
76.88
4.04
41.86
55.80
4.82
Matryoshka
72.65
89.30
1.85
63.08
79.27
2.49
42.93
57.08
2.99
Table 11 : Ablation study on retrieval and efficiency performance, as measured by R@2, R@5, and Res (computed as the sum of retrieval and answer generation time) on the three benchmarks. The Embedding column indicates whether the system uses Matryoshka truncated (Matryoshka) or full-dimensional (Full) embeddings during retrieval. Values represent the results computed over 5 samples composed of 1,000 different questions. For each column, the optimal value is shown in bold.