Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
Figures & tables
Figure 1: An illustrative example of a hierarchical tree structure associated with a document and its JSON string representation.
Figure 2: Overview of RIT-RAG. (a) Offline, we index the text of the document trees Ti , with each chunk c belonging to one node node(c) . (b) Each query q~t retrieves top- k′ chunks; their top n nodes and ancestors induce a sub-forest Fq~t . (c) In each round, the LLM calls GetToC , reasons over S(Fq~t) , reads selected nodes St , and either answers or continues.
Corpus
WixQA
QASPER
FinanceBench
Family
Method
Haiku
GPT-5-mini
Luna
Haiku
GPT-5-mini
Luna
Haiku
GPT-5-mini
Luna
Vanilla
RAG ( k=5 )
79.5
86.0
82.5
72.0
77.2
73.7
41.3
49.3
48.0
Graph
HyperGraphRAG
87.5
94.0
87.5
44.8
52.6
52.6
29.3
35.3
47.3
Cog-RAG
87.0
95.5
92.0
44.4
55.5
52.6
32.0
43.3
45.3
Agentic
Agentic Search
84.5
91.0
91.0
86.2
85.6
85.7
74.0
76.0
75.3
Interact-RAG
71.5
79.5
85.0
74.2
73.4
78.3
74.7
74.0
80.7
Table 1: Answer accuracy (%) on WixQA, QASPER, and FinanceBench, by generation model. Best per column in bold .
WixQA
QASPER
FinanceBench
Method
P
R
F1
P
R
F1
P
R
F1
Vanilla
RAG ( k=5 )
20.5
64.7
31.1
21.6
67.7
32.8
8.0
34.7
13.0
Agentic
AS
15.1
70.8
24.9
17.3
85.5
28.8
11.5
71.0
19.8
I-RAG
9.1
60.9
15.8
13.8
63.9
22.7
8.8
66.6
15.5
Table 4: Precision (P), recall (R), and F1 (%) of the passages that each method reads, measured against the gold evidence, with Haiku as the LLM. AS: Agentic Search; I-RAG: Interact-RAG. Best P, R, and F1 per dataset in bold .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Corpus
Node =
Docs
Chunks
Tokens
Q
WixQA
article section
6,221
32,032
3.09 M
200
FinanceBench
PDF page span
364
129,973
48.5 M
150
EntDocs
doc. webpage
2,839,892
8,993,388
3.81 B
317
QASPER
paper section
1,585
31,014
10.6 M
1372
Appendix
Table 5: Corpus and evaluation-set statistics. Docs is the number of source pdfs or webpages indexed; Chunks is the number of retrieval chunks after node-aligned chunking; Q is the number of questions evaluated.
Corpus
Total
Max
Node size (tok)
nodes
depth
mean
med.
p95
WixQA
31,733
5
84
60
219
FinanceBench
79,658
8
602
274
2,368
EntDocs
2,934,928
17
1,020
182
1,960
QASPER
24,350
4
362
262
977
Appendix
Table 6: Hierarchical document structure per corpus. Depth is from the document root; node size is a node’s own-text token count.
Table 7: Models and decoding settings used in the experiments. Role identifies how each model is used in the pipeline. Temperature and Max tokens report the generation settings. Input/Output price gives the effective API cost in USD per one million tokens.
Figure 3: Answer accuracy vs. mean API cost (US cents/question) with Haiku. Each dataset has its own panel/axes. Graph-based methods, PageIndex, and RIT-RAG use the Section 5.1 preprocessing. RIT-RAG is the star; error bars show Wilson 95% CIs, and dashed lines indicate the Pareto frontier.
Figure 4: Answer accuracy versus the mean LLM latency per question (seconds), with Haiku as the LLM. Each dataset has its own panel and axes. The graph-based methods, PageIndex, and RIT-RAG use the preprocessing of Section 5.1 , viz. , Haiku on WixQA, and Qwen3-32B on QASPER and FinanceBench. RIT-RAG is the star, the error bars are Wilson 95% confidence intervals, and the dashed line is the Pareto frontier.
WixQA ( n=200 )
QASPER ( n=1372 )
FinanceBench ( n=150 )
Method
Acc.
¢/q
s/q
Acc.
¢/q
s/q
Acc.
¢/q
s/q
RAG ( k=5 )
79.5
0.2
3
72.0
0.3
9
41.3
0.3
4
HyperGraphRAG
87.5
2.6
14
44.8
2.9
39
29.3
3.0
24
Cog-RAG
87.0
6.2
29
44.4
5.9
29
32.0
6.8
54
Agentic Search
84.5
1.0
25
86.2
4.1
55
74.0
4.8
23
Interact-RAG
71.5
2.8
57
74.2
6.2
79
74.7
6.5
42
Appendix
Table 8: Per-query online trade-offs (Haiku generation). Acc.: binary-judge accuracy (%); ¢/q: mean API cost (US cents); s/q: mean LLM latency (s).
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.427
0.740
0.821
0.779
0.658
0.745
0.948
0.491
GPT-5-mini
0.411
0.792
0.883
0.796
0.739
0.851
0.929
0.652
Luna
0.256
0.778
0.850
0.831
0.632
0.742
0.907
0.528
HyperGraphRAG
Haiku
0.632
0.858
0.926
0.860
0.823
0.889
0.945
0.781
GPT-5-mini
0.655
0.969
0.989
0.975
0.959
0.978
0.997
0.932
Luna
0.514
0.919
0.974
0.935
0.882
0.922
0.976
0.792
Appendix
Table 10: WixQA answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.370
0.791
0.933
0.771
0.881
0.835
0.988
0.433
GPT-5-mini
0.278
0.705
0.910
0.733
0.843
0.837
0.985
0.455
Luna
0.255
0.667
0.918
0.677
0.803
0.760
0.993
0.407
HyperGraphRAG
Haiku
0.386
0.515
0.807
0.546
0.682
0.687
0.959
0.599
GPT-5-mini
0.296
0.574
0.828
0.588
0.645
0.648
0.737
0.407
Luna
0.338
0.693
0.935
0.683
0.697
0.737
0.778
0.401
Appendix
Table 11: FinanceBench answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.392
0.601
0.698
0.626
0.607
0.557
0.928
0.379
GPT-5-mini
0.288
0.709
0.801
0.721
0.655
0.697
0.880
0.457
Luna
0.218
0.655
0.734
0.664
0.564
0.608
0.827
0.385
Agentic Search
Haiku
0.559
0.827
0.868
0.832
0.777
0.839
0.918
0.470
GPT-5-mini
0.369
0.834
0.883
0.851
0.759
0.836
0.929
0.502
Luna
0.374
0.852
0.891
0.863
0.790
0.845
0.924
0.548
Appendix
Table 12: EntQABench answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.680
0.836
0.925
0.823
0.869
0.822
0.995
0.537
GPT-5-mini
0.611
0.833
0.913
0.828
0.835
0.846
0.987
0.574
Luna
0.569
0.789
0.885
0.792
0.779
0.795
0.978
0.516
HyperGraphRAG
Haiku
0.612
0.651
0.765
0.653
0.650
0.694
0.944
0.646
GPT-5-mini
0.530
0.729
0.841
0.734
0.693
0.760
0.939
0.603
Luna
0.542
0.752
0.838
0.756
0.707
0.776
0.934
0.605
Appendix
Table 13: QASPER answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Dataset
WixQA
QASPER
FinanceBench
Family
Method
P
R
F1
P
R
F1
P
R
F1
Vanilla
RAG ( k=5 )
20.5
64.7
31.1
21.6
67.7
32.8
8.0
34.7
13.0
Agentic
Agentic Search
17.1
76.9
28.0
19.0
85.0
31.1
12.9
68.6
21.7
Interact-RAG
7.2
75.4
13.1
9.6
61.1
16.6
6.9
63.2
12.4
PageIndex
28.7
72.0
41.0
50.4
49.7
50.0
34.5
58.0
43.3
Ours
RIT-RAG
31.8
80.0
45.5
32.3
86.1
47.0
29.6
78.1
42.9
Appendix
Table 14: Precision (P), recall (R), and F1 (%) of the passages that each method reads, measured against the gold evidence, with GPT-5-mini as the LLM. Best P, R, and F1 per dataset in bold .
Dataset
WixQA
QASPER
FinanceBench
Family
Method
P
R
F1
P
R
F1
P
R
F1
Vanilla
RAG ( k=5 )
20.5
64.7
31.1
21.6
67.7
32.8
8.0
34.7
13.0
Agentic
Agentic Search
16.6
75.4
27.2
19.8
85.6
32.2
13.1
72.8
22.2
Interact-RAG
12.6
74.8
21.6
13.4
68.1
22.4
10.3
67.3
17.9
PageIndex
28.7
72.0
41.0
49.3
43.9
46.4
31.6
50.4
38.8
Ours
RIT-RAG
32.4
80.1
46.1
34.4
86.4
49.2
30.1
77.1
43.3
Appendix
Table 15: Precision (P), recall (R), and F1 (%) of the passages that each method reads, measured against the gold evidence, with Luna as the LLM. Best P, R, and F1 per dataset in bold .
Figure 5: Which questions each method answers correctly (Haiku generation), as an outline Venn. Each colored ring is one method’s set of correct questions and region labels are question counts; “none correct” (all three wrong) is noted per panel.
Figure 6: Multi-document question. PageIndex commits to a single filing before searching and can never see the FY2022 revenue. RIT-RAG retrieves a corpus-level table of contents, locates the income statement in both 10-Ks, and answers in two tool calls. Traces abridged ( […] ); quoted text is verbatim model output.
Figure 7: Structure fragmentation. Agentic RAG retrieves the right balance sheet, but its fixed-size chunk ends mid-table, before the total-current-liabilities line. After ten searches it sums the visible rows plus a figure from a note in a different filing, and answers 2.37. RIT-RAG navigates to the Consolidated Balance Sheets node and reads the full statement, so the total is available and the ratio is correct. Traces abridged ( […] ).
Figure 8: Accuracy of RIT-RAG against the number of nodes n that the GetToC tool keeps, on held-out questions of WixQA (simulated split) and QASPER (validation split), with Haiku for preprocessing and generation.
Figure 9: Failure modes of graph-based retrieval. (a) Entity nodes such as Word2Vec are shared across the whole corpus, so a hyperedge from an unrelated paper is retrieved as evidence and the model reports FastText. (b) Table cells are extracted as free-floating numeric entities without their company or fiscal period. The retrieved subgraph holds several contradictory “FY2022 free cash flow” values, and the answer is built on a conversion ratio that does not belong to Adobe. In both cases the gold evidence is absent from ∼ 80–90k characters of retrieved knowledge.
Mean calls per query
Benchmark
n
Total
ToC
Cont.
Nodes
N/call
WixQA
200
2.99
1.33
1.66
5.5
3.3
QASPER
1372
2.75
1.43
1.32
3.3
2.5
FinanceBench
150
2.70
1.46
1.24
2.7
2.2
EntQABench
317
2.99
1.47
1.51
3.8
2.5
Appendix
Table 16: Per-query tool-call statistics for RIT-RAG (Haiku generation). ToC / Cont. : mean calls to get_toc_structure / get_content ; Total : their sum; Nodes : mean content nodes per query; N/call : nodes per get_content call.