Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
Figures & tables
Figure 1: An illustrative example of a hierarchical tree structure associated with a document and its JSON string representation.
Figure 2: Overview of RIT-RAG. (a) Offline, we index the text of the document trees Ti , with each chunk c belonging to one node node(c) . (b) Each query q~t retrieves top- k′ chunks; their top n nodes and ancestors induce a sub-forest Fq~t . (c) In each round, the LLM calls GetToC , reasons over S(Fq~t) , reads selected nodes St , and either answers or continues.
Corpus
WixQA
QASPER
FinanceBench
Family
Method
Haiku
GPT-5-mini
Luna
Haiku
GPT-5-mini
Luna
Haiku
GPT-5-mini
Luna
Vanilla
RAG ( k=5 )
79.5
86.0
82.5
72.0
77.2
73.7
41.3
49.3
48.0
Graph
HyperGraphRAG
87.5
94.0
87.5
44.8
52.6
52.6
29.3
35.3
47.3
Cog-RAG
87.0
95.5
92.0
44.4
55.5
52.6
32.0
43.3
45.3
Agentic
Agentic Search
84.5
91.0
91.0
86.2
85.6
85.7
74.0
76.0
75.3
Interact-RAG
71.5
79.5
85.0
74.2
73.4
78.3
74.7
74.0
80.7
Table 1: Answer accuracy (%) on WixQA, QASPER, and FinanceBench, by generation model. Best per column in bold .
WixQA
QASPER
FinanceBench
Method
P
R
F1
P
R
F1
P
R
F1
Vanilla
RAG ( k=5 )
20.5
64.7
31.1
21.6
67.7
32.8
8.0
34.7
13.0
Agentic
AS
15.1
70.8
24.9
17.3
85.5
28.8
11.5
71.0
19.8
I-RAG
9.1
60.9
15.8
13.8
63.9
22.7
8.8
66.6
15.5
Table 4: Precision (P), recall (R), and F1 (%) of the passages that each method reads, measured against the gold evidence, with Haiku as the LLM. AS: Agentic Search; I-RAG: Interact-RAG. Best P, R, and F1 per dataset in bold .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Corpus
Node =
Docs
Chunks
Tokens
Q
WixQA
article section
6,221
32,032
3.09 M
200
FinanceBench
PDF page span
364
129,973
48.5 M
150
EntDocs
doc. webpage
2,839,892
8,993,388
3.81 B
317
QASPER
paper section
1,585
31,014
10.6 M
1372
Appendix
Table 5: Corpus and evaluation-set statistics. Docs is the number of source pdfs or webpages indexed; Chunks is the number of retrieval chunks after node-aligned chunking; Q is the number of questions evaluated.
Corpus
Total
Max
Node size (tok)
nodes
depth
mean
med.
p95
WixQA
31,733
5
84
60
219
FinanceBench
79,658
8
602
274
2,368
EntDocs
2,934,928
17
1,020
182
1,960
QASPER
24,350
4
362
262
977
Appendix
Table 6: Hierarchical document structure per corpus. Depth is from the document root; node size is a node’s own-text token count.
Table 7: Models and decoding settings used in the experiments. Role identifies how each model is used in the pipeline. Temperature and Max tokens report the generation settings. Input/Output price gives the effective API cost in USD per one million tokens.
Figure 3: Answer accuracy vs. mean API cost (US cents/question) with Haiku. Each dataset has its own panel/axes. Graph-based methods, PageIndex, and RIT-RAG use the Section 5.1 preprocessing. RIT-RAG is the star; error bars show Wilson 95% CIs, and dashed lines indicate the Pareto frontier.
Figure 4: Answer accuracy versus the mean LLM latency per question (seconds), with Haiku as the LLM. Each dataset has its own panel and axes. The graph-based methods, PageIndex, and RIT-RAG use the preprocessing of Section 5.1 , viz. , Haiku on WixQA, and Qwen3-32B on QASPER and FinanceBench. RIT-RAG is the star, the error bars are Wilson 95% confidence intervals, and the dashed line is the Pareto frontier.
WixQA ( n=200 )
QASPER ( n=1372 )
FinanceBench ( n=150 )
Method
Acc.
¢/q
s/q
Acc.
¢/q
s/q
Acc.
¢/q
s/q
RAG ( k=5 )
79.5
0.2
3
72.0
0.3
9
41.3
0.3
4
HyperGraphRAG
87.5
2.6
14
44.8
2.9
39
29.3
3.0
24
Cog-RAG
87.0
6.2
29
44.4
5.9
29
32.0
6.8
54
Agentic Search
84.5
1.0
25
86.2
4.1
55
74.0
4.8
23
Interact-RAG
71.5
2.8
57
74.2
6.2
79
74.7
6.5
42
Appendix
Table 8: Per-query online trade-offs (Haiku generation). Acc.: binary-judge accuracy (%); ¢/q: mean API cost (US cents); s/q: mean LLM latency (s).
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.427
0.740
0.821
0.779
0.658
0.745
0.948
0.491
GPT-5-mini
0.411
0.792
0.883
0.796
0.739
0.851
0.929
0.652
Luna
0.256
0.778
0.850
0.831
0.632
0.742
0.907
0.528
HyperGraphRAG
Haiku
0.632
0.858
0.926
0.860
0.823
0.889
0.945
0.781
GPT-5-mini
0.655
0.969
0.989
0.975
0.959
0.978
0.997
0.932
Luna
0.514
0.919
0.974
0.935
0.882
0.922
0.976
0.792
Appendix
Table 10: WixQA answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.370
0.791
0.933
0.771
0.881
0.835
0.988
0.433
GPT-5-mini
0.278
0.705
0.910
0.733
0.843
0.837
0.985
0.455
Luna
0.255
0.667
0.918
0.677
0.803
0.760
0.993
0.407
HyperGraphRAG
Haiku
0.386
0.515
0.807
0.546
0.682
0.687
0.959
0.599
GPT-5-mini
0.296
0.574
0.828
0.588
0.645
0.648
0.737
0.407
Luna
0.338
0.693
0.935
0.683
0.697
0.737
0.778
0.401
Appendix
Table 11: FinanceBench answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.392
0.601
0.698
0.626
0.607
0.557
0.928
0.379
GPT-5-mini
0.288
0.709
0.801
0.721
0.655
0.697
0.880
0.457
Luna
0.218
0.655
0.734
0.664
0.564
0.608
0.827
0.385
Agentic Search
Haiku
0.559
0.827
0.868
0.832
0.777
0.839
0.918
0.470
GPT-5-mini
0.369
0.834
0.883
0.851
0.759
0.836
0.929
0.502
Luna
0.374
0.852
0.891
0.863
0.790
0.845
0.924
0.548
Appendix
Table 12: EntQABench answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Method
Gen
Recall
Corr
Rel
Fact
Compr
Know
Coher
Div
RAG ( k=5 )
Haiku
0.680
0.836
0.925
0.823
0.869
0.822
0.995
0.537
GPT-5-mini
0.611
0.833
0.913
0.828
0.835
0.846
0.987
0.574
Luna
0.569
0.789
0.885
0.792
0.779
0.795
0.978
0.516
HyperGraphRAG
Haiku
0.612
0.651
0.765
0.653
0.650
0.694
0.944
0.646
GPT-5-mini
0.530
0.729
0.841
0.734
0.693
0.760
0.939
0.603
Luna
0.542
0.752
0.838
0.756
0.707
0.776
0.934
0.605
Appendix
Table 13: QASPER answer quality by method and generation model. G-E dimensions (each normalized 0 – 1 ): Corr(ectness), Rel(evance), Fact(uality), Compr(ehensiveness), Know(ledgeability), Coher(ence, logical), Div(ersity). Recall is lexical token recall. Best per column in bold .
Dataset
WixQA
QASPER
FinanceBench
Family
Method
P
R
F1
P
R
F1
P
R
F1
Vanilla
RAG ( k=5 )
20.5
64.7
31.1
21.6
67.7
32.8
8.0
34.7
13.0
Agentic
Agentic Search
17.1
76.9
28.0
19.0
85.0
31.1
12.9
68.6
21.7
Interact-RAG
7.2
75.4
13.1
9.6
61.1
16.6
6.9
63.2
12.4
PageIndex
28.7
72.0
41.0
50.4
49.7
50.0
34.5
58.0
43.3
Ours
RIT-RAG
31.8
80.0
45.5
32.3
86.1
47.0
29.6
78.1
42.9
Appendix
Table 14: Precision (P), recall (R), and F1 (%) of the passages that each method reads, measured against the gold evidence, with GPT-5-mini as the LLM. Best P, R, and F1 per dataset in bold .
Dataset
WixQA
QASPER
FinanceBench
Family
Method
P
R
F1
P
R
F1
P
R
F1
Vanilla
RAG ( k=5 )
20.5
64.7
31.1
21.6
67.7
32.8
8.0
34.7
13.0
Agentic
Agentic Search
16.6
75.4
27.2
19.8
85.6
32.2
13.1
72.8
22.2
Interact-RAG
12.6
74.8
21.6
13.4
68.1
22.4
10.3
67.3
17.9
PageIndex
28.7
72.0
41.0
49.3
43.9
46.4
31.6
50.4
38.8
Ours
RIT-RAG
32.4
80.1
46.1
34.4
86.4
49.2
30.1
77.1
43.3
Appendix
Table 15: Precision (P), recall (R), and F1 (%) of the passages that each method reads, measured against the gold evidence, with Luna as the LLM. Best P, R, and F1 per dataset in bold .
Figure 5: Which questions each method answers correctly (Haiku generation), as an outline Venn. Each colored ring is one method’s set of correct questions and region labels are question counts; “none correct” (all three wrong) is noted per panel.
Figure 6: Multi-document question. PageIndex commits to a single filing before searching and can never see the FY2022 revenue. RIT-RAG retrieves a corpus-level table of contents, locates the income statement in both 10-Ks, and answers in two tool calls. Traces abridged ( […] ); quoted text is verbatim model output.
Figure 7: Structure fragmentation. Agentic RAG retrieves the right balance sheet, but its fixed-size chunk ends mid-table, before the total-current-liabilities line. After ten searches it sums the visible rows plus a figure from a note in a different filing, and answers 2.37. RIT-RAG navigates to the Consolidated Balance Sheets node and reads the full statement, so the total is available and the ratio is correct. Traces abridged ( […] ).
Figure 8: Accuracy of RIT-RAG against the number of nodes n that the GetToC tool keeps, on held-out questions of WixQA (simulated split) and QASPER (validation split), with Haiku for preprocessing and generation.
Figure 9: Failure modes of graph-based retrieval. (a) Entity nodes such as Word2Vec are shared across the whole corpus, so a hyperedge from an unrelated paper is retrieved as evidence and the model reports FastText. (b) Table cells are extracted as free-floating numeric entities without their company or fiscal period. The retrieved subgraph holds several contradictory “FY2022 free cash flow” values, and the answer is built on a conversion ratio that does not belong to Adobe. In both cases the gold evidence is absent from ∼ 80–90k characters of retrieved knowledge.
Mean calls per query
Benchmark
n
Total
ToC
Cont.
Nodes
N/call
WixQA
200
2.99
1.33
1.66
5.5
3.3
QASPER
1372
2.75
1.43
1.32
3.3
2.5
FinanceBench
150
2.70
1.46
1.24
2.7
2.2
EntQABench
317
2.99
1.47
1.51
3.8
2.5
Appendix
Table 16: Per-query tool-call statistics for RIT-RAG (Haiku generation). ToC / Cont. : mean calls to get_toc_structure / get_content ; Total : their sum; Nodes : mean content nodes per query; N/call : nodes per get_content call.
Retrieval-augmented generation (RAG) systems must balance retrieval granularity with contextual coherence, a challenge that existing methods address through LLM-guided chunking, single-level context expansion, or hierarchical summarization. These approaches variously depend on costly LLM calls during indexing or retrieval, limit context aggregation to a single granularity level, or introduce information loss through summarization. We present SproutRAG, an attention-guided hierarchical RAG framework that addresses this trade-off by organizing sentence-level chunks into progressively larger but semantically coherent units, using learned inter-sentence attention to construct a binary chunking tree. Unlike prior approaches that rely on external LLMs, fixed context expansion, or lossy summarization, SproutRAG learns which attention heads and layers best capture semantic document structure, enabling multi-granularity retrieval without additional LLM calls or compressed summaries. At retrieval time, SproutRAG uses hierarchical beam search to retrieve candidates at multiple granularities, capturing multi-sentence relevance beyond flat retrieval. The framework is trained end-to-end with a joint objective that improves both embeddings and tree structure. Experiments across four benchmarks spanning scientific, legal, and open-domain settings demonstrate that SproutRAG improves information efficiency (IE) by 6.1% on average over the strongest baseline. Code is available on https://github.com/AmirAbaskohi/SproutRAG.
Amirhossein Abaskohi, Issam H. Laradji, Peter West +1
University of British Columbia · ServiceNow Research
Retrieval-augmented generation (RAG) enhances large language models with external knowledge, and tree-based RAG organizes documents into hierarchical indexes to support queries at multiple granularities. However, existing Tree-RAG methods designed for single-document retrieval face critical challenges in scaling to cross-document multi-hop questions: (1) poor distribution adaptability, where k-means clustering introduces noise due to rigid distribution assumptions; (2) structural isolation, as tree indexes lack explicit cross-document connections; and (3) coarse abstraction, which obscures fine-grained details. To address these limitations, we propose Ψ-RAG, a tree-RAG framework with two key components. First, a hierarchical abstract tree index built through an iterative "merging and collapse" process that adapts to data distributions without a priori assumption. Second, a multi-granular retrieval agent that intelligently interacts with the knowledge base with reorganized queries and an agent-powered hybrid retriever. Ψ-RAG supports diverse tasks from token-level question answering to document-level summarization. On cross-document multi-hop QA benchmarks, it outperforms RAPTOR by 25.9% and HippoRAG 2 by 7.4% in average F1 score. Code is available at https://github.com/Newiz430/Psi-RAG.
Ziwen Zhao, Menglin Yang
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China.
Answering complex questions over large document collections requires assembling complementary evidence across sections and documents. GraphRAG offers structured retrieval but typically uses fixed traversal, while agentic RAG operates over weakly structured interfaces. Our key insight is that agents should navigate document structure within and across documents rather than repeatedly search from scratch. We introduce DocNavRAG, which organizes document hierarchies and cross-region relations into a navigable graph, exposes graph operations for locating, navigating, expanding, and fetching, and maintains an evolving evidence state to guide retrieval until sufficient evidence is collected. Across four long- and multi-document QA benchmarks, DocNavRAG improves answer quality and context sufficiency over the strongest baseline by 7.8% and 17.7% on average.
Dongyang Xie, Yao Tian, Hao Zhang +5
School of Computer Science, Wuhan University · The Hong Kong University of Science and Technology · The Chinese University of Hong Kong +1