Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo---heading paths that are structurally valid but shuffled across documents. We score retrieval with a coverage-aware nDCG and test four pre-registered contrasts via document-clustered bootstrap with Holm correction, on two distant corpora: 200 Wikipedia Featured Articles (951 queries) and 1,585 QASPER papers (4,303 questions). Organization helps, and the cause is content, not tokens: structure-aligned chunks with real heading paths beat contextualized fixed windows (+0.022 / +0.012 cov-nDCG@10) and the placebo (+0.010 / +0.016). Naive two-stage hierarchical retrieval hurts (-0.033 / -0.015), traceable to first-stage section recall. Gold structure beats LLM-induced structure on Wikipedia but not on QASPER. Effects are small (dz 0.06-0.11) but Holm-significant and consistent across corpora.
Figures & tables
Figure 1: The four document-structure mechanisms under test. (a) Structure-aligned chunking cuts at section boundaries (A2) instead of ignoring them (A0). (b) Chunk contextualization prepends an LLM-generated blurb to a fixed-size window (A1). (c) Structure-metadata injection prepends the chunk’s real heading path (A3); the placebo (P) prepends an equally-shaped but shuffled, semantically null path. (d) Hierarchical retrieval routes queries to sections before ranking chunks (A4); flat retrieval (A3) searches the full chunk index directly. A4 underperforms because stage 1 frequently drops the gold section before stage 2 ever sees it (Section 6.4 ).
Contrast
Corpus
Δ
95% CI
p (Holm)
dz
p (MixedLM)
C1: A3 − A1
Wikipedia-FA
+0.0218
[+0.0052, +0.0380]
0.017
+0.09
0.120
QASPER
+0.0115
[+0.0055, +0.0175]
< 0.001
+0.06
0.008
C2: A3 − P
Wikipedia-FA
+0.0099
[+0.0024, +0.0173]
0.017
+0.08
0.487
QASPER
+0.0161
[+0.0114, +0.0207]
< 0.001
+0.11
< 0.001
C3: G3 − A3
Wikipedia-FA
+0.0129
[+0.0038, +0.0228]
0.010
+0.10
0.373
QASPER
+0.0011
[ − 0.0020, +0.0043]
0.485 (n.s.)
+0.01
0.801
Table 1: Pre-registered contrasts, cov-nDCG@10, strong embedder, no reranker. Bootstrap CIs and Holm-corrected p -values. MixedLM Seabold and Perktold (2010) ; Seabold et al. (2025) is the conservative secondary model (See sec. 5.2 ).
Corpus
A0
A1
A2
A3
P
A4
G2
G3
G4
Wikipedia-FA
0.362
0.391
0.403
0.413
0.403
0.381
0.417
0.426
0.400
QASPER
0.098
0.114
0.117
0.125
0.109
0.111
0.120
0.126
0.111
Table 2: Per-condition cov-nDCG@10 (strong embedder, no reranker).
Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
Retrieval Augmented Generation (RAG) is a key component for generating accurate and hallucination free answers using Large Language Models (LLMs). LLMs are improving at handling long context, but still suffer from "lost in the middle" problem. Thus, precise and accurate retrieval is important. Current retrievers chunk long context into length-based manageable chunks - in the process throwing away rich and informative semantic global structure in the corpus. We introduce a novel retrieval system STAIR that empowers an LLM to exploit global structure in a corpus such as a Table of Contents (ToC) to efficiently store and retrieve information from its model parameters. Our thorough and careful ablation studies with a finetuned Differentiable Search Index (DSI) system show that ToC helps build a low hallucination (less than 0.05%) generative Information Retrieval (IR) system and can generalize to examples where very few training samples are available. To further research in this novel direction of ToC based retrieval we release SearchTome - a diverse benchmark created from 18 books across 6 diverse domains to further research in this novel direction. STAIR achieves a high Recall@1 score of 82.6% on SearchTome as compared to DSI (76.9%), where the difference is found to be statistically significant. STAIR easily beats other strong baselines such as BM25 (59.5%), DPR (68.7%) and out-of-the-box Mistral (13.8%).
Retrieval-augmented generation (RAG) systems must balance retrieval granularity with contextual coherence, a challenge that existing methods address through LLM-guided chunking, single-level context expansion, or hierarchical summarization. These approaches variously depend on costly LLM calls during indexing or retrieval, limit context aggregation to a single granularity level, or introduce information loss through summarization. We present SproutRAG, an attention-guided hierarchical RAG framework that addresses this trade-off by organizing sentence-level chunks into progressively larger but semantically coherent units, using learned inter-sentence attention to construct a binary chunking tree. Unlike prior approaches that rely on external LLMs, fixed context expansion, or lossy summarization, SproutRAG learns which attention heads and layers best capture semantic document structure, enabling multi-granularity retrieval without additional LLM calls or compressed summaries. At retrieval time, SproutRAG uses hierarchical beam search to retrieve candidates at multiple granularities, capturing multi-sentence relevance beyond flat retrieval. The framework is trained end-to-end with a joint objective that improves both embeddings and tree structure. Experiments across four benchmarks spanning scientific, legal, and open-domain settings demonstrate that SproutRAG improves information efficiency (IE) by 6.1% on average over the strongest baseline. Code is available on https://github.com/AmirAbaskohi/SproutRAG.
Amirhossein Abaskohi, Issam H. Laradji, Peter West +1
University of British Columbia · ServiceNow Research