Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo---heading paths that are structurally valid but shuffled across documents. We score retrieval with a coverage-aware nDCG and test four pre-registered contrasts via document-clustered bootstrap with Holm correction, on two distant corpora: 200 Wikipedia Featured Articles (951 queries) and 1,585 QASPER papers (4,303 questions). Organization helps, and the cause is content, not tokens: structure-aligned chunks with real heading paths beat contextualized fixed windows (+0.022 / +0.012 cov-nDCG@10) and the placebo (+0.010 / +0.016). Naive two-stage hierarchical retrieval hurts (-0.033 / -0.015), traceable to first-stage section recall. Gold structure beats LLM-induced structure on Wikipedia but not on QASPER. Effects are small (dz 0.06-0.11) but Holm-significant and consistent across corpora.
Figures & tables
Figure 1: The four document-structure mechanisms under test. (a) Structure-aligned chunking cuts at section boundaries (A2) instead of ignoring them (A0). (b) Chunk contextualization prepends an LLM-generated blurb to a fixed-size window (A1). (c) Structure-metadata injection prepends the chunk’s real heading path (A3); the placebo (P) prepends an equally-shaped but shuffled, semantically null path. (d) Hierarchical retrieval routes queries to sections before ranking chunks (A4); flat retrieval (A3) searches the full chunk index directly. A4 underperforms because stage 1 frequently drops the gold section before stage 2 ever sees it (Section 6.4 ).
Contrast
Corpus
Δ
95% CI
p (Holm)
dz
p (MixedLM)
C1: A3 − A1
Wikipedia-FA
+0.0218
[+0.0052, +0.0380]
0.017
+0.09
0.120
QASPER
+0.0115
[+0.0055, +0.0175]
< 0.001
+0.06
0.008
C2: A3 − P
Wikipedia-FA
+0.0099
[+0.0024, +0.0173]
0.017
+0.08
0.487
QASPER
+0.0161
[+0.0114, +0.0207]
< 0.001
+0.11
< 0.001
C3: G3 − A3
Wikipedia-FA
+0.0129
[+0.0038, +0.0228]
0.010
+0.10
0.373
QASPER
+0.0011
[ − 0.0020, +0.0043]
0.485 (n.s.)
+0.01
0.801
Table 1: Pre-registered contrasts, cov-nDCG@10, strong embedder, no reranker. Bootstrap CIs and Holm-corrected p -values. MixedLM Seabold and Perktold (2010) ; Seabold et al. (2025) is the conservative secondary model (See sec. 5.2 ).
Corpus
A0
A1
A2
A3
P
A4
G2
G3
G4
Wikipedia-FA
0.362
0.391
0.403
0.413
0.403
0.381
0.417
0.426
0.400
QASPER
0.098
0.114
0.117
0.125
0.109
0.111
0.120
0.126
0.111
Table 2: Per-condition cov-nDCG@10 (strong embedder, no reranker).