Large Language Models (LLMs) enable us to better understand text documents, including PDFs and Word documents. However, LLMs, as well as more modern LLM agents, i.e., those with tool-calling abilities, typically treat such documents as plain text, ignoring the fact that they are often organized hierarchically into sections and subsections. Extracting this structure, while difficult, can improve efficiency and effectiveness for agents (and humans)---since only sections relevant to a given task need to be processed. Unfortunately, prior work on structure extraction provides no formal guarantees on how well the inferred structure matches the true one. Instead, we target a robust and compact variant that is feasible to infer and useful in practice. Robustness ensures that the text under each subsection header is a superset of the text under the same header in the true structure. Compactness seeks to minimize this superset, reducing agentic cost (or human cognitive load). We propose SHED, a two-stage workflow for inferring a robust and compact structure. The first stage is pluggable with an infinite family of approaches, each guaranteeing robustness for a specific document class. We theoretically characterize the document space using these classes and their hierarchical relationships. Empirically, SHED improves F-1 scores (measuring the robustness--compactness trade-off) by 13%--68% over non-LLM baselines and 9%--15% over expensive LLM-based approaches. Finally, we show how SHED-inferred structures are valuable for agentic document analysis: agents using SHED outperform baselines, achieving 3%--23% higher accuracy while being up to 10x cheaper.
Figures & tables
Figure 1 . Hierarchical structures of a financial document Walmart Inc. form 10-K. The node vi in an SHT corresponds to the header phrase pi .
Figure 2 . Our contribution: Hierarchy of robustness classes.
Figure 3 . The overall workflow of SHED for SHT Inference.
Figure 4 . Intuition of SD Inference.
Figure 5 . An illustration of SHT inference.
Algorithm 1 Infer-SD1(V,C) // Local-First
Algorithm 2 Infer-SD2(V,C) // Global-First
Algorithm 3 Assemble-SHT ( V,C,sd )
Figure 6 . An example in which local-first generates a robust SHT while global-first does not (Figures 6(a) , 6(b) , 6(c) ), and vice versa (Figures 6(d) , 6(e) , 6(f) ). Colors indicate the cluster membership of each node in 6(g) .
Statistics
Civic
Contracts
Finance
Papers
Avg. doc size
7156
7863
178,905
17,352
#Docs (#Qs)
107
248
100
500
Table 1. Statistics of our evaluation datasets. Avg. doc size: average word count per document in each dataset. #Docs (#Qs): number of documents (questions) in each dataset.
Approach
Recall
Precision
F-1
Total Cost (USD)
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Tot.
GROBID
0.14
0.21
0.61
0.85
0.45
0.73
0.61
0.58
0.85
0.69
0.24
0.30
0.58
0.84
0.49
0
0
0
0
0
LLM-text
0.91
0.46
0.56
0.86
0.70
0.84
0.80
0.52
0.88
0.76
0.87
0.56
0.48
0.87
0.70
5.7
10.7
104.0
36.3
156.6
LLM-vision
0.91
0.37
0.72
0.85
0.71
0.82
0.73
0.48
0.85
0.72
0.85
0.47
0.55
0.85
0.68
12.0
19.9
88.6
43.4
164.0
SHED
0.93
0.85
0.95
0.92
0.91
0.89
0.93
0.60
0.90
0.83
0.91
0.88
0.72
0.91
0.86
0
0
0
0
0
Table 2 . Recall, precision, and F-1 scores for identified headers relative to the true headers, averaged per dataset and then across datasets (stratified average Avg. ), as well as the total cost (USD) of SHT inference per dataset ( Tot. : summed across datasets). SHED , Deep, and Wide use the same set of headers extracted with Huridocs; we report SHED only as the other two are identical on header inference and cost. Green : highest recall, precision, F-1 scores, and lowest cost per dataset
Approach
Recall (Robustness)
Precision (Compactness)
F-1 (Trade-off)
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Top-down: M↓(T)=avgv∈T′f(v) , where f(v) calculates recall, precision, and F-1 score of ts(v) relative to ts′(v) .
Deep
0.94
0.99
1.00
0.97
0.98
0.21
0.15
0.09
0.20
0.16
0.29
0.20
0.13
0.27
0.22
Wide
0.70
0.81
0.81
0.78
0.78
0.93
0.90
0.97
0.97
0.94
0.72
0.74
0.81
0.79
0.77
GROBID
0.48
0.84
0.74
0.76
0.71
0.25
0.45
0.75
0.83
0.57
0.26
0.47
0.66
0.76
0.54
LLM-text
0.96
0.93
0.76
0.96
0.90
0.98
0.67
0.61
0.96
0.81
0.96
0.67
0.57
0.95
0.79
Table 3 . M↓ and M↑ results, averaged per dataset and evaluated with f∈{REC,PREC,F-1} for recall, precision, and F-1 score, representing robustness, compactness, and their trade-off, respectively. Avg. : stratified averages across datasets. Green : highest per dataset per metric; costs for each scheme can be found in Table 2 .
Strategy
Civic
Contracts
Finance
Papers
Avg. (Tot.)
Vanilla-in-context
0.59
(3.1)
0.71
(6.7)
0.64
(112.5)
0.77
(32.6)
0.67
(154.9)
Vanilla-grep agent
0.57
(3.4)
0.45
( 4.9 )
0.63
( 2.3 )
0.53
(11.9)
0.55
(22.5)
SHT-aug-context
0.78
(4.1)
0.70
(9.0)
0.71
(118.2)
0.77
(34.1)
0.74
(165.3)
SHT-based agent
0.72
( 2.1 )
0.75
(5.1)
0.71
(3.9)
0.76
( 6.3 )
0.74
( 17.4 )
Table 4 . Average accuracy and total cost (USD, inside parenthesis) of four strategies by dataset. The last column reports stratified average accuracy ( Avg. ) and total cost ( Tot. ) across four datasets. Green : highest accuracy (lowest cost) per dataset.
Agent
Civic
Contracts
Finance
Papers
Avg. Acc. (Tot. Cost)
Vanilla-embed
0.38
(4.5)
0.39
(8.5)
0.64
(2.6)
0.64
(9.1)
0.51
(24.7)
Vanilla-grep
0.57
(3.4)
0.45
( 4.9 )
0.63
( 2.3 )
0.53
(11.9)
0.55
(22.5)
SHT-grep
0.62
(5.5)
0.70
(12.5)
0.65
(5.6)
0.61
(11.3)
0.65
(34.8)
SHT-embed
0.75
(4.9)
0.74
(15.2)
0.64
(5.8)
0.69
(9.2)
0.71
(35.0)
SHT-based
0.72
( 2.1 )
0.75
(5.1)
0.71
(3.9)
0.76
( 6.3 )
0.74
( 17.4 )
Table 5 . Average accuracy and total cost in parenthesis of all agents: baselines (vanilla-grep from Tab. 4 , and the 3 other approaches from Sec. 7.3 ) and SHT-based (Tab. 4 ). Green : highest acc. (lowest cost).
Figure 7 . Total cost (USD) vs. average accuracy of SHT-based agents using different SHTs across datasets.
Civic
Contracts
Finance
Papers
Avg. (Tot.)
Deep
0.59
(12.2)
0.63
(27.1)
0.20
(95.5)
0.75
(39.7)
0.54
(174.5)
Wide
0.57
(2.0)
0.61
( 4.9 )
0.67
( 3.0 )
0.76
( 5.3 )
0.65
( 15.1 )
GROBID
0.31
( 1.3 )
0.70
(5.8)
0.63
(5.0)
0.74
(5.3)
0.60
(17.4)
LLM-text
0.68
(7.9)
0.60
(15.9)
0.59
(108.4)
0.74
(40.4)
0.65
(172.6)
LLM-vision
0.75
(14.0)
0.50
(25.0)
0.69
(93.3)
0.74
(47.9)
0.67
(180.3)
SHED
0.67
(2.2)
0.74
(6.4)
0.75
(5.4)
0.75
(6.2)
0.73
(20.1)
Table 6 . Average accuracy and total cost (USD, inside parenthesis) of SHT-based agents using different SHTs by dataset. The last column reports stratified average accuracy ( Avg. ) and total cost ( Tot. ) across four datasets. Green : highest accuracy (lowest cost) per dataset.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Algorithm 4 Infer-SDK(V,C)
Figure 18
Datasets
Src. #Docs
Src. #Qs
#Const. Docs
Src. Query Example
Modified Query Example
Civic ( Lin et al., 2025b ; Center, 2024 )
19
–
4 ( k=3 )
Return a list of project names for all projects whose status matches the status of project ‘Westward Beach Road Shoulder Repairs (CalOES Project)’
According to the report for the meeting on January 26, 2022: Return a list of project names for all projects whose status matches the status of project ‘Westward Beach Road Shoulder Repairs (CalOES Project)’
Contracts ( Koreeda and Manning, 2021b )
73
1241
5 ( k=4 )
Determine the relationship between contract ‘064-19 Non Disclosure Agreement 2019’ and a hypothesis H (one of ‘Entailment’, ‘Contradiction’, or ‘NotMentioned’)
Return a list of contract names whose relationship to H is the same as that of the contract ‘064-19 Non Disclosure Agreement 2019’.
Finance ( Islam et al., 2023 )
84
150
2 ( k=1 )
What is the FY2018 capital expenditure amount (in USD millions) for 3M?
Same as the source query.
Papers ( Dasigi et al., 2021 )
416
1451
3 ( k=2 )
How big is the ANTISCAM dataset?
According to the paper ‘End-to-End Trainable Non-Collaborative Dialog System’: How big is the ANTISCAM dataset?
Appendix
Table 7 . Datasets are collected from existing sources and cited accordingly. Src. #Docs : number of documents in the source. Src. #Qs : number of questions in the source. For Civic , metadata are provided, allowing flexible question curation; we therefore report –. #Const. Docs ( k ) : number of source documents constituting each composite document (i.e., k+1 ) used in our evaluation ( Section 5.1 ). Src. Query Example : an example question on the original source (single document). Modified Query Example : its modified version for our agentic document QA evaluation. Red : modified text.
Dataset
Top-down: M↓
Bottom-up: M↑
Recall
Precision
F-1
Recall
Precision
F-1
Civic
0.96
(+6)
0.96
(+5)
0.95
(+6)
0.94
(+2)
0.93
(+0)
0.93
(+0)
Contracts
0.99
(+2)
0.99
(+11)
0.98
(+11)
0.94
(+12)
0.97
(+5)
0.94
(+11)
Finance
0.96
(+2)
0.95
(+0)
0.92
(+2)
0.71
(+2)
0.91
(+20)
0.76
(+11)
Papers
0.96
(+0)
0.96
(+1)
0.95
(+1)
0.86
(+4)
0.94
(+1)
0.88
(+4)
Avg.
0.97
(+3)
0.96
(+4)
0.95
(+5)
0.86
(+5)
0.94
(+6)
0.88
(+6)
Appendix
Table 8 . Robustness (Recall), compactness (Precision), and their F-1 of SHTs inferred by SHED from true headers. Parentheses indicate the percentage difference from SHED with extracted headers.
Figure 8 . Compactness vs. depth relative to true SHT. A point shows average compactness of SHTs in a depth ratio bin, from extracted (blue) or true (green) headers.
Approach
Recall (Robustness)
Precision (Compactness)
F-1 (Trade-off)
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Top-down: M↓(T)=avgv∈T′f(v) , where f(v) calculates recall, precision, and F-1 score of ts(v) relative to ts′(v) .
Deep
0.93
0.99
0.99
0.97
0.97
0.28
0.29
0.15
0.35
0.27
0.37
0.37
0.19
0.44
0.34
Wide
0.67
0.81
0.80
0.78
0.76
0.92
0.94
0.95
0.97
0.95
0.69
0.79
0.79
0.79
0.76
GROBID
0.35
0.79
0.71
0.78
0.66
0.17
0.34
0.71
0.83
0.51
0.18
0.38
0.62
0.77
0.49
LLM-text
0.96
0.75
0.92
0.97
0.90
0.97
0.60
0.89
0.97
0.86
0.96
0.58
0.86
0.96
0.84
Appendix
Table 9 . M↓ and M↑ results at composition size 1 (a single original document), in the same format as Table 3 .
Approach
Recall (Robustness)
Precision (Compactness)
F-1 (Trade-off)
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Civic
Contracts
Finance
Papers
Avg.
Top-down: M↓(T)=avgv∈T′f(v) , where f(v) calculates recall, precision, and F-1 score of ts(v) relative to ts′(v) .
Deep
0.94
0.99
1.00
0.97
0.97
0.20
0.12
0.08
0.14
0.13
0.28
0.16
0.10
0.20
0.19
Wide
0.72
0.82
0.82
0.78
0.78
0.93
0.90
0.97
0.97
0.94
0.74
0.73
0.82
0.79
0.77
GROBID
0.55
0.84
0.76
0.76
0.73
0.30
0.48
0.76
0.82
0.59
0.30
0.48
0.67
0.75
0.55
LLM-text
0.97
0.94
0.79
0.96
0.91
0.97
0.66
0.55
0.95
0.79
0.96
0.67
0.51
0.95
0.77
Appendix
Table 10 . M↓ and M↑ results at composition size 2(k+1) , in the same format as Table 3 .
Figure 9 . Average accuracy and total cost (USD) as the composition size (i.e., the number of concatenated documents) grows from the original (1) to twice the size used in Section 7 ( 2(k+1) ), where k varies by dataset ( Table 7 in Appendix D ). The panels replicate the experiments in Sections 7.2 (Figures 9(a) , 9(c) ) and 7.4 (Figures 9(b) , 9(d) ).
Datasets
#Docs
Avg. Doc Size
#Qs
Structure
Query Template
CFR ( Office of the Federal Register, National Archives and Records Administration, 2025 )
30
415,115
40
agency → regime → regulation
List all agencies whose regulations in regime R (e.g., FOIA, Privacy) relate to hypothesis H the same way (Entailment / Contradiction / NotMentioned) as agency X .
ETSI ( European Telecommunications Standards Institute, 2026 )
49
30,485
49
equipment category → characteristic → attribute
List all characteristics, spanning ≥2 equipment categories (transmitter / receiver / duplex), whose attribute A (Definition / Method-of-measurement / Limits) satisfies predicate P .
FERC ( Federal Energy Regulatory Commission, 2026 )
40
49,126
53
resource type → environmental effect → staff analysis
List all environmental effects whose resource type (e.g., aquatic, terrestrial, T&E species, recreation, cultural) differs from that of effect X but whose staff-analysis sentiment (positive / negative / neutral) is the same as X .
Appendix
Table 11 . Three new datasets, collected from inherently complex real-world documents. Avg. Doc Size: average word count per document in each dataset. #Docs (#Qs): number of documents (questions) in each dataset. Structure: document’s hierarchical structure, represented as section → subsection → subsubsection. Query template: natural-language template instantiated with entities in documents.
Approach
Recall
Precision
F-1
CFR
ETSI
FERC
Avg.
CFR
ETSI
FERC
Avg.
CFR
ETSI
FERC
Avg.
Top-down: M↓(T)
Wide
0.71
0.77
0.68
0.72
0.47
0.96
0.93
0.78
0.42
0.76
0.69
0.62
SHED
0.79
0.89
0.83
0.84
0.43
0.95
0.91
0.76
0.43
0.89
0.79
0.70
Bottom-up: M↑(T)
Wide
0.19
0.19
0.43
0.27
0.75
0.98
0.91
0.88
0.28
0.29
0.54
0.37
Appendix
Table 12 . Robustness (Recall), compactness (Precision), and their F-1 of SHTs inferred by Wide and SHED on the three new datasets in Section 8.3 . Green : better of the two approaches per dataset per metric.
Strategy
CFR
ETSI
FERC
Avg. (Tot.)
Vanilla-in-context
0.65
(111.1)
0.87
(4.8)
0.21
(12.3)
0.58
(128.3)
Vanilla-grep agent
0.67
(4.8)
0.66
(6.5)
0.23
(18.0)
0.52
(29.2)
SHT-based agent (Wide)
0.71
( 3.6 )
0.83
( 1.1 )
0.34
(3.0)
0.63
( 7.7 )
SHT-based agent ( SHED )
0.78
(4.3)
0.90
(1.7)
0.42
( 2.4 )
0.70
(8.5)
Appendix
Table 13 . Average QA accuracy and total cost (USD, inside parenthesis) on the three new datasets in Section 8.3 . Green : highest accuracy (lowest cost) per dataset.
Approach
Civic
Contracts
Finance
Papers
Est. Tot. (h)
GROBID
14
15
52
8
4
LLM-text
8
5
33
3
2
LLM-vision
13
9
93
11
5
SHED
22
18
48
13
5
SmolDocling-256M-preview
229
308
5,074
747
273
Appendix
Table 14 . SHT inference latency (average seconds per document), measured on 20 sampled composite documents per dataset. Est. Tot. : total latency (hours) for all 955 documents in Table 1 , as extrapolated estimates via the samples.
Dataset
Top-down: M↓
Bottom-up: M↑
Recall
Precision
F-1
Recall
Precision
F-1
Civic
0.66
( − 24)
0.88
( − 3)
0.69
( − 20)
0.26
( − 66)
0.86
( − 7)
0.40
( − 53)
Contracts
0.76
( − 21)
0.43
( − 45)
0.40
( − 47)
0.24
( − 58)
0.74
( − 18)
0.33
( − 50)
Finance
0.70
( − 24)
0.89
( − 6)
0.70
( − 20)
0.22
( − 47)
0.88
(+17)
0.33
( − 32)
Papers
0.63
( − 33)
0.83
( − 12)
0.66
( − 28)
0.26
( − 56)
0.87
( − 6)
0.35
( − 49)
Avg.
0.69
( − 25)
0.76
( − 16)
0.61
( − 29)
0.25
( − 56)
0.84
( − 3)
0.35
( − 46)
Appendix
Table 15 . Robustness (Recall), compactness (Precision), and their F-1 of SHTs inferred by SmolDocling-256M-preview, on 80 sampled documents (20 per dataset). Parentheses indicate the percentage difference from SHED ( Table 3 ).
Dataset
Top-down: M↓
Bottom-up: M↑
Recall
Precision
F-1
Recall
Precision
F-1
Civic
0.90
(+0)
0.91
(+0)
0.89
(+0)
0.92
(+0)
0.93
(+0)
0.93
(+0)
Contracts
0.97
(+0)
0.88
(+0)
0.87
(+0)
0.82
(+0)
0.92
(+0)
0.84
(+1)
Finance
0.94
(+0)
0.95
(+0)
0.90
(+0)
0.69
(+0)
0.70
( − 1)
0.65
(+0)
Papers
0.96
(+0)
0.95
(+0)
0.94
(+0)
0.82
(+0)
0.93
(+0)
0.84
(+0)
Avg.
0.94
(+0)
0.92
(+0)
0.90
(+0)
0.81
(+0)
0.87
(+0)
0.82
(+1)
Appendix
Table 16 . Robustness (Recall), compactness (Precision), and their trade-off (F-1) of SHTs inferred by SHED with the global-first SD inference. Parentheses: difference from local-first ( Table 3 ), in percentage points. Avg. : stratified averages across datasets.
Approach
Top-down: M↓
Bottom-up: M↑
Recall
Precision
F-1
Recall
Precision
F-1
global-first
0.97
0.94
0.92
0.94
0.81
0.84
local-first
1.00
0.96
0.96
1.00
0.84
0.89
Appendix
Table 17 . Robustness (Recall), compactness (Precision), and their F-1 of SHTs inferred by SHED with local- and global-first SD inference over 40 real-world documents ( et al.(2021), PI ) .
Leveraging large language models (LLMs) to analyze complex documents -- such as academic papers, technical manuals, and financial reports -- has emerged as a mainstream and critical task in both research and industry. In practice, users must first filter relevant documents from large collections and then conduct in-depth analysis (e.g. question answering) over the selected subset, yet existing systems flatten documents into plain-text chunks, discarding the rich hierarchical structures (sections, tables, figures, equations) and degrading downstream performance. We present DocMaster, a hierarchical structure-aware document analysis system. DocMaster parses documents into hierarchical document trees preserving original layouts and constructs a structure-aware semantic index that enables accurate document filtering and in-depth analysis. We demonstrate DocMaster through an interactive web interface that enables users to upload document collections, construct tree-based and multi-view semantic indices, filter relevant documents via natural-language conditions, and perform follow-up question answering over the filtered results. The source code, data, and demo are available at https://doc-master.github.io/.
Ziqi Chen, Yingli Zhou, Fangyuan Zhang +3
The Chinese University of Hong Kong, Shenzhen · The Chinese University of Hong Kong · OceanBase, AntGroup
Research artifacts are distributed primarily as reader-oriented documents like PDFs. This creates a bottleneck for increasingly agent-assisted and agent-native research workflows, in which LLM agents need to infer fine-grained, task-relevant information from lengthy full documents, a process that is expensive, repetitive, and unstable at scale. We introduce Knows, a lightweight companion specification that binds structured claims, evidence, provenance, and verifiable relations to existing research artifacts in a form LLM agents can consume directly. Knows addresses the gap with a thin YAML sidecar (KnowsRecord) that coexists with the original PDF, requiring no changes to the publication itself, and validated by a deterministic schema linter. We evaluate Knows on 140 comprehension questions across 20 papers spanning 14 academic disciplines, comparing PDF-only, sidecar-only, and hybrid conditions across six LLM agents of varying capacity. Weak models (0.8B--2B parameters) improve from 19--25% to 47--67% accuracy (+29 to +42 percentage points) when reading sidecar instead of PDF, while consuming 29--86% fewer input tokens; an LLM-as-judge re-scoring confirms that weak-model sidecar accuracy (75--77%) approaches stronger-model PDF accuracy (78--83%). Beyond this controlled evaluation, a community sidecar hub at https://knows.academy/ has already indexed over ten thousand publications and continues to grow daily, providing independent evidence that the format is adoption-ready at scale.
Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.