Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.
Figures & tables
EnterpriseRAG-Bench
WixQA
HERB
Overall
Method
Doc. Rec. ↑
Correct. ↑
Complete. ↑
Tokens ↓
Ctx. Rec. ↑
Fact. ↑
Tokens ↓
Content ↑
Tokens ↓
Quality ↑
Rel. Tok. ↓
GPT-5.5
Raw Corpus
61.62 ± 1.66
62.08 ± 4.12
73.11 ± 0.98
206.5k
73.31 ± 0.75
67.51 ± 0.39
337.2k
62.31 ± 0.59
598.4k
66.11
1.00 ×
Document Page
64.18 ± 1.55
57.50 ± 1.02
68.97 ± 1.83
175.6k
77.85 ± 1.18
66.98 ± 1.90
288.0k
63.26 ± 0.14
680.9k
66.41
0.94 ×
Group Page
64.75 ± 0.96
48.75 ± 2.70
62.62 ± 2.58
423.0k
60.86 ± 3.68
62.24 ± 4.45
1,125.8k
62.97 ± 1.66
392.9k
61.07
1.65 ×
LLM Wiki
63.15 ± 1.55
56.67 ± 2.36
67.74 ± 1.77
192.4k
73.73 ± 0.52
65.19 ± 0.26
140.9k
61.49 ± 0.76
387.0k
64.49
0.63 ×
Corpus2Skill
49.39 ± 2.32
47.50 ± 2.70
57.29 ± 1.01
88.8k
65.93 ± 2.07
60.34 ± 0.39
76.0k
21.94 ± 0.79
180.9k
45.49
0.31 ×
Table 1: Main results across EnterpriseRAG-Bench, WixQA, and HERB with GPT-5.5 and GPT-5.6 models, as means ± standard deviations over three runs. Overall reports dataset-balanced quality and geometric-mean input-token ratios to Raw Corpus within each LLM. Best and second-best effectiveness scores per LLM among retrieval methods are bolded and underlined .
DeepSeek
MAI
Method
Quality
Tokens
Quality
Tokens
Raw Corpus
68.02
1,349.4k
34.06
129.6k
Document Page
43.18
4,047.0k
38.56
208.4k
Group Page
59.24
1,965.5k
32.81
421.7k
LLM Wiki
57.99
3,523.8k
35.57
109.8k
Corpus2Skill
49.72
628.7k
21.33
120.4k
Table 2: Overall Quality with LLMs from other model families on EnterpriseRAG-Bench.
DeepSeek
MAI
Method
Quality
Tokens
Quality
Tokens
Raw Corpus
68.02
1,349.4k
34.06
129.6k
Document Page
43.18
4,047.0k
38.56
208.4k
Group Page
59.24
1,965.5k
32.81
421.7k
LLM Wiki
57.99
3,523.8k
35.57
109.8k
Corpus2Skill
49.72
628.7k
21.33
120.4k
Table 2: Overall Quality with LLMs from other model families on EnterpriseRAG-Bench.
Quality
Method
GPT-5.5
Luna
BM25
63.66
60.78
Dense
65.05
58.84
HippoRAG
64.66
61.74
GraphRAG
47.95
43.46
Raw Corpus
65.60
63.31
Table 3: Retrieval-based approaches on EnterpriseRAG-Bench.
Map Builder
GPT-5.5
Sol
Luna
DeepSeek
Cost
Raw Corpus
65.60
69.54
63.31
68.02
–
GPT-5.5
76.60
75.31
73.79
68.51
≤ $4,732.14
Sol
77.39
78.03
73.60
71.06
≤ $2,681.76
Luna
73.59
74.68
73.20
70.90
≤ $74.65
DeepSeek
73.06
73.11
69.42
71.11
$310.58
Table 4: Overall Quality of map reuse across LLMs on EnterpriseRAG-Bench. Rows denote the map builder, columns the answering LLM, and Cost the one-time construction cost.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Doc. Rec. ↑
Correct. ↑
Complete. ↑
Tokens ↓
Raw Corpus
89.74
89.05
86.42
250.7k
CorpusMap (Ours)
95.38
94.76
89.58
171.7k
Appendix
Table 5: Results on the remaining questions of EnterpriseRAG-Bench with GPT-5.5. Doc. Rec. is computed over the questions with gold documents.
Baseline
GPT-5.5
Luna
Terra
Sol
Raw Corpus
+6.45 [4.30, 8.62]
+11.74 [9.40, 14.07]
+6.76 [4.34, 9.19]
+7.00 [5.05, 8.94]
Document Page
+6.14 [4.05, 8.26]
+9.86 [7.63, 12.08]
+9.60 [7.28, 11.93]
+8.37 [6.24, 10.50]
Group Page
+11.48 [8.93, 14.01]
+18.03 [15.37, 20.70]
+28.67 [25.82, 31.48]
+43.46 [40.62, 46.26]
LLM Wiki
+8.07 [5.81, 10.35]
+11.20 [8.61, 13.75]
+13.79 [11.05, 16.52]
+9.54 [7.23, 11.85]
Corpus2Skill
+27.07 [24.37, 29.80]
+28.32 [25.19, 31.41]
+28.65 [25.69, 31.61]
+30.25 [27.38, 33.12]
Appendix
Table 6: Gains in Overall Quality of CorpusMap over each baseline in Table 1 , with 95% confidence intervals from a paired bootstrap over questions. All gains are significant with Holm-corrected p<10−4 .
EnterpriseRAG-Bench
WixQA
HERB
Method
Correct.
Complete.
Ctx. Rec.
Fact.
Content
Raw Corpus
83.75
78.11
50.95
67.09
65.04
Document Page
83.33
73.61
58.12
67.72
65.56
Group Page
76.25
68.90
35.13
62.55
65.53
LLM Wiki
80.42
73.10
58.54
67.09
64.45
Corpus2Skill
67.50
61.76
60.55
61.71
29.72
Appendix
Table 7: LLM-judged metrics of the GPT-5.5 answers in Table 1 , judged by DeepSeek-V4-Pro instead of GPT-5.6 Sol. The last row reports Kendall’s τ between the rankings of the retrieval methods under the two judges.
Relevant
Random
Method
Quality
Tokens
Quality
Tokens
Raw Corpus
65.60
206.5k
–
–
Raw Corpus w/ Candidates
66.19
107.4k
–
–
CorpusMap
w/o Candidates
69.04
496.4k
–
–
w/ Candidate Entity Pages
76.68
247.1k
68.65
404.6k
w/ Candidate Entity-Linked Docs
76.60
88.1k
60.76
413.5k
Appendix
Table 8: Overall Quality with candidate file paths on EnterpriseRAG-Bench with GPT-5.5.
Question
List every internal communication thread (email, Slack, and meeting notes) about the Redwood Private upgrade ‘rollback loop’ bug (including references to RRB-17 or ‘stuck rollback’).
Gold Documents
D1: #eng thread (Slack) D2: INC-2147 thread (Slack) D3: customer email (Gmail) D4: RRB-17 root cause and patch plan (Gmail) D5: escalation meeting (Fireflies)
Path: RRB-17 → workaround deletes → installer-rollback-lock Extracted from D4: “Workaround (current): delete CM installer-rollback-lock and restart installer-controller.”
Appendix
Table 9: Case study of entity–entity edges on EnterpriseRAG-Bench. Blue and orange boxes denote documents and Entity Pages, as in Figure 2 .
Question
In the Deterministic Playback Manifest v1, how is manifest signing/integrity represented (signature vs. integrity fields)?
Gold Answer
In v1, the manifest does not embed a signature blob; integrity is expressed via an optional integrity field and an integrity_ref URI. The earlier draft instead listed an embedded signature field. Gold documents: D1: manifest draft (Google Drive) D2: manifest v1 (Confluence)
Corpus2Skill
Explored: ROOT → all four top-level skills → several clusters below them; D1 and D2 both sit under one cluster that the agent does not open. Lookups: “Deterministic Playback” (no match), “manifest” (an unrelated cluster), “integrity” (no match). Answer: “I couldn’t locate a document for ‘Deterministic Playback Manifest v1’ in the available corpus.” ✗
CorpusMap (Ours)
Path: ENG-8192 → D1 → Serving Runtime → D2 (1) Searches Entity Pages for “playback” and opens ENG-8192 , whose linked documents include D1. (2) Reads D1 : “signature: optional signed blob for integrity verification.” (3) Searches Entity Pages with these terms and opens Serving Runtime , whose linked documents include D2. (4) Reads D2 : “ signature is no longer a direct embedded blob in v1.” Answer: v1 uses an optional integrity field and an integrity_ref URI rather than an embedded signature. ✓
Appendix
Table 10: Case study comparing CorpusMap with Corpus2Skill on EnterpriseRAG-Bench. Blue and orange boxes denote documents and Entity Pages, as in Figure 2 .