As language-model agents become more capable of iterative search, corpus access is shifting from retrieval toward interaction. Agents can explore the corpus, inspect documents, and use newly discovered evidence to decide what to examine next. Yet accessible evidence may still fail to become usable within a finite interaction budget. We call this progressive failure Evidence Blindness: supporting documents may never enter view, may remain unopened, or may fail to expose the decisive evidence even after being opened. A key reason is that agents often have to infer useful evidence directions during interaction, spending limited budget on deciding where to search next. Existing approaches either leave corpus structure largely implicit or reconstruct useful directions at query time. AtlasNav instead organizes reusable cross-document structure before any query arrives. It builds a persistent multi-view Corpus Atlas, which each query can navigate adaptively while still accessing the original documents directly. On BrowseComp-Plus, AtlasNav outperforms the previous state-of-the-art interactive corpus access method across different backbones. On DeepSeek, it improves strict accuracy by 7.47 points while reducing query-time inference cost by 30.22%.AtlasNav also reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.
Figures & tables
Figure 1: Evidence realization on BrowseComp-Plus with DeepSeek-V4-Flash. (a) Complete evidence realized per checkpoint. (b) Strict accuracy given complete vs. incomplete realization.
Figure 2: AtlasNav and Evidence Blindness. (a) Offline construction of the persistent multi-view Corpus Atlas. (b) Query-adaptive finite-budget navigation. (c) Evidence Blindness diagnosis across the four evidence-realization checkpoints.
Strict Accuracy (%) ↑
Normalized Query-Time Agent Cost ↓
Backbone
DCI
DR-DCI
AtlasNav
Ref.
Δ Acc.
DCI
DR-DCI
AtlasNav
Cost save (%)
GPT-5.6
87.35
87.83
91.81
96.51
+3.98
1.718
1.173
1.000
14.74
DeepSeek
81.45
84.58
92.05
96.51
+7.47
2.347
1.433
1.000
30.22
Qwen
36.99
50.96
72.53
95.42
+21.57
2.243
1.271
1.000
21.31
MiMo
62.89
71.08
79.04
96.14
+7.96
2.579
1.006
1.000
0.62
Table 1: Accuracy and recorded query-time agent inference cost on BrowseComp-Plus. Costs are normalized within each backbone to AtlasNav =1.000 ; accuracy gains and cost savings are relative to DR-DCI.
Surface ↓
Open ↓
Locate ↓
Backbone
Interface
EBAnyS
EBAllS
EBAnyO
EBAllO
EBAnyL
EBAllL
GPT-5.6
DCI
6.63
6.99
12.89
13.01
15.90
16.27
DR-DCI
9.40
9.76
10.48
11.08
15.90
16.63
AtlasNav
5.66
6.02
9.04
9.52
11.93
12.65
DeepSeek
DCI
13.13
13.37
14.70
15.42
17.83
18.80
DR-DCI
13.49
14.10
16.51
17.11
23.61
24.46
Table 2: Evidence Blindness (EB) on BrowseComp-Plus (%); lower is better.
Figure 3: Finite-budget empirical reference gaps ΔI(B) under turn and query-time inference-cost budgets for GPT-5.6 and DeepSeek; lower is better. Results for Qwen and MiMo are provided in Appendix B .
Figure 6
System
Overall ↑
Correct. (%) ↑
Complete. (%) ↑
Doc. Recall (%) ↑
Invalid Extra ↓
AtlasNav
87.00
94.40
89.14
87.05
0.66
Mixedbread (+ Opus 5)
86.58
89.80
90.62
94.28
0.39
Prism (aurait.ai)
80.46
85.60
86.26
87.81
3.31
metor.com
80.34
82.00
86.22
85.53
4.96
ZNV AgentCube
80.26
82.40
86.56
86.10
2.94
Skyller AI
79.30
81.60
87.39
86.50
8.74
Table 4: EnterpriseRAG-Bench metrics with representative public systems.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Interface
Surface
Open
Locate
GPT-5.6-Luna
DCI
6.81
12.93
16.02
DR-DCI
9.58
10.78
16.27
AtlasNav
5.84
9.28
12.29
DeepSeek
DCI
13.25
15.04
18.29
DR-DCI
13.80
16.81
24.04
AtlasNav
4.90
7.73
11.08
Appendix
Table 5: Mean endpoint Evidence Blindness on BrowseComp-Plus (%); lower is better. Any-slot and All-slot results are reported in the main text. Construction is omitted because it is saturated for all evaluated interfaces.
Strict Accuracy (%) ↑
EBAllS/O/L (%) ↓
Backbone
Cost
DCI
DR-DCI
AtlasNav
DCI
DR-DCI
AtlasNav
GPT-5.6-Luna
0.025
45.66
65.42
66.39
24.70/36.14/40.60
17.47/ 20.12 /29.88
16.51 /23.61/ 26.75
0.05
57.83
74.82
75.90
20.24/30.00/34.46
14.10/ 16.02 /23.73
13.37 /18.67/ 21.57
0.15
69.52
81.20
82.05
13.49/22.77/26.51
12.77/ 14.46 /20.84
10.12 /15.30/ 18.07
0.65
84.46
86.39
91.81
7.95/14.58/18.07
10.12/11.57/17.47
6.02 / 9.52 / 12.65
DeepSeek
0.15
49.16
50.36
70.96
43.73/49.88/54.46
29.40/36.02/45.18
23.25 / 30.36 / 35.90
Appendix
Table 7: BrowseComp-Plus results at within-backbone query-time agent-cost checkpoints. Each row is one cost budget for one backbone; the first block reports strict accuracy (%) and the second checkpoint-wise EBAllS/O/L (%), with each EB entry as Surface / Open / Locate. Lower is better for EB.
Scale
Files
Interface
EBAllO↓
Surface-to-Open (%) ↑
10K
10,059
DCI
38.0
93.94
DR-DCI
76.5
95.92
AtlasNav
45.5
80.15
50K
50,324
DCI
52.0
84.21
DR-DCI
76.5
79.66
AtlasNav
52.5
77.24
Appendix
Table 8: PhantomWiki evidence realization beyond the main text. Questions and supporting evidence are fixed across nested corpus sizes. Surface-to-Open is the fraction of complete-Surface questions that also reach complete Open.
Category
n
Overall ↑
Correct. ↑
Complete. ↑
Doc. Recall ↑
Invalid Extra ↓
Basic
175
95.16
98.86
95.16
93.71
0.26
Semantic
125
80.73
89.60
82.25
80.00
0.80
Intra-document reasoning
40
96.46
100.00
96.46
97.50
0.05
Project related
40
59.09
87.50
67.88
65.48
2.77
Constrained
30
94.57
96.67
96.79
95.00
0.70
Conflicting information
20
92.69
100.00
92.69
80.00
0.20
Appendix
Table 9: Complete AtlasNav metrics by EnterpriseRAG-Bench question category.
Source tag
n
Overall ↑
Correct. ↑
Complete. ↑
Doc. Recall ↑
Invalid Extra ↓
Confluence
115
78.00
91.30
81.68
80.92
1.37
Fireflies
24
87.22
95.83
87.22
74.58
0.38
GitHub
60
82.80
93.33
86.07
82.67
0.98
Gmail
55
86.74
94.55
89.58
86.00
0.45
Google Drive
59
88.70
96.61
89.53
83.92
1.05
HubSpot
33
92.42
96.97
92.93
96.97
0.03
Appendix
Table 10: Complete AtlasNav metrics by overlapping EnterpriseRAG-Bench source tag. Counts do not sum to the benchmark size because questions may have multiple source tags.
Item
Count
Questions
830
Required evidence slots
889
Accepted canonical fragments
1,753
Canonical documents represented
1,445
Questions passing automatic validation
828
Questions requiring manual adjudication
2
Appendix
Table 11: Frozen BrowseComp-Plus fragment-level Qrel used to operationalize Answer-Evidence Locate.
multiplex Topic (1.0) + Identity (0.75); 77 parent regions
Fine hierarchy
conditional Episode + Relation within each parent; 443 leaves
Addressing
100,195/100,195 documents receive one parent–leaf address; canonical files remain directly searchable and openable
Appendix
Table 12: Frozen Atlas construction configuration on BrowseComp-Plus.
Item
Setting
Tasks
1,913 single; 4,300 pair; 950 triple; 7,163 total
Parent-disjoint split
4,972 train; 1,167 validation; 1,024 test
Input/model
816 features; 4,902-parameter linear three-head router
Output
Topic, Identity, Episode, Relation, and BM25 channel weights
Objective
all-positive ranking + worst-positive bottleneck and calibration losses
Safety calibration
0.05 floor on each view’s share of semantic mass; low confidence shrinks toward uniform semantic weights and a 1:1 semantic/BM25 mass
Appendix
Table 13: Corpus-derived router data and frozen model configuration.
Figure 5: BrowseComp-Plus Locate Evidence Blindness EBAllL over interaction turns for DCI, DR-DCI, and AtlasNav; lower is better. Exact checkpoint values are reported in Appendix B .
Backbone
AtlasNav Acc.
DR-DCI Acc.
Difference
Paired bootstrap 95% CI
Exact McNemar
DeepSeek
92.05%
84.58%
+7.47 pp
[4.94,10.00] pp
p=1.22×10−8
GPT-5.6-Luna
91.81%
87.83%
+3.98 pp
[1.69,6.39] pp
p=0.00119
Appendix
Table 14: Paired uncertainty for the AtlasNav–DR-DCI endpoint comparison on the two backbones with retained per-question records. Confidence intervals are fixed-seed paired bootstrap intervals; p -values are from exact McNemar tests.
Figure 6: Finite-budget empirical reference gaps ΔI(B) for Qwen and MiMo under turn and query-time inference-cost budgets. These panels complement Figure 3 in the main text; lower is better.
Parent views
Child views
Accuracy ↑
Mean turns ↓
Topic + Identity
Episode + Relation
95.78
27.43
Topic + Episode
Identity + Relation
89.16
37.33
Topic + Relation
Identity + Episode
91.57
35.10
Identity + Episode
Topic + Relation
91.57
35.88
Identity + Relation
Topic + Episode
89.76
31.89
Episode + Relation
Topic + Identity
88.55
36.02
Appendix
Table 15: Hierarchical view-assignment ablation on a shared 166-question subset. All variants use the same corpus, view graphs, router, agent, and observation budget and differ only in parent–child view assignment.
Interface
P@10
L@10
P@20
L@20
P@50
L@50
DCI
7.77
8.64
12.96
15.77
22.83
33.84
DR-DCI
3.95
5.50
5.97
8.85
9.89
15.80
AtlasNav
9.00
9.99
16.54
18.64
23.53
31.31
Appendix
Table 16: Mean number of distinct Atlas parent (P) and leaf (L) regions under matched surfaced-file budgets on BrowseComp-Plus.
Case
Evidence state
Outcome
Diagnostic
8 ( Lush Life )
AtlasNav and DCI realize the complete annotated evidence; DR-DCI does not realize the annotated chain.
AtlasNav correct; DCI and DR-DCI wrong.
Evidence access and downstream reasoning are distinct failure modes.
394 (wing three-quarter)
AtlasNav reaches only part of the annotated supporting-document set but realizes the complete decision-relevant slot-level evidence; the baselines do not.
AtlasNav correct; baselines wrong.
Complete support-document coverage is not required once the decisive evidence slots are realized.
417 (Owlman)
AtlasNav realizes the complete annotated evidence; DR-DCI realizes only part of it.
AtlasNav wrong; DR-DCI correct.
Complete annotated evidence does not guarantee correct synthesis, and partial annotated realization does not preclude a correct answer.
Appendix
Table 18: Representative BrowseComp-Plus trajectories illustrating the distinction between evidence realization and final-answer correctness.
Question category
Overall
Source tag
Overall
Miscellaneous
96.08
HubSpot
92.42
Intra-document
96.46
Gmail
86.74
Constrained
94.57
Jira
81.21
Basic
95.16
Google Drive
88.70
Conflicting information
92.69
Fireflies
87.22
Information not found
100.00
Linear
84.23
Appendix
Table 19: EnterpriseRAG-Bench Overall score by question category and overlapping source tag. Source tags are not mutually exclusive; the full-set Overall score is 87.00. Complete metric matrices are reported in Appendix B .
Paradigm
Representative methods
Persistent structure
Query-specific structure
Continued interaction
Primary role of structure
Sparse/dense retrieval
BM25; dense retrievers
Index only
No
Limited
Select top- k evidence
Retrieval + reranking
ReasonRank ( Liu et al., 2026b ) ; neural rerankers
Index
No
Limited
Refine evidence ranking
Hierarchical/ graph RAG
RAPTOR; HippoRAG; SiReRAG; LinearRAG ( Zhuang et al., 2026a )
Yes
Usually no
Retrieval-mediated
Structured evidence selection
Trained retrieval agents
Search-R1; R1-Searcher ( Song et al., 2025 ) ; ASearcher ( Gao et al., 2025 )
Retrieval index
Adaptive querying
Via search API
Improve search policy
Relevance-guided interaction
RARG ( Li et al., 2026a )
Retrieval index
Query-conditioned priority
Yes
Guide traversal and local match visibility
Raw direct interaction
DCI
No
No
Yes
Direct corpus operations
Appendix
Table 20: Conceptual corpus-access design space. The taxonomy positions methods by interface semantics rather than comparing performance across heterogeneous experimental stacks.
Interface
Strict Acc. ↑
Cost ↓
Turns ↓
DCI
95.50
22.59
4,117
DR-DCI
94.50
28.01
5,540
AtlasNav
94.50
10.89
3,135
Appendix
Table 21: Exact 2Wiki-Global results on 400 questions over a shared 56,684-document corpus. Cost is total recorded query-time agent inference cost in CNY.
Interface
Loose ↑
Strict ↑
Cost ↓
Turns ↓
DCI
79.65
39.35
73.35
6,596
DR-DCI
81.42
44.19
43.24
6,493
AtlasNav
80.10
40.32
36.49
4,218
Appendix
Table 22: Exact FanOutQA endpoint results on 310 questions. Loose and Strict are the official answer-coverage metrics; cost is recorded query-time agent inference cost in CNY.
Figure 7: Deliverable nDCG@10 over frozen, data-driven turn checkpoints on TREC-COVID. Unfinished queries score zero at each checkpoint. The evidence-supplied reference receives the positive-document pool but not graded labels or the ideal ranking.
Symbol
Meaning
Symbol
Meaning
q
Question or query.
Q
Set of evaluation questions.
Eq={e1,…,emq}
Required evidence slots for q ; mq is their number.
τ∈T={C,S,O,L}
Construction, Surface, Open, or Locate checkpoint.
Eqτ
Evidence slots realized by checkpoint τ .
rqτ
Fraction of required slots realized at τ .
EBAny/Mean/Allτ
Evidence Blindness metrics for total failure, mean loss, and incomplete realization.
C
Original accessible corpus.
M
Language-model agent.
I
Corpus interface.
B
Interaction budget.
AI(B)
Strict task performance of I under budget B .
Appendix
Table 23: Principal notation used in Sections 3–4.
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.
Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang +1
Large Language Model (LLM) search agents have shown strong promise on knowledge-intensive tasks through iterative reasoning and retrieval. Most existing systems rely on retrievers that return ranked documents from a pre-built index. We explore a complementary paradigm in which the agent treats the corpus as the search environment and finds evidence through executable shell commands. We introduce GrepSeek, an optimized direct corpus interaction (DCI) agent that learns to find, filter, and compose evidence over large text corpora. To stabilize reinforcement learning (RL) over large corpora, we train in two stages: first, we initialize the policy using verified, causally grounded search trajectories generated by an answer-aware Tutor and an answer-blind Planner; then, we refine the policy using Group Relative Policy Optimization (GRPO). To make DCI practical at scale, we introduce two semantics-preserving execution optimizations: Pruned Adaptive Command Execution, which reduces shell-based search latency by up to 77× on a 14GB corpus with 21 million documents using a compact auxiliary structure, and Sharded-Parallel Corpus Search, which achieves up to 7.6× speedup without additional preprocessing; both preserve equivalence with sequential execution. Across eight open-domain QA benchmarks, GrepSeek achieves the strongest overall performance, with a statistically significant relative improvement of 5.7% over the best baseline. Our analysis shows how DCI-optimized agents conduct flexible and effective compositional search through direct corpus interaction.
Alireza Salemi, Chang Zeng, Atharva Nijasure +4
University of Massachusetts Amherst · Princeton University · Carnegie Mellon University
Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction (DCI) enables such fine-grained operations through grep-style exploration, but its relevance-agnostic search can expose useful clues late and delay convergence. Recent advances use relevance to narrow the corpus into a working space for interaction. Once interaction begins, however, relevance still does not directly guide which documents grep searches first or distinguish informative excerpts from a broad set of matches to let LLMs see them first. We introduce the Relevance-Aware RipGrep Search Agent (RARG), which turns relevance into an execution prior for corpus interaction. RARG provides coarse-to-fine relevance guidance: it orders documents for sequential 'ripgrep' traversal to expose globally relevant clues earlier, initializes promising entry points with query-relevant paragraphs, and reranks grep matches to surface informative excerpts that document-level ranking may otherwise obscure. Across challenging browse question answering and reasoning-intensive retrieval, RARG improves the accuracy--efficiency frontier over retrieval-based and direct-interaction agents. These results demonstrate that relevance-aware interaction enables faster and more reliable search convergence.