AtlasNav: Mitigating Evidence Blindness with Persistent Corpus Navigation
Abstract
As language-model agents become more capable of iterative search, corpus access is shifting from retrieval toward interaction. Agents can explore the corpus, inspect documents, and use newly discovered evidence to decide what to examine next. Yet accessible evidence may still fail to become usable within a finite interaction budget. We call this progressive failure Evidence Blindness: supporting documents may never enter view, may remain unopened, or may fail to expose the decisive evidence even after being opened. A key reason is that agents often have to infer useful evidence directions during interaction, spending limited budget on deciding where to search next. Existing approaches either leave corpus structure largely implicit or reconstruct useful directions at query time. AtlasNav instead organizes reusable cross-document structure before any query arrives. It builds a persistent multi-view Corpus Atlas, which each query can navigate adaptively while still accessing the original documents directly. On BrowseComp-Plus, AtlasNav outperforms the previous state-of-the-art interactive corpus access method across different backbones. On DeepSeek, it improves strict accuracy by 7.47 points while reducing query-time inference cost by 30.22%.AtlasNav also reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.
Figures & tables
| Strict Accuracy (%) | Normalized Query-Time Agent Cost | ||||||||
| Backbone | DCI | DR-DCI | AtlasNav | Ref. | Acc. | DCI | DR-DCI | AtlasNav | Cost save (%) |
| GPT-5.6 | 87.35 | 87.83 | 91.81 | 96.51 | 1.718 | 1.173 | 1.000 | 14.74 | |
| DeepSeek | 81.45 | 84.58 | 92.05 | 96.51 | 2.347 | 1.433 | 1.000 | 30.22 | |
| Qwen | 36.99 | 50.96 | 72.53 | 95.42 | 2.243 | 1.271 | 1.000 | 21.31 | |
| MiMo | 62.89 | 71.08 | 79.04 | 96.14 | 2.579 | 1.006 | 1.000 | 0.62 | |
| Surface | Open | Locate | |||||
|---|---|---|---|---|---|---|---|
| Backbone | Interface | ||||||
| GPT-5.6 | DCI | 6.63 | 6.99 | 12.89 | 13.01 | 15.90 | 16.27 |
| DR-DCI | 9.40 | 9.76 | 10.48 | 11.08 | 15.90 | 16.63 | |
| AtlasNav | 5.66 | 6.02 | 9.04 | 9.52 | 11.93 | 12.65 | |
| DeepSeek | DCI | 13.13 | 13.37 | 14.70 | 15.42 | 17.83 | 18.80 |
| DR-DCI | 13.49 | 14.10 | 16.51 | 17.11 | 23.61 | 24.46 | |
| System | Overall | Correct. (%) | Complete. (%) | Doc. Recall (%) | Invalid Extra |
|---|---|---|---|---|---|
| AtlasNav | 87.00 | 94.40 | 89.14 | 87.05 | 0.66 |
| Mixedbread (+ Opus 5) | 86.58 | 89.80 | 90.62 | 94.28 | 0.39 |
| Prism (aurait.ai) | 80.46 | 85.60 | 86.26 | 87.81 | 3.31 |
| metor.com | 80.34 | 82.00 | 86.22 | 85.53 | 4.96 |
| ZNV AgentCube | 80.26 | 82.40 | 86.56 | 86.10 | 2.94 |
| Skyller AI | 79.30 | 81.60 | 87.39 | 86.50 | 8.74 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Backbone | Interface | Surface | Open | Locate |
|---|---|---|---|---|
| GPT-5.6-Luna | DCI | 6.81 | 12.93 | 16.02 |
| DR-DCI | 9.58 | 10.78 | 16.27 | |
| AtlasNav | 5.84 | 9.28 | 12.29 | |
| DeepSeek | DCI | 13.25 | 15.04 | 18.29 |
| DR-DCI | 13.80 | 16.81 | 24.04 | |
| AtlasNav | 4.90 | 7.73 | 11.08 |
| Strict Accuracy (%) | (%) | ||||||
| Backbone | Cost | DCI | DR-DCI | AtlasNav | DCI | DR-DCI | AtlasNav |
| GPT-5.6-Luna | 0.025 | 45.66 | 65.42 | 66.39 | 24.70/36.14/40.60 | 17.47/ 20.12 /29.88 | 16.51 /23.61/ 26.75 |
| 0.05 | 57.83 | 74.82 | 75.90 | 20.24/30.00/34.46 | 14.10/ 16.02 /23.73 | 13.37 /18.67/ 21.57 | |
| 0.15 | 69.52 | 81.20 | 82.05 | 13.49/22.77/26.51 | 12.77/ 14.46 /20.84 | 10.12 /15.30/ 18.07 | |
| 0.65 | 84.46 | 86.39 | 91.81 | 7.95/14.58/18.07 | 10.12/11.57/17.47 | 6.02 / 9.52 / 12.65 | |
| DeepSeek | 0.15 | 49.16 | 50.36 | 70.96 | 43.73/49.88/54.46 | 29.40/36.02/45.18 | 23.25 / 30.36 / 35.90 |
| Scale | Files | Interface | Surface-to-Open (%) | |
|---|---|---|---|---|
| 10K | 10,059 | DCI | 38.0 | 93.94 |
| DR-DCI | 76.5 | 95.92 | ||
| AtlasNav | 45.5 | 80.15 | ||
| 50K | 50,324 | DCI | 52.0 | 84.21 |
| DR-DCI | 76.5 | 79.66 | ||
| AtlasNav | 52.5 | 77.24 |
| Category | Overall | Correct. | Complete. | Doc. Recall | Invalid Extra | |
|---|---|---|---|---|---|---|
| Basic | 175 | 95.16 | 98.86 | 95.16 | 93.71 | 0.26 |
| Semantic | 125 | 80.73 | 89.60 | 82.25 | 80.00 | 0.80 |
| Intra-document reasoning | 40 | 96.46 | 100.00 | 96.46 | 97.50 | 0.05 |
| Project related | 40 | 59.09 | 87.50 | 67.88 | 65.48 | 2.77 |
| Constrained | 30 | 94.57 | 96.67 | 96.79 | 95.00 | 0.70 |
| Conflicting information | 20 | 92.69 | 100.00 | 92.69 | 80.00 | 0.20 |
| Source tag | Overall | Correct. | Complete. | Doc. Recall | Invalid Extra | |
|---|---|---|---|---|---|---|
| Confluence | 115 | 78.00 | 91.30 | 81.68 | 80.92 | 1.37 |
| Fireflies | 24 | 87.22 | 95.83 | 87.22 | 74.58 | 0.38 |
| GitHub | 60 | 82.80 | 93.33 | 86.07 | 82.67 | 0.98 |
| Gmail | 55 | 86.74 | 94.55 | 89.58 | 86.00 | 0.45 |
| Google Drive | 59 | 88.70 | 96.61 | 89.53 | 83.92 | 1.05 |
| HubSpot | 33 | 92.42 | 96.97 | 92.93 | 96.97 | 0.03 |
| Item | Count |
|---|---|
| Questions | 830 |
| Required evidence slots | 889 |
| Accepted canonical fragments | 1,753 |
| Canonical documents represented | 1,445 |
| Questions passing automatic validation | 828 |
| Questions requiring manual adjudication | 2 |
| Component | Setting |
|---|---|
| View signatures | Topic (6,144 characters); Identity, Episode, Relation (4,096 characters each) |
| Encoding | 2,560-dimensional unit vectors; independent 192-dimensional PCA projections fitted on 50,000 deterministic samples |
| Sparse graphs | cosine/IP HNSW; , , efConstruction , efSearch |
| Coarse hierarchy | multiplex Topic (1.0) + Identity (0.75); 77 parent regions |
| Fine hierarchy | conditional Episode + Relation within each parent; 443 leaves |
| Addressing | 100,195/100,195 documents receive one parent–leaf address; canonical files remain directly searchable and openable |
| Item | Setting |
|---|---|
| Tasks | 1,913 single; 4,300 pair; 950 triple; 7,163 total |
| Parent-disjoint split | 4,972 train; 1,167 validation; 1,024 test |
| Input/model | 816 features; 4,902-parameter linear three-head router |
| Output | Topic, Identity, Episode, Relation, and BM25 channel weights |
| Objective | all-positive ranking + worst-positive bottleneck and calibration losses |
| Safety calibration | 0.05 floor on each view’s share of semantic mass; low confidence shrinks toward uniform semantic weights and a 1:1 semantic/BM25 mass |
| Backbone | AtlasNav Acc. | DR-DCI Acc. | Difference | Paired bootstrap 95% CI | Exact McNemar |
|---|---|---|---|---|---|
| DeepSeek | 92.05% | 84.58% | pp | pp | |
| GPT-5.6-Luna | 91.81% | 87.83% | pp | pp |
| Parent views | Child views | Accuracy | Mean turns |
|---|---|---|---|
| Topic + Identity | Episode + Relation | 95.78 | 27.43 |
| Topic + Episode | Identity + Relation | 89.16 | 37.33 |
| Topic + Relation | Identity + Episode | 91.57 | 35.10 |
| Identity + Episode | Topic + Relation | 91.57 | 35.88 |
| Identity + Relation | Topic + Episode | 89.76 | 31.89 |
| Episode + Relation | Topic + Identity | 88.55 | 36.02 |
| Interface | P@10 | L@10 | P@20 | L@20 | P@50 | L@50 |
|---|---|---|---|---|---|---|
| DCI | 7.77 | 8.64 | 12.96 | 15.77 | 22.83 | 33.84 |
| DR-DCI | 3.95 | 5.50 | 5.97 | 8.85 | 9.89 | 15.80 |
| AtlasNav | 9.00 | 9.99 | 16.54 | 18.64 | 23.53 | 31.31 |
| Case | Evidence state | Outcome | Diagnostic |
|---|---|---|---|
| 8 ( Lush Life ) | AtlasNav and DCI realize the complete annotated evidence; DR-DCI does not realize the annotated chain. | AtlasNav correct; DCI and DR-DCI wrong. | Evidence access and downstream reasoning are distinct failure modes. |
| 394 (wing three-quarter) | AtlasNav reaches only part of the annotated supporting-document set but realizes the complete decision-relevant slot-level evidence; the baselines do not. | AtlasNav correct; baselines wrong. | Complete support-document coverage is not required once the decisive evidence slots are realized. |
| 417 (Owlman) | AtlasNav realizes the complete annotated evidence; DR-DCI realizes only part of it. | AtlasNav wrong; DR-DCI correct. | Complete annotated evidence does not guarantee correct synthesis, and partial annotated realization does not preclude a correct answer. |
| Question category | Overall | Source tag | Overall |
|---|---|---|---|
| Miscellaneous | 96.08 | HubSpot | 92.42 |
| Intra-document | 96.46 | Gmail | 86.74 |
| Constrained | 94.57 | Jira | 81.21 |
| Basic | 95.16 | Google Drive | 88.70 |
| Conflicting information | 92.69 | Fireflies | 87.22 |
| Information not found | 100.00 | Linear | 84.23 |
| Paradigm | Representative methods | Persistent structure | Query-specific structure | Continued interaction | Primary role of structure |
|---|---|---|---|---|---|
| Sparse/dense retrieval | BM25; dense retrievers | Index only | No | Limited | Select top- evidence |
| Retrieval + reranking | ReasonRank ( Liu et al., 2026b ) ; neural rerankers | Index | No | Limited | Refine evidence ranking |
| Hierarchical/ graph RAG | RAPTOR; HippoRAG; SiReRAG; LinearRAG ( Zhuang et al., 2026a ) | Yes | Usually no | Retrieval-mediated | Structured evidence selection |
| Trained retrieval agents | Search-R1; R1-Searcher ( Song et al., 2025 ) ; ASearcher ( Gao et al., 2025 ) | Retrieval index | Adaptive querying | Via search API | Improve search policy |
| Relevance-guided interaction | RARG ( Li et al., 2026a ) | Retrieval index | Query-conditioned priority | Yes | Guide traversal and local match visibility |
| Raw direct interaction | DCI | No | No | Yes | Direct corpus operations |
| Interface | Strict Acc. | Cost | Turns |
|---|---|---|---|
| DCI | 95.50 | 22.59 | 4,117 |
| DR-DCI | 94.50 | 28.01 | 5,540 |
| AtlasNav | 94.50 | 10.89 | 3,135 |
| Interface | Loose | Strict | Cost | Turns |
|---|---|---|---|---|
| DCI | 79.65 | 39.35 | 73.35 | 6,596 |
| DR-DCI | 81.42 | 44.19 | 43.24 | 6,493 |
| AtlasNav | 80.10 | 40.32 | 36.49 | 4,218 |
| Symbol | Meaning | Symbol | Meaning |
|---|---|---|---|
| Question or query. | Set of evaluation questions. | ||
| Required evidence slots for ; is their number. | Construction, Surface, Open, or Locate checkpoint. | ||
| Evidence slots realized by checkpoint . | Fraction of required slots realized at . | ||
| Evidence Blindness metrics for total failure, mean loss, and incomplete realization. | Original accessible corpus. | ||
| Language-model agent. | Corpus interface. | ||
| Interaction budget. | Strict task performance of under budget . |