ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces
Organizations: University of Illinois Urbana-Champaign
Abstract
Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what the query still requires. We introduce ANTMAN, an adaptive coordination framework that treats evolving unresolved information needs as the unit of runtime coordination. ANTMAN maintains a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress, and uses this state to control worker selection, routing, and task-local recovery as new evidence is discovered. By separating the coordination policy from substrate-specific search interfaces, the same need-conditioned mechanism can operate across different information spaces. Experiments across multi-document question answering, controlled long-context scaling, and realistic structured navigation show that ANTMAN remains effective across settings, including when execution is delegated to substantially smaller worker models. Under a 16x increase in searchable context, ANTMAN increases active coordination by only 1.23x, compared with more than 15x for partition-driven baselines, while preserving strong answer quality.
Figures & tables
| Method | HotpotQA | 2Wiki | MuSiQue | Avg. | ||||
| EM | F1 | EM | F1 | EM | F1 | EM | F1 | |
| Retrieval and RAG Baselines | ||||||||
| Iterative Sparse Retrieval (BM25) | .533 | .684 | .367 | .515 | .433 | .582 | .444 | .594 |
| Dense Retrieval (DPR) ( Karpukhin et al., 2020 ) EMNLP’20 | .533 | .696 | .267 | .379 | .200 | .292 | .333 | .456 |
| ChainRAG ( Zhu et al., 2025 ) ACL’25 | .567 | .711 | .733 | .791 | .500 | .654 | .600 | .719 |
| S2G-RAG ( Li et al., 2026 ) ACL’26 | .667 | .810 | .800 | .858 | .633 | .730 | .700 | .799 |
| Required Evidence Units | Active Coordination | Calls / Query | Cost / Query | Complete Evidence | Answer Acc. |
| ANTMAN | |||||
| 1 | 5.0 | 35.5 | $0.107 | 100% | 100% |
| 4 | 4.3 | 43.6 | $0.111 | 100% | 100% |
| 16 | 8.0 | 86.2 | $0.293 | 100% | 100% |
| Method | RepoProbe | SWE-QA-Pro | GAIA-Text-103 | |||
| L1 | L2 | L3 | Overall | |||
| Retrieval and Iterative Retrieval Baselines | ||||||
| Direct | 31.08 | 55.74 | 28.2 | 17.3 | 8.3 | 20.4 |
| Iterative Sparse Retrieval | 33.15 | 60.90 | 53.8 | 42.3 | 33.3 | 45.6 |
| Dense Retrieval ( Karpukhin et al., 2020 ) | 31.51 | 58.30 | 43.6 | 26.9 | 25.0 | 33.0 |
| S2G-RAG ( Li et al., 2026 ) ACL’26 | 22.87 | 60.40 | 25.6 | 19.2 | 0.0 | 19.4 |
| Variant | 512K F1 | SWE-QA-Pro Score | GAIA Accuracy |
| Adaptive Coordination Ablations | |||
| Static ANTMAN | 40.32 | 68.58 | 50.00 |
| Graph-free Adaptive | 64.01 | 78.88 | 50.00 |
| w/o Need Revision | 74.44 | 75.32 | 42.50 |
| w/o Adaptive Rerouting | 82.20 | 72.13 | 60.00 ( 0.00%) |
| w/o Recovery | 60.12 | 76.52 | 52.50 |
| Method | 32K | 64K | 128K | Middle Gap | ||||||
| E | M | L | E | M | L | E | M | L | ||
| Direct Long-Context Reference | ||||||||||
| Full Context | .911 | .590 | .813 | .767 | .667 | .733 | .733 | .613 | .870 | |
| Ours | ||||||||||
| ANTMAN | .741 | .783 | .800 | .747 | .826 | .780 | .743 | .748 | .813 | +.015 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Coord. Growth | Calls Growth | Cost Growth | F1 |
| Multi-Agent Baselines | ||||
| LongAgent ( Zhao et al., 2024 ) | 15.29 | 12.81 | 14.22 | |
| CoA ( Zhang et al., 2024 ) | 15.33 | 11.30 | 16.11 | |
| Ours | ||||
| ANTMAN | 1.23 | 1.21 | 6.90 | |
| Method | Length | Active Coord. | Calls / Query | Cost / Query | EM | F1 |
| Multi-Agent Baselines | ||||||
| LongAgent ( Zhao et al., 2024 ) | 32K | 19.00 | 22.83 | $0.0983 | .533 | .731 |
| 64K | 37.00 | 39.50 | $0.1797 | .533 | .737 | |
| 128K | 73.10 | 75.60 | $0.3549 | .500 | .670 | |
| 512K | 290.50 | 292.50 | $1.3974 | .567 | .754 | |
| CoA ( Zhang et al., 2024 ) | 32K | 5.10 | 7.10 | $0.0809 | .567 | .750 |
| Method | RepoProbe-Python / 10 | SWE-QA-Pro / 50 |
| Retrieval and Iterative Retrieval Baselines | ||
| Direct | 3.108 | 27.87 |
| Iterative Sparse Retrieval | 3.315 | 30.45 |
| Dense Retrieval ( Karpukhin et al., 2020 ) | 3.151 | 29.15 |
| S2G-RAG ( Li et al., 2026 ) | 2.287 | 30.20 |
| Agentic and Repository-Aware Baselines | ||
| Variant | Explicit Need State | Need Revision | Adaptive Rerouting | Recovery |
| Static ANTMAN | ✓ | |||
| Graph-free Adaptive | – | ✓ | ✓ | |
| w/o Need Revision | ✓ | ✓ | ✓ | |
| w/o Adaptive Rerouting | ✓ | ✓ | ✓ | |
| w/o Recovery | ✓ | ✓ | ✓ | |
| Full ANTMAN | ✓ | ✓ | ✓ | ✓ |
| Variant | Answer Quality | Active Coordination | Cost ($/query) | LLM Calls (/query) |
| Adaptive Coordination Ablations | ||||
| Static ANTMAN | 40.32 | 1.97 | 0.156 | 28.7 |
| Graph-free Adaptive | 64.01 | 3.20 | 0.354 | 20.7 |
| w/o Need Revision | 74.44 | 3.10 | 0.418 | 21.0 |
| w/o Adaptive Rerouting | 82.20 | 2.00 | 0.451 | 28.2 |
| w/o Recovery | 60.12 | 2.07 | 0.353 | 25.9 |
| Variant | Answer Quality | Active Coordination | Cost ($/query) | LLM Calls (/query) |
| Adaptive Coordination Ablations | ||||
| Static ANTMAN | 68.58 | 1.35 | 0.273 | 36.5 |
| Graph-free Adaptive | 78.88 | 2.45 | 0.165 | 22.1 |
| w/o Need Revision | 75.32 | 3.50 | 0.298 | 32.4 |
| w/o Adaptive Rerouting | 72.13 | 2.20 | 0.407 | 48.6 |
| w/o Recovery | 76.52 | 2.15 | 0.522 | 56.5 |
| Variant | Answer Quality | Cost ($/query) | LLM Calls (/query) |
| Adaptive Coordination Ablations | |||
| Static ANTMAN | 50.00 | 0.196 | 42.8 |
| Graph-free Adaptive | 50.00 | 0.179 | 54.0 |
| w/o Need Revision | 42.50 | 0.115 | 28.7 |
| w/o Adaptive Rerouting | 60.00 | 0.221 | 47.4 |
| w/o Recovery | 52.50 | 0.325 | 69.5 |
| Setting | Information access |
| Controlled long context | Lexical and semantic retrieval with bounded reading of selected document regions. |
| Repository navigation | Retrieval and bounded reading augmented with structural navigation over symbols, imports, references, inheritance, and call relations. |
| GAIA-Text-103 | Web search, webpage text extraction, and isolated Python execution for tool-assisted information seeking. |
| Required Evidence | Worker Activity | Relevant Activity | Irrelevant Activity |
| ANTMAN | |||
| 1 unit | 5.0 | 0.9 | 4.1 |
| 4 units | 4.3 | 3.2 | 1.1 |
| 16 units | 8.2 | 8.0 | 0.2 |
| Method | LLM Calls | Tool Calls | Cost / Query | RepoProbe Score |
| Retrieval and Iterative Retrieval Baselines | ||||
| Direct | 1.00 | 0.00 | $0.0067 | 31.08 |
| Iterative Sparse Retrieval | 3.19 | 2.16 | $0.0217 | 33.15 |
| Dense Retrieval ( Karpukhin et al., 2020 ) | 1.00 | 1.00 | $0.0124 | 31.51 |
| S2G-RAG ( Li et al., 2026 ) | 37.12 | 24.00 | $0.1788 | 22.87 |
| Agentic and Repository-Aware Baselines | ||||