ICMAPE: In-Context Multiagent Pure Exploration
Organizations: Boston University · Broad Institute of MIT and Harvard
Abstract
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-specified models, while there is currently a gap for practical multi-agent methods that can perform active sequential testing. We fill this gap with ICMAPE, a Bayesian learning-based framework for decentralized multi-agent pure-exploration driven by inference objectives. ICMAPE converts the fixed-confidence identification objective into a reward derived from inference confidence, so that standard reinforcement learning machinery can be applied to decentralized pure exploration. It jointly learns a centralized neural inference network that estimates a posterior distribution over hypotheses from global trajectory data, and decentralized policies that select actions from local observation histories and learn when to stop collecting data once the target confidence is reached. On two synthetic benchmarks and a Maryland nitrate concentration monitoring task based on real-world data, ICMAPE-TD3 achieves target accuracy with fewer exploration steps.
Figures & tables
| Variant | Inf. Reward | Learned Stop | Step Cost | Accuracy | Avg. Stop Time |
|---|---|---|---|---|---|
| Full ICMAPE-TD3 | ✓ | ✓ | ✓ | 0.91 | 13.5 |
| Environment-reward TD3 | ✗ | ✓ | ✓ | 0.21 | 47.0 |
| Fixed-threshold stopping | ✓ | ✗ | ✓ | 0.70 | 37.3 |
| No sampling cost | ✓ | ✓ | ✗ | 0.90 | 50.0 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
|---|---|
| 10 | |
| Year | Raw Data Points | Avg. Var. | Avg. Sigma |
|---|---|---|---|
| 2000 | 10,693 | 0.265 | 0.515 |
| 2001 | 10,146 | 0.287 | 0.536 |
| 2002 | 11,063 | 0.249 | 0.499 |
| 2003 | 11,542 | 0.264 | 0.514 |
| 2004 | 11,769 | 0.421 | 0.649 |
| 2005 | 10,326 | 0.238 | 0.488 |
| Symbol / Parameter | Value | Description |
|---|---|---|
| grid | Discretized Maryland sensing domain | |
| Grid resolution along each spatial dimension | ||
| valid grid cells | Cells inside the region of interest with historical monitoring support | |
| Number of valid sensing locations | ||
| Number of agents | ||
| Maximum episode horizon |
| Confidence Bin | Mean Confidence | Empirical Accuracy | Samples ( ) | |
|---|---|---|---|---|
| 0.03 | 0.03 | 0.00 | 69,777 | |
| 0.14 | 0.13 | 0.01 | 13,284 | |
| 0.23 | 0.22 | 0.01 | 8,793 | |
| 0.34 | 0.33 | 0.01 | 6,280 | |
| 0.46 | 0.44 | 0.02 | 8,143 | |
| 0.53 | 0.53 | 0.00 | 1,127 |
| Joint action | Count | % | ||
|---|---|---|---|---|
| 0 | 10,000 | 10,000 | 100.0 | |
| 1 | 10,000 | 10,000 | 100.0 | |
| 2 | 10,000 | 10,000 | 100.0 | |
| 3 | 10,000 | 10,000 | 100.0 | |
| 4 | 10,000 | 10,000 | 100.0 | |
| 5 | 10,000 | 10,000 | 100.0 |
| Combination | Count | % vol. stops | % all envs |
|---|---|---|---|
| 1 player stopping simultaneously | |||
| 1,778 | 17.8 | 17.8 | |
| 1,056 | 10.6 | 10.6 | |
| 947 | 9.5 | 9.5 | |
| 10 | 0.1 | 0.1 | |
| 2 players stopping simultaneously | |||
| Agent 1 | Agent 2 | Agent 3 | ||
|---|---|---|---|---|
| 0 | dn (100.0%) | dn (100.0%) | rt ( 0 99.9%) | 10,000 |
| 1 | dn (100.0%) | dn (100.0%) | rt ( 0 94.6%) | 9,994 |
| 2 | rt (100.0%) | dn (100.0%) | rt ( 0 92.6%) | 9,459 |
| 3 | up (100.0%) | dn (100.0%) | up ( 0 95.2%) | 8,759 |
| 4 | up (100.0%) | dn (100.0%) | up ( 0 93.0%) | 8,338 |
| 5 | up (100.0%) | dn (100.0%) | up ( 0 87.5%) | 7,755 |
| Combination | Count | % vol. stops | % all envs |
|---|---|---|---|
| 1 robot stopping simultaneously | |||
| 5,095 | 50.9 | 50.9 | |
| 4,165 | 41.6 | 41.6 | |
| 2 robots stopping simultaneously | |||
| 740 | 7.4 | 7.4 | |