When Should Agents Check External State? Budgeting Observations for Stored Intentions
Organizations: School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China · School of Distance Education, Xi’an Jiaotong University, Xi’an, China
Abstract
Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems decide when intentions require attention, but do not allocate the resulting observations under a shared budget. We introduce the first resource-allocation formulation for the external observations required by stored intentions under a shared episode budget. BudgetPM offers two policy variants that share a hard-budget executor. BudgetPM-Static uses a lightweight Logistic scorer to learn whether a check improves the current decision. BudgetPM-Sequential distills full-episode hindsight schedules into a lightweight policy that decides when to spend or reserve capacity using only pre-query information at deployment. We evaluate BudgetPM against two public memory-agent systems, five matched controls, and four hand-designed monitoring or budget-adaptation rules. Across two benchmarks and three backbones, BudgetPM-Static outperforms adapted Mem0 and PMA workflows. On PM-Bench, its Logistic scorer reaches competitive quality--cost operating points alongside higher-capacity scorers and retains 99.9--100% of unconstrained quality with 42--54% fewer observations. Under severe scarcity and the same hard caps, BudgetPM-Sequential exceeds the strongest tested natural monitoring schedule by 1.92--2.58 Set F1 points. It reaches the same Set F1 and on-time recall with 16--33% fewer observations. Matched attribution, exact-cost analysis, and a fixed-budget load intervention link this gain to competition between present and future opportunities. These results yield a demand--capacity design rule: local gating works when capacity covers demand, while future-aware supervision adds value when observations compete across time.
Figures & tables
| Backbone | System | PM-Bench Set F1 | GoodAI Macro Acc. |
|---|---|---|---|
| Llama-3-8B | Mem0 | 0.201 | 0.000 |
| PMA | 0.398 | 0.026 | |
| BudgetPM (Static) | 0.811 | 0.946 | |
| Qwen2.5-14B | Mem0 | 0.484 | 0.450 |
| PMA | 0.490 | 0.513 | |
| BudgetPM (Static) | 0.852 | 0.946 |
| B20 | B40 | B60 | B80 | |||||
|---|---|---|---|---|---|---|---|---|
| Selection / scorer | F1 | Cost | F1 | Cost | F1 | Cost | F1 | Cost |
| Relevance heuristic | 0.592 | 12.33 | 0.592 | 12.47 | 0.785 | 114.30 | 0.806 | 138.37 |
| Matched Due + Logistic | 0.773 | 34.37 | 0.808 | 60.97 | 0.816 | 89.27 | 0.819 | 115.27 |
| Outcome + LambdaMART | 0.782 | 35.23 | 0.812 | 64.90 | 0.821 | 88.83 | 0.820 | 109.73 |
| Outcome + Ensemble | 0.761 | 31.67 | 0.809 | 66.93 | 0.819 | 88.73 | 0.820 | 115.07 |
| Outcome + Logistic (Static) | 0.771 | 31.17 | 0.814 | 61.37 | 0.819 | 89.47 | 0.820 | 113.23 |
| Test | Budget | Deadline | Deadline+Backoff | Periodic |
|---|---|---|---|---|
| Severe | B5 | 1.50 | 1.50 | 4.70 |
| B10 | 1.20 | 1.40 | 4.08 | |
| Boundary | B5 | 1.50 | 1.50 | 4.19 |
| B10 | 1.19 | 1.39 | 3.49 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Role | Seed range | Size |
|---|---|---|
| Static and standard-budget development | 2000–2019 | 20 weeks |
| System comparison and Static analysis | 4000–4029 | 30 weeks |
| Static confirmation | 6000–6029 | 30 weeks |
| Severe-scarcity test | 8000–8029 | 30 weeks |
| Boundary confirmation | 9000–9029 | 30 weeks |
| Controlled-load development | 20000–20049 | 50 pairs |
| Benchmark | Labels | B20 | B40 | B60 | B80 |
|---|---|---|---|---|---|
| PM-Bench | Outcome | 0.51608 | 0.33971 | 0.12289 | 0.00283 |
| Due | 0.62312 | 0.32715 | 0.15360 | 0.03700 | |
| GoodAI | Outcome | 0.78747 | 0.28710 | 0.08381 | 0.02749 |
| Due | 0.80578 | 0.29739 | 0.08343 | 0.02729 |
| Backbone | System | PM-Bench Set F1 | GoodAI Macro Acc. (Pros. / Trigger) |
|---|---|---|---|
| Llama-3-8B | Mem0 | 0.201 | 0.000 (0.000/0.000) |
| PMA | 0.398 | 0.026 (0.013/0.038) | |
| Qwen2.5-14B | Mem0 | 0.484 | 0.450 (0.013/0.887) |
| PMA | 0.490 | 0.513 (0.027/1.000) | |
| Qwen3-14B | Mem0 | 0.433 | 0.500 (0.000/1.000) |
| PMA | 0.432 | 0.500 (0.013/0.987) |
| PM-Bench | GoodAI | |||
|---|---|---|---|---|
| Model | B40 | B60 | B40 | B60 |
| Llama-3-8B | 0.802 / 62.97 | 0.811 / 90.90 | 0.927 / 20.17 | 0.946 / 30.13 |
| Qwen2.5-14B | 0.821 / 76.63 | 0.852 / 105.10 | 0.927 / 20.17 | 0.946 / 30.13 |
| Qwen3-14B | – | 0.885 / 121.87 | – | 0.944 / 30.13 |
| Method | 40 | 60 | 90 | 100 |
|---|---|---|---|---|
| Relevance heuristic | 0.6643 | 0.7168 | 0.7722 | 0.7774 |
| Task-specific Due/Logistic | 0.7690 | 0.7877 | 0.8070 | 0.8103 |
| Matched Due/Logistic | 0.7818 | 0.8067 | 0.8165 | 0.8191 |
| Outcome/LambdaMART | 0.7897 | 0.8090 | 0.8211 | 0.8205 |
| Outcome/Ensemble | 0.7761 | 0.8018 | 0.8194 | 0.8201 |
| Outcome/Logistic (Static) | 0.7842 | 0.8127 | 0.8190 | 0.8200 |
| Method | B60 | B80 |
|---|---|---|
| Task-specific Due gate | 0.932 / 30.37 | 0.934 / 40.07 |
| Static | 0.946 / 30.13 | 0.951 / 39.63 |
| Static gain (pp) | +1.33 | +1.67 |
| Benchmark | Method | B20 | B40 | B60 | B80 |
|---|---|---|---|---|---|
| PM-Bench | Random feasible | 0.613 / 38.58 | 0.655 / 74.70 | 0.696 / 105.30 | 0.735 / 128.16 |
| All relevant | 0.820 / 195.03 (unconstrained) | ||||
| GoodAI | Random feasible | 0.579 / 9.37 | 0.664 / 20.04 | 0.756 / 30.40 | 0.836 / 40.03 |
| All semantic | 0.993 / 57.67 (unconstrained) | ||||
| Benchmark | Supervision | B20 | B40 | B60 | B80 |
|---|---|---|---|---|---|
| PM-Bench | Due state | 0.762 / 36.33 | 0.798 / 63.53 | 0.808 / 92.90 | 0.810 / 114.57 |
| Outcome | 0.760 / 30.57 | 0.802 / 62.97 | 0.811 / 90.90 | 0.811 / 115.60 | |
| GoodAI | Due state | 0.787 / 9.83 | 0.924 / 20.17 | 0.941 / 30.20 | 0.948 / 39.73 |
| Outcome | 0.789 / 9.87 | 0.927 / 20.17 | 0.946 / 30.13 | 0.951 / 39.63 |
| Benchmark | Budget | Static | Pacing |
|---|---|---|---|
| PM-Bench | B20 | 0.760 / 30.57 | 0.770 / 40.87 |
| B40 | 0.802 / 62.97 | 0.804 / 79.47 | |
| B60 | 0.811 / 90.90 | 0.811 / 112.00 | |
| B80 | 0.811 / 115.60 | 0.811 / 130.40 | |
| GoodAI | B20 | 0.789 / 9.87 | 0.790 / 9.93 |
| B40 | 0.927 / 20.17 | 0.928 / 20.63 |
| Benchmark | Method | B30 | B50 | B70 |
|---|---|---|---|---|
| PM-Bench | Relevance | 0.580 / 14.70 | 0.753 / 81.43 | 0.793 / 133.03 |
| Task-specific Due gate | 0.777 / 61.50 | 0.803 / 96.10 | 0.807 / 113.67 | |
| Budget-tuned ens. | 0.767 / 43.40 | 0.784 / 71.47 | 0.806 / 96.27 | |
| Static | 0.789 / 50.10 | 0.807 / 71.60 | 0.811 / 96.73 | |
| GoodAI | Similarity | 0.637 / 12.23 | 0.677 / 24.97 | 0.672 / 35.73 |
| Task-specific Due gate | 0.906 / 15.40 | 0.959 / 23.83 | 0.941 / 35.10 |
| Budget | Reconstructed Static cost | Total oracle gap (DP–Static) [95% CI] | Exhaustion rate | Later positives |
|---|---|---|---|---|
| B5 | 10.00 | +0.0457 [0.0406, 0.0508] | 1.00 | 17.60 |
| B10 | 21.00 | +0.0794 [0.0718, 0.0873] | 1.00 | 7.93 |
| B20 | 30.00 | +0.0741 [0.0680, 0.0804] | 0.00 | 0.00 |
| B40 | 56.93 | +0.0342 [0.0296, 0.0391] | 0.00 | 0.00 |
| B60 | 89.93 | +0.0253 [0.0208, 0.0299] | 0.00 | 0.00 |
| B80 | 115.50 | +0.0245 [0.0202, 0.0289] | 0.00 | 0.00 |
| Budget | F1 [95% CI] | cost [95% CI] | Wins | Ties |
|---|---|---|---|---|
| B5 | +0.0183 [0.0132, 0.0233] | 0 [0, 0] | 25 | 2 |
| B10 | +0.0085 [0.0051, 0.0120] | [ , ] | 19 | 9 |
| Policy | B5 | B10 |
|---|---|---|
| Static | 0.6225 / 10.00 | 0.6944 / 21.00 |
| Pacing | 0.6225 / 10.00 | 0.6944 / 21.00 |
| Pacing-LB | 0.6435 / 10.00 | 0.7159 / 20.93 |
| Periodic-Outcome | 0.5620 / 10.00 | 0.5874 / 20.73 |
| Myopic | 0.6367 / 10.00 | 0.7159 / 20.90 |
| Sequential | 0.6550 / 10.00 | 0.7243 / 20.67 |
| Budget | F1 [95% CI] | cost [95% CI] | Wins | Ties |
|---|---|---|---|---|
| B5 | +0.0137 [0.0072, 0.0201] | 0 [0, 0] | 21 | 5 |
| B10 | +0.0105 [0.0067, 0.0146] | [ , 0] | 22 | 6 |
| B15 | 0 [0, 0] | 0 [0, 0] | 0 | 30 |
| B20 | 0 [0, 0] | 0 [0, 0] | 0 | 30 |
| Policy | B5 | B10 | B15 | B20 |
|---|---|---|---|---|
| Static | 0.6236 / 10.00 | 0.6970 / 21.00 | 0.7459 / 29.57 | 0.7574 / 32.13 |
| Pacing | 0.6236 / 10.00 | 0.6970 / 21.00 | 0.7472 / 30.70 | 0.7669 / 41.23 |
| Pacing-LB | 0.6441 / 10.00 | 0.7151 / 21.00 | – | – |
| Periodic-Outcome | 0.5662 / 10.00 | 0.5975 / 20.63 | 0.6256 / 30.00 | 0.6461 / 36.37 |
| Myopic | 0.6340 / 10.00 | 0.7144 / 21.00 | 0.7486 / 29.23 | 0.7583 / 39.80 |
| Sequential | 0.6477 / 10.00 | 0.7250 / 20.93 | 0.7486 / 29.23 | 0.7583 / 39.80 |
| Control | Test | B5 | B10 |
|---|---|---|---|
| Pacing-LB | Severe | +0.0115 [0.0037, 0.0197] | +0.0084 [0.0047, 0.0121] |
| Boundary | +0.0036 [ , 0.0096] | +0.0099 [0.0072, 0.0128] | |
| Periodic | Severe | +0.0930 [0.0851, 0.1009] | +0.1370 [0.1260, 0.1481] |
| Boundary | +0.0816 [0.0753, 0.0879] | +0.1275 [0.1196, 0.1353] | |
| Deadline | Severe | +0.0258 [0.0204, 0.0314] | +0.0194 [0.0140, 0.0248] |
| Boundary | +0.0192 [0.0123, 0.0259] | +0.0192 [0.0137, 0.0249] |
| Test | Target | Sequential | Deadline | Deadline+Backoff | Periodic |
|---|---|---|---|---|---|
| Severe | B5 | 10.00 | 15.00 | 15.00 | 47.03 |
| B10 | 20.67 | 24.80 | 29.00 | 84.40 | |
| Boundary | B5 | 10.00 | 15.00 | 15.00 | 41.93 |
| B10 | 20.93 | 24.93 | 29.10 | 73.13 |
| Budget | Full | Across-step allocation | Within-step choice |
|---|---|---|---|
| B5 | +0.00737 | +0.00737 [0.00433, 0.01071] | 0 |
| B10 | +0.00367 | +0.00367 [0.00230, 0.00512] | 0 |
| B15 | 0 | 0 [0, 0] | 0 |
| B20 | 0 | 0 [0, 0] | 0 |
| Condition | Learned F1 (policy costs) | Structural F1 (exact cost) |
|---|---|---|
| +0.00407 [0.00224, 0.00599] | 0 [0, 0] | |
| +0.00836 [0.00353, 0.01320] | +0.00652 [0.00580, 0.00726] | |
| +0.00429 [ , +0.00922] | – |
| Policy or structural reference | ||
|---|---|---|
| Myopic | 0.7131 / 18.95 | 0.5670 / 19.96 |
| Sequential | 0.7171 / 19.90 | 0.5754 / 20.86 |
| Current-greedy | 0.7384 / 14.58 | 0.6763 / 21.00 |
| Hindsight DP (exact cost) | 0.7384 / 14.58 | 0.6828 / 21.00 |
| Retained set | B5 F1 [95% CI] | B10 F1 [95% CI] | |
|---|---|---|---|
| Full set | 30 | +0.0137 [0.0072, 0.0201] | +0.0105 [0.0067, 0.0146] |
| No current-positive exception | 16 | +0.0171 [0.0088, 0.0259] | +0.0078 [0.0023, 0.0143] |
| No future-linked exception | 15 | +0.0142 [0.0059, 0.0232] | +0.0096 [0.0033, 0.0165] |
| No mutation-bearing exception | 12 | +0.0140 [0.0044, 0.0245] | +0.0080 [0.0006, 0.0163] |
| No exception of any kind | 11 | +0.0148 [0.0044, 0.0261] | +0.0071 [ , 0.0162] |
| Cost profile | Target / scorer | B40 | B60 |
|---|---|---|---|
| Unit | Binary Outcome / Logistic | 0.802 / 62.83 | 0.811 / 90.90 |
| Magnitude-weighted Outcome / Logistic | 0.802 / 62.77 | 0.811 / 91.23 | |
| Hurdle expected gain | 0.800 / 62.97 | 0.809 / 93.27 | |
| Direct effect / Ridge | 0.776 / 66.70 | 0.797 / 94.63 | |
| Pairwise gain / cost | 0.804 / 68.83 | 0.809 / 95.77 | |
| Mild | Relevance / cost | 0.765 / 83.48 | 0.794 / 125.58 |