When Should Agents Check External State? Budgeting Observations for Stored Intentions
Authors: Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu, Bo Dong
Organizations: School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China · School of Distance Education, Xi’an Jiaotong University, Xi’an, China
Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems decide when intentions require attention, but do not allocate the resulting observations under a shared budget. We introduce the first resource-allocation formulation for the external observations required by stored intentions under a shared episode budget. BudgetPM offers two policy variants that share a hard-budget executor. BudgetPM-Static uses a lightweight Logistic scorer to learn whether a check improves the current decision. BudgetPM-Sequential distills full-episode hindsight schedules into a lightweight policy that decides when to spend or reserve capacity using only pre-query information at deployment. We evaluate BudgetPM against two public memory-agent systems, five matched controls, and four hand-designed monitoring or budget-adaptation rules. Across two benchmarks and three backbones, BudgetPM-Static outperforms adapted Mem0 and PMA workflows. On PM-Bench, its Logistic scorer reaches competitive quality--cost operating points alongside higher-capacity scorers and retains 99.9--100% of unconstrained quality with 42--54% fewer observations. Under severe scarcity and the same hard caps, BudgetPM-Sequential exceeds the strongest tested natural monitoring schedule by 1.92--2.58 Set F1 points. It reaches the same Set F1 and on-time recall with 16--33% fewer observations. Matched attribution, exact-cost analysis, and a fixed-budget load intervention link this gain to competition between present and future opportunities. These results yield a demand--capacity design rule: local gating works when capacity covers demand, while future-aware supervision adds value when observations compete across time.
Figures & tables
Figure 1: Overview of BudgetPM. Frozen development trajectories provide cached observation results for constructing (a) local Outcome labels and (b) hindsight-DP schedule labels under an episode budget. At deployment (c), Static or Sequential uses only pre-query inputs; thresholded proposals pass through a shared hard-cap executor that enforces the per-step and episode limits. Future information is used only for offline supervision.
Backbone
System
PM-Bench Set F1
GoodAI Macro Acc.
Llama-3-8B
Mem0
0.201
0.000
PMA
0.398
0.026
BudgetPM (Static)
0.811
0.946
Qwen2.5-14B
Mem0
0.484
0.450
PMA
0.490
0.513
BudgetPM (Static)
0.852
0.946
Table 1: Complete-workflow task quality averaged over seeds 4000–4029. BudgetPM uses Static at B60; Mem0 and PMA retain their adapted native workflows. Controlled allocation tests follow; full protocol and GoodAI component scores appear in Appendix B.1 .
B20
B40
B60
B80
Selection / scorer
F1
Cost
F1
Cost
F1
Cost
F1
Cost
Relevance heuristic
0.592
12.33
0.592
12.47
0.785
114.30
0.806
138.37
Matched Due + Logistic
0.773
34.37
0.808
60.97
0.816
89.27
0.819
115.27
Outcome + LambdaMART
0.782
35.23
0.812
64.90
0.821
88.83
0.820
109.73
Outcome + Ensemble
0.761
31.67
0.809
66.93
0.819
88.73
0.820
115.07
Outcome + Logistic (Static)
0.771
31.17
0.814
61.37
0.819
89.47
0.820
113.23
Table 2: Matched PM-Bench local-selection controls on 30 paired weeks (seeds 6000–6029). Row names identify the selection signal and scorer. Each budget reports mean Set F1 ( ↑ ) and realized query cost ( ↓ ); unconstrained All relevant achieves 0.820 F1 at cost 195.03.
Test
Budget
Deadline
Deadline+Backoff
Periodic
Severe
B5
1.50 ×
1.50 ×
4.70 ×
B10
1.20 ×
1.40 ×
4.08 ×
Boundary
B5
1.50 ×
1.50 ×
4.19 ×
B10
1.19 ×
1.39 ×
3.49 ×
Table 3: Observation cost required by each natural schedule to match Sequential’s Set F1 and on-time recall, relative to Sequential ( 1.00× ; lower is better).
Figure 2: Future-aware gains follow the demand–capacity boundary. a: Sequential-minus-Myopic Set F1 gains on two disjoint tests and exact-cost structural headroom are positive at B5/B10, where observations compete, and zero at B15/B20, where capacity covers demand. b: At fixed B10, the controlled-load test shows that increasing opportunity load creates 0.00652 exact-cost headroom while Sequential remains above Myopic. Error bars show paired-bootstrap 95% CIs.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Role
Seed range
Size
Static and standard-budget development
2000–2019
20 weeks
System comparison and Static analysis
4000–4029
30 weeks
Static confirmation
6000–6029
30 weeks
Severe-scarcity test
8000–8029
30 weeks
Boundary confirmation
9000–9029
30 weeks
Controlled-load development
20000–20049
50 pairs
Appendix
Table 4: PM-Bench split roles. A pair contains one k=0 and one k=3 week.
Benchmark
Labels
B20
B40
B60
B80
PM-Bench
Outcome
0.51608
0.33971
0.12289
0.00283
Due
0.62312
0.32715
0.15360
0.03700
GoodAI
Outcome
0.78747
0.28710
0.08381
0.02749
Due
0.80578
0.29739
0.08343
0.02729
Appendix
Table 5: Frozen Static thresholds selected on development data.
Backbone
System
PM-Bench Set F1
GoodAI Macro Acc. (Pros. / Trigger)
Llama-3-8B
Mem0
0.201
0.000 (0.000/0.000)
PMA
0.398
0.026 (0.013/0.038)
Qwen2.5-14B
Mem0
0.484
0.450 (0.013/0.887)
PMA
0.490
0.513 (0.027/1.000)
Qwen3-14B
Mem0
0.433
0.500 (0.000/1.000)
PMA
0.432
0.500 (0.013/0.987)
Appendix
Table 6: Complete Mem0 and PMA quality matrix on seeds 4000–4029. The GoodAI column reports macro accuracy with prospective-memory / trigger-response component accuracy in parentheses.
PM-Bench
GoodAI
Model
B40
B60
B40
B60
Llama-3-8B
0.802 / 62.97
0.811 / 90.90
0.927 / 20.17
0.946 / 30.13
Qwen2.5-14B
0.821 / 76.63
0.852 / 105.10
0.927 / 20.17
0.946 / 30.13
Qwen3-14B
–
0.885 / 121.87
–
0.944 / 30.13
Appendix
Table 7: Zero-adaptation transfer. Each cell is quality / mean query cost over 30 held-out seeds 4000–4029. Qwen3 reports B60, the complete-system operating point.
Method
40
60
90
100
Relevance heuristic
0.6643
0.7168
0.7722
0.7774
Task-specific Due/Logistic
0.7690
0.7877
0.8070
0.8103
Matched Due/Logistic
0.7818
0.8067
0.8165
0.8191
Outcome/LambdaMART
0.7897
0.8090
0.8211
0.8205
Outcome/Ensemble
0.7761
0.8018
0.8194
0.8201
Outcome/Logistic (Static)
0.7842
0.8127
0.8190
0.8200
Appendix
Table 8: Interpolated F1 at fixed mean realized-query costs on held-out PM-Bench weeks 6000–6029. The 13 operating points are fixed from development data; values interpolate adjacent points on each frozen policy curve.
Method
B60
B80
Task-specific Due gate
0.932 / 30.37
0.934 / 40.07
Static
0.946 / 30.13
0.951 / 39.63
Static gain (pp)
+1.33
+1.67
Appendix
Table 9: GoodAI gate comparison on 30 paired analysis trajectories. The first two rows give macro accuracy / mean realized semantic-verifier queries; the last gives the accuracy gain computed before rounding.
Benchmark
Method
B20
B40
B60
B80
PM-Bench
Random feasible
0.613 / 38.58
0.655 / 74.70
0.696 / 105.30
0.735 / 128.16
All relevant
0.820 / 195.03 (unconstrained)
GoodAI
Random feasible
0.579 / 9.37
0.664 / 20.04
0.756 / 30.40
0.836 / 40.03
All semantic
0.993 / 57.67 (unconstrained)
Appendix
Table 10: Random feasible selection and unconstrained query references. Random entries average three repetitions per seed (PM-Bench: confirmation seeds 6000–6029; GoodAI: seeds 4000–4029) and are quality / mean realized cost. The unconstrained row has no target budget.
Benchmark
Supervision
B20
B40
B60
B80
PM-Bench
Due state
0.762 / 36.33
0.798 / 63.53
0.808 / 92.90
0.810 / 114.57
Outcome
0.760 / 30.57
0.802 / 62.97
0.811 / 90.90
0.811 / 115.60
GoodAI
Due state
0.787 / 9.83
0.924 / 20.17
0.941 / 30.20
0.948 / 39.73
Outcome
0.789 / 9.87
0.927 / 20.17
0.946 / 30.13
0.951 / 39.63
Appendix
Table 11: Matched Outcome-versus-Due control. Every cell uses identical data volume, features, logistic capacity, calibration, executor, and hard cap. Entries are quality / mean realized cost on analysis seeds 4000–4029.
Benchmark
Budget
Static
Pacing
PM-Bench
B20
0.760 / 30.57
0.770 / 40.87
B40
0.802 / 62.97
0.804 / 79.47
B60
0.811 / 90.90
0.811 / 112.00
B80
0.811 / 115.60
0.811 / 130.40
GoodAI
B20
0.789 / 9.87
0.790 / 9.93
B40
0.927 / 20.17
0.928 / 20.63
Appendix
Table 12: Static versus remaining-budget-paced Outcome gating. Each entry is quality / mean realized cost on analysis seeds 4000–4029.
Benchmark
Method
B30
B50
B70
PM-Bench
Relevance
0.580 / 14.70
0.753 / 81.43
0.793 / 133.03
Task-specific Due gate
0.777 / 61.50
0.803 / 96.10
0.807 / 113.67
Budget-tuned ens.
0.767 / 43.40
0.784 / 71.47
0.806 / 96.27
Static
0.789 / 50.10
0.807 / 71.60
0.811 / 96.73
GoodAI
Similarity
0.637 / 12.23
0.677 / 24.97
0.672 / 35.73
Task-specific Due gate
0.906 / 15.40
0.959 / 23.83
0.941 / 35.10
Appendix
Table 13: Unseen-budget interpolation on analysis seeds 4000–4029. Each cell is mean quality / realized cost.
Figure 3: Held-out quality versus Outcome-label data. Error bars are paired bootstrap 95% CIs over analysis seeds 4000–4029. GoodAI reaches an early performance plateau.
Budget
Reconstructed Static cost
Total oracle gap (DP–Static) [95% CI]
Exhaustion rate
Later positives
B5
10.00
+0.0457 [0.0406, 0.0508]
1.00
17.60
B10
21.00
+0.0794 [0.0718, 0.0873]
1.00
7.93
B20
30.00
+0.0741 [0.0680, 0.0804]
0.00
0.00
B40
56.93
+0.0342 [0.0296, 0.0391]
0.00
0.00
B60
89.93
+0.0253 [0.0208, 0.0299]
0.00
0.00
B80
115.50
+0.0245 [0.0202, 0.0289]
0.00
0.00
Appendix
Table 14: Exact-cost hindsight analysis on seeds 4000–4029, using 30 dedicated frozen full-subset collection trajectories. “Reconstructed Static cost” is evaluated on the collection trajectory; “Later positives” counts positive singleton opportunities strictly after the reconstructed Static rule first exhausts the episode cap, averaged over weeks. The total oracle gap can reflect both current-value ranking and temporal allocation; it is not a direct measure of temporal headroom.
Budget
Δ F1 [95% CI]
Δ cost [95% CI]
Wins
Ties
B5
+0.0183 [0.0132, 0.0233]
0 [0, 0]
25
2
B10
+0.0085 [0.0051, 0.0120]
−0.23 [ −0.47 , −0.07 ]
19
9
Appendix
Table 15: Primary sequential-minus-matched-myopic contrasts on weeks 8000–8029. Intervals are paired complete-week bootstrap intervals with 100,000 resamples.
Policy
B5
B10
Static
0.6225 / 10.00
0.6944 / 21.00
Pacing
0.6225 / 10.00
0.6944 / 21.00
Pacing-LB
0.6435 / 10.00
0.7159 / 20.93
Periodic-Outcome
0.5620 / 10.00
0.5874 / 20.73
Myopic
0.6367 / 10.00
0.7159 / 20.90
Sequential
0.6550 / 10.00
0.7243 / 20.67
Appendix
Table 16: Absolute policy performance on severe-scarcity weeks 8000–8029. Entries are mean episode Set F1 / realized cost over 30 weeks. Pacing-LB is calibrated at B5/B10; Pacing retains the original standard-budget thresholds.
Budget
Δ F1 [95% CI]
Δ cost [95% CI]
Wins
Ties
B5
+0.0137 [0.0072, 0.0201]
0 [0, 0]
21
5
B10
+0.0105 [0.0067, 0.0146]
−0.07 [ −0.17 , 0]
22
6
B15
0 [0, 0]
0 [0, 0]
0
30
B20
0 [0, 0]
0 [0, 0]
0
30
Appendix
Table 17: Primary boundary-confirmation contrasts on weeks 9000–9029. Intervals are paired complete-week bootstrap intervals with 100,000 resamples.
Policy
B5
B10
B15
B20
Static
0.6236 / 10.00
0.6970 / 21.00
0.7459 / 29.57
0.7574 / 32.13
Pacing
0.6236 / 10.00
0.6970 / 21.00
0.7472 / 30.70
0.7669 / 41.23
Pacing-LB
0.6441 / 10.00
0.7151 / 21.00
–
–
Periodic-Outcome
0.5662 / 10.00
0.5975 / 20.63
0.6256 / 30.00
0.6461 / 36.37
Myopic
0.6340 / 10.00
0.7144 / 21.00
0.7486 / 29.23
0.7583 / 39.80
Sequential
0.6477 / 10.00
0.7250 / 20.93
0.7486 / 29.23
0.7583 / 39.80
Appendix
Table 18: Absolute policy performance on boundary weeks 9000–9029. Entries are mean episode Set F1 / realized cost over 30 weeks. Pacing-LB is calibrated at B5/B10; Pacing retains the original standard-budget thresholds.
Control
Test
B5
B10
Pacing-LB
Severe
+0.0115 [0.0037, 0.0197]
+0.0084 [0.0047, 0.0121]
Boundary
+0.0036 [ −0.0023 , 0.0096]
+0.0099 [0.0072, 0.0128]
Periodic
Severe
+0.0930 [0.0851, 0.1009]
+0.1370 [0.1260, 0.1481]
Boundary
+0.0816 [0.0753, 0.0879]
+0.1275 [0.1196, 0.1353]
Deadline
Severe
+0.0258 [0.0204, 0.0314]
+0.0194 [0.0140, 0.0248]
Boundary
+0.0192 [0.0123, 0.0259]
+0.0192 [0.0137, 0.0249]
Appendix
Table 19: Sequential minus practical controls at matched B5/B10 caps. Entries are Set F1 differences with paired 95% CIs.
Test
Target
Sequential
Deadline
Deadline+Backoff
Periodic
Severe
B5
10.00
15.00
15.00
47.03
B10
20.67
24.80
29.00
84.40
Boundary
B5
10.00
15.00
15.00
41.93
B10
20.93
24.93
29.10
73.13
Appendix
Table 20: Minimum observed cost that matches both Sequential’s mean Set F1 and on-time recall. Values summarize the 20-cap sweep.
Budget
Full DP−CG
Across-step allocation
Within-step choice
B5
+0.00737
+0.00737 [0.00433, 0.01071]
0
B10
+0.00367
+0.00367 [0.00230, 0.00512]
0
B15
0
0 [0, 0]
0
B20
0
0 [0, 0]
0
Appendix
Table 21: Exact-cost headroom decomposition on 30 frozen PM-Bench boundary weeks. Brackets give paired-bootstrap 95% confidence intervals for the across-step component. Within-step choice is zero in every week.
Condition
Learned Δ F1 (policy costs)
Structural Δ F1 (exact cost)
k=0
+0.00407 [0.00224, 0.00599]
0 [0, 0]
k=3
+0.00836 [0.00353, 0.01320]
+0.00652 [0.00580, 0.00726]
IL=G3−G0
+0.00429 [ −0.00064 , +0.00922]
–
Appendix
Table 22: Paired controlled-load evaluation at B10 on 240 held-out base seeds. Learned entries are Sequential minus Myopic. Structural headroom is exact-cost DP minus the future-blind current-greedy oracle. Brackets give paired bootstrap 95% confidence intervals.
Policy or structural reference
k=0
k=3
Myopic
0.7131 / 18.95
0.5670 / 19.96
Sequential
0.7171 / 19.90
0.5754 / 20.86
Current-greedy
0.7384 / 14.58
0.6763 / 21.00
Hindsight DP (exact cost)
0.7384 / 14.58
0.6828 / 21.00
Appendix
Table 23: Absolute performance at fixed B10 on controlled-load test seeds 21000–21239. Entries are mean episode Set F1 / realized cost. Learned policies are evaluated with their deployed histories; the two structural references share frozen full-subset trajectories and are matched at exact cost.
Retained set
n
B5 Δ F1 [95% CI]
B10 Δ F1 [95% CI]
Full set
30
+0.0137 [0.0072, 0.0201]
+0.0105 [0.0067, 0.0146]
No current-positive exception
16
+0.0171 [0.0088, 0.0259]
+0.0078 [0.0023, 0.0143]
No future-linked exception
15
+0.0142 [0.0059, 0.0232]
+0.0096 [0.0033, 0.0165]
No mutation-bearing exception
12
+0.0140 [0.0044, 0.0245]
+0.0080 [0.0006, 0.0163]
No exception of any kind
11
+0.0148 [0.0044, 0.0261]
+0.0071 [ −0.0005 , 0.0162]
Appendix
Table 24: Boundary robustness to compiler-event exclusions. Each row excludes whole paired weeks; intervals use 100,000 complete-week bootstrap resamples. “Future-linked” means an inferred affected benchmark object has a later positive instance.
Cost profile
Target / scorer
B40
B60
Unit
Binary Outcome / Logistic
0.802 / 62.83
0.811 / 90.90
Magnitude-weighted Outcome / Logistic
0.802 / 62.77
0.811 / 91.23
Hurdle expected gain
0.800 / 62.97
0.809 / 93.27
Direct effect / Ridge
0.776 / 66.70
0.797 / 94.63
Pairwise gain / cost
0.804 / 68.83
0.809 / 95.77
Mild
Relevance / cost
0.765 / 83.48
0.794 / 125.58
Appendix
Table 25: Marginal-target and heterogeneous-cost comparisons. Each cell is PM-Bench F1 / mean realized cost on analysis seeds 4000–4029.
While agents are increasingly spending more resources, today agent cost is mostly measured only after execution. A Budget-Aware Agent (BAGEN) should treat budget as an active control signal, rather than a passive cost metric. We first systematically define budget estimation as internal budgets (from agent computation) and external budgets (from agent actions). We then formalize budget-awareness as progressive interval estimation: at each step of a plan, an agent should predict an upper and lower bound on remaining budget, and alert when completion is unlikely. Scoring with a rollout-replay protocol, we find consistent failure patterns on four environments and five frontier agents: (1) strong agents do not necessarily have strong budget-awareness, with correlation r=0.35. (2) frontier models are consistently over-optimistic, continue spending on tasks that are unlikely to succeed, instead of alerting the user early. (3) budget-aware signal is actionable and trainable. Early stop saves 28-64% tokens on failed trajectories, and SFT+RL strengthens early stop and alert behavior. (4) precise interval calibration remains challenging, with interval coverage capping at 47% after SFT+RL. Project page: https://ragen-ai.github.io/bagen/
Yuxiang Lin, Zihan Wang, Mengyang Liu +9
1Northwestern University · 2O2 Lab · 3Independent +5
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents. Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due. We compare eight state-of-the-art LLMs on PM-Bench under eight different agent configurations. PM-Bench proves challenging across all settings: the best method, a GPT-5.4 agent, reaches only 65.1% F1 score under our evaluation. Furthermore, no single strategy for improving prospective memory dominates across models. We release PM-Bench as a controlled testbed for diagnosing these failures and developing training or inference-time interventions that support reliable prospective behavior.
Prospective memory means carrying out a deferred intention at the right future cue while other work continues. Benchmarks now isolate it as an agent skill, yet frontier LLMs still struggle: the best published PM-Bench scaffold reaches only 65.1% Set-F1. We argue that this loop is schema-constrained state tracking rather than open-ended reasoning, and that small models can execute it when the action space is typed. We propose the Prospective Intention Store (PIS) that puts lifecycle logic in code and scoped language work on the model. The scaffold is agentic and training-free: no selector fine-tuning and no trajectory distillation. On PM-Bench, DeepSeek-Chat with PIS reaches 82.9% Set-F1. On Gemma-E2B, Set-F1 is only 4.2% without a store and at most 6.6% under seven retrospective memories, while PIS reaches 66.2%. PIS further reaches 70.1% Set-F1, where retrospective memory methods stay at most 54.4%. PIS sets a new state of the art on this benchmark and enables small models to surpass the published large-model scaffold.