Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
Organizations: KAIST · University of Minnesota
Abstract
Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.
Figures & tables
| Situation. The user takes a language class after work and connects the class app and calendar to the agent. It can book a 15-minute review session and set a reminder 10 minutes beforehand. | |
| Standing request (Day 1). “Always get my confirmation before you add any review session or reminder.” The user repeats this requirement on Day 3. | |
| Day 2: approval before action | Day 4: action without approval |
| App notification. A unit covers café ordering phrases; tomorrow’s 20:00–20:15 slot is free. | App notification. A unit covers asking directions; tomorrow’s 19:30–19:45 slot is free. |
| Agent. “I can book a slot tomorrow from 20:00 to 20:15 to review the café ordering phrases, with a reminder at 19:50 , 10 minutes before it starts. Shall I add it? I haven’t changed your calendar or any reminder yet.” | Agent. “ I’ve added a slot tomorrow from 19:30 to 19:45 to review the phrases for asking directions, with a reminder at 19:20 . I haven’t changed any other events.” |
| User. “I’ve checked it. Go ahead with this one. ” Agent. “I got your confirmation and applied exactly what I showed you.” | Action. The agent adds both entries without asking first , then reports its action. There are no scheduling conflicts. |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Work | System / application | TC | TA | TR | Basis for assignment |
| Proactive systems | |||||
| ContextAgent ( Yang et al., 2025b ) | LLM agent | Anticipates needs and selects tools; no TA , with only persona-aware thresholding for TR . | |||
| ProAct ( Hu et al., 2026 ) | LLM agent | Selects useful preparation within idle-time and compute budgets, but no compute allocation between idle and active time; no TR objective. | |||
| Benchmarks and simulation environments for LLM/VLM assistance | |||||
| ProactiveBench ( Lu et al., 2025 ) | Desktop assistant benchmark | Scores proposed tasks and triggers, not completed assistance; no TA or TR measure. | |||
| ProAgentBench ( Tang et al., 2026b ) | LLM/VLM assistant benchmark | Scores timing and query prediction; no deliverable-quality, TA , or TR evaluation. | |||
| Time | Trigger | What to infer or verify | What to do |
|---|---|---|---|
| Mon. 18:00 | User message “Is that M27 really £150?” | • Whether the monitor fits Jordan’s desk and connects to the laptop. • Jordan needs an HDMI cable to connect the monitor to the laptop. The preferred new cable costs £10, bringing the total to £160. • Stock can be checked after Tuesday 09:00. | • Answer the price question and explain the complete setup. • Prepare the purchase plan and save it if permitted. • Schedule a stock check if the purchase remains wanted. |
| Tue. 09:00–09:30 | Agent-scheduled follow-up Check stock | • Whether stock is available and an order already exists. • A 10-minute purchase review and a 20-minute warranty discussion cannot both fit Jordan’s 20-minute phone session. • Orders close at noon; warranty coverage can be added within 30 days, so that decision can wait. | • Now: review the purchase and place the order if authorized. • Later: defer the warranty comparison to Sunday’s available session and arrange to return to it. |
| Fri. 19:10 | File notification Jordan added delivery photos | • The photos show a DisplayPort cable, although the invoice specifies HDMI. • The delivered cable cannot connect to the laptop, and a replacement would arrive after Saturday’s class. • Jordan’s working HDMI spare can keep the monitor usable for class. | • Prepare a free replacement request with the invoice and photos; ask before submitting it. • Explain how to connect the existing spare for Saturday’s class. • Keep the working monitor rather than returning the whole order. |
| Sun. 12:00 | Product notification The saved listing has changed | • The identical in-stock monitor now costs £135; the cable remains £10. • Whether a paid order qualifies under the seven-day price-adjustment policy. • Tuesday’s £150 monitor purchase would support a £15 claim; without a purchase, no refund is owed. | • If Tuesday’s order was paid: prepare the £15 claim and ask before submitting it. • If no purchase was made: present the new £145 total as a purchase option. |
| Context | Comment |
|---|---|
| P27: Task Capability. Preferred Agent B over Agent A, where Agent A has limited TC . | “B also points out information that could easily be overlooked, making it more helpful.” |
| P12: Temporal Allocation. Preferred sleep-time assistance despite possible revisions. | “Even if it needs correction tomorrow, I should focus on what matters now and delegate as much as possible to AI.” |
| P04: TA – TC tradeoff. Accepted imperfect sleep-time outputs when correction still allowed an overall time saving. | “I would revise it if the time spent prompting the AI and fixing its answer, excluding time waiting for the AI, were sufficiently shorter than making the material myself from scratch.” |
| P09: Temporal Allocation. Preferred interaction-time assistance when both options met the deadline. | “If the deadline for the most important task can be met, I prefer the AI to work when I can check it myself. Working during sleep is more efficient, but I cannot correct it midway if it takes the wrong direction.” |
| P21: Over-intervention. The agent chose execute despite the user’s preference for suggest ; task outcomes were correct. | “I do not think there was any major harm in the end, but my trust declined because it handled things differently from what was requested.” |
| P04: Under-intervention. The agent repeatedly chose suggest despite the user’s preference for execute . | “I stopped trusting it after it asked the user again twice, despite being told not to seek approval.” |