Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive execution, leaving unclear whether agents can formulate, execute, and revise an explicit plan while preserving user preferences throughout long-term interaction. To address this gap, we introduce \textbf{PDEU-Bench} (\textbf{P}ersonalized plan \textbf{D}efinition, plan \textbf{E}xecution, and plan \textbf{U}pdate \textbf{Bench}mark), a benchmark for evaluating the complete planning lifecycle of personalized tool-using agents. PDEU-Bench comprises 214 long-horizon interaction tasks spanning 12 everyday domains and 94 tools, with stage-specific assessments of preference adherence and plan quality. Extensive evaluations of 15 representative open-source and closed-source LLMs reveal a pronounced gap between local tool execution and dynamic planning: LLMs can often instantiate preferences in individual calls, yet struggle to construct coherent plan definition and plan update. We further evaluate mainstream personalization and memory-augmentation methods. Although these methods improve particular stages, none of the evaluated methods reliably propagates user preferences throughout the complete lifecycle, and their gains frequently fail to transfer to subsequent execution. Fine-grained error analysis further reveals that preference omissions and conflicts persist throughout the planning lifecycle, highlighting the need for future research to parameterize LLMs with preference-aware information retrieval and memory capabilities. We provide the relevant code and data in the appendix to support future research.
Figures & tables
Figure 1: Motivation for PDEU-Bench: a benchmark for personalized tool-use planning across the whole planning lifecycle. (a) Single-turn evaluation focuses on an isolated personalized single execution. (b) Multi-turn reactive evaluation extends interaction through repeated reasoning and plan execution. (c) PDEU-Bench treats the plan as an explicit, evolving object and evaluates its core lifecycle through plan definition, plan execution (Tool Call), and feedback-driven plan update.
Existing Benchmarks
Multi-turn Planning
Plan Action
Definition
Execution
Update
PTBench
✗
✗
✓
✗
ToolSpectrum
✗
✗
✓
✗
PEToolBench
✗
✗
✓
✗
Claw-Anything
✗
✗
✓
✗
ASTRA-Bench
✗
✗
✓
✗
Table 1: Comparison of existing benchmarks in terms of multi-turn planning and plan action capabilities.
Figure 2: Overview of the proposed benchmark PDEU-Bench
Models
Pref. Score
Qual. Score
Plan Definition
Plan Update
Plan Execution
Pref. Acc.
Qual. Acc.
Joint Acc.
Pref. Acc.
Qual. Acc.
Joint Acc.
Pref. Acc.
Closed-Source Models
gpt-5.6-terra
0.6632
0.4350
0.7908
0.7329
0.6027
0.3425
0.1370
0.1370
0.8562
claude-sonnet-5
0.6134
0.3219
0.7055
0.5616
0.4452
0.2192
0.0822
0.0753
0.9157
grok-4.5
0.6112
0.3767
0.8151
0.7123
0.5959
0.0858
0.0411
0.0068
0.9328
gpt-5.4
0.5551
0.3630
0.8082
0.6918
0.5959
0.1781
0.0342
0.0205
0.6791
Table 2: Average accuracy across models for plan definition, update, and execution on “easy” and “hard” tasks. Bold and underlined values mark the best and second best results.
Models
Methods
Pref. Score
Qual. Score
Plan Definition
Plan Update
Plan Execution
Pref. Acc.
Qual. Acc.
Joint Acc.
Pref. Acc.
Qual. Acc.
Joint Acc.
Pref. Acc.
Gemini
Baseline
0.4414
0.1164
0.2329
0.1507
0.0411
0.1849
0.0822
0.0616
0.9066
Reminder
0.3949( ↓ )
0.1506( ↑ )
0.3219( ↑ )
0.2260( ↑ )
0.0753( ↑ )
0.1644( ↓ )
0.0753( ↓ )
0.0685( ↑ )
0.6986( ↓ )
RAG-based Methods
RAG
0.4497( ↑ )
0.1301( ↑ )
0.3836( ↑ )
0.2192( ↑ )
0.1370( ↑ )
0.2123( ↑ )
0.0411( ↓ )
0.0342( ↓ )
0.7534( ↓ )
PAG
0.4908( ↑ )
0.2226( ↑ )
0.5753( ↑ )
0.3699( ↑ )
0.2945( ↑ )
0.2055( ↑ )
0.0753( ↓ )
0.0753( ↑ )
0.6918( ↓ )
Table 3: Average accuracy on “easy” and “hard” tasks for gemini-3.1-flash-lite (Gemini) and deepSeek-v4-flash (Deepseek). " ↓ " and " ↑ " indicate decreases and improvements relative to the baseline respectively.
Model
Error Type
Planning Stage
Update
Definition
Execution
Deepseek v4-pro
E1: Preference omission
79.9
43.9
23.4
E2: Preference conflict
89.3
10.7
38.3
E3: Goal forgetting
86.0
7.5
–
E4: Logic error
93.5
47.2
–
Gemini-3.1 flash-lite
E1: Preference omission
81.3
75.0
22.2
Table 4: Error-type prevalence among failed instances (%) for the two baseline LLMs.“–” denotes not applicable.
Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent evaluations often report only end-to-end success, making it difficult to determine whether failures stem from planning or execution. We introduce Agent Planning Benchmark (APB), a planning-specific diagnostic benchmark with 4,209 multimodal cases across 22 domains and five settings, covering holistic planning, feedback-conditioned step-wise planning, and robustness under extraneous tools, broken tools, and unsolvable tasks. Across 12 MLLMs, APB reveals systematic weaknesses in long-horizon planning, tool-noise robustness, calibrated refusal, and inference-time refinement. We further validate APB on 200 ToolSandbox tasks and 200 τ2-bench tasks, where APB-guided refinement consistently improves plan correctness, plan grade, and downstream execution metrics across three representative models. APB thus serves as an upstream diagnostic complement to execution benchmarks. The APB benchmark and code are available in this URL.
Haoyu Sun, Wenxuan Wang, Mingyang Song +5
Tongji University · Shanghai AI Laboratory · Harbin Institute of Technology +5
LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons. However, existing benchmarks rarely evaluate planning under retrieval-limited tool visibility. To address this gap, we introduce PlanBench-XL, an interactive benchmark of 327 retail tasks over 1,665 tools that tests whether agents can iteratively retrieve usable tools, invoke them to uncover intermediate evidence for subsequent calls toward the final goal. PlanBench-XL further features an optional blocking mechanism that simulates real-world unpredictability through missing, failing, or distracting tool functions, forcing agents to detect disrupted paths and adapt at runtime. Experiments on ten leading LLMs show that massive-tool planning remains challenging: while GPT-5.4 achieves 51.90% accuracy in block-free settings, it collapses to 11.36% under the most severe blocking condition. Further analysis shows that agents are especially vulnerable when failures lack explicit error signals or when recovery requires longer alternative tool-use paths. These results establish PlanBench-XL as a testbed for diagnosing agentic planning failures and highlight the need for robust adaptive planning in long-horizon tasks with large, imperfect tool environments.
Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user-aligned tool-call trajectories under incomplete information. The benchmark is built from privacy-sanitized real interaction traces and combines structured persona profiles, public API-style tool ecosystems, and long-horizon multi-turn trajectories. It includes 10 user profiles, 36 tool sets, 1,065 turns, 170 unique tools, and evaluation-focused task types covering lack-of-information, single-tool, and multi-tool settings. Experiments with strong tool-use LLMs show that current models still have difficulty with personalized delegation. Multi-tool coordination, missing-constraint inference, and long-horizon behavioral consistency remain major bottlenecks. These results suggest that personalization evaluation should move beyond asking whether outputs sound user-specific and instead ask whether LLMs make correct decisions for the users they represent.
Xuexiong Yin, Zechuan Chen, Yongsen Zheng +5
Sun Yat-Sen University · Nanyang Technological University · Huawei Noah’s Ark Lab