CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection
Authors: Yu Li, Yunlu Wan, Zijian Zhu, Han Luo, Chao Ren, Long-Fei Li, Lei Feng
Organizations: School of Computer Science and Engineering, Southeast University, Nanjing, China · KTH Royal Institute of Technology · Huawei Noah’s Ark Lab, China
Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools and composite skills coexist, skill use becomes a policy problem: the agent must decide whether the current state requires atomic fine control or skill-level abstraction. In this paper, we argue that effective skill use should be studied as adaptive tool granularity selection. The most direct training signal for this problem is to compare the consequences of atomic and skill choices available from the same state. Based on this view, we propose CIPO, a Counterfactual Imagination Policy Optimization framework for adaptive tool granularity. CIPO constructs executable skills through budget-constrained mining of successful tool-use trajectories and trains granularity decisions with counterfactual branch rollouts. For each base rollout, CIPO branches at the first eligible granularity decision and replaces the chosen action with a feasible atomic or skill alternative. The paired outcome difference serves as a supplementary reward for policy optimization. Experiments across multiple benchmarks and model backbones show that CIPO improves task success and decision efficiency over baselines. Further analyses show that CIPO learns effective skill use by improving the choice between atomic tools and composite skills based on the current state, without simply increasing skill frequency.
Figures & tables
Figure 1 : Pilot analyses on tool-use granularity. (a) Successful TOOLATHLON trajectories contain recurrent tool patterns. (b) Looser mining constraints increase both the skill library size and candidate overlap, motivating budget-constrained skill construction. (c) Frequent subsequences often require runtime execution support, including parameter binding, candidate selection, dataflow handling, and failure recovery. (d) Granularity preference varies across states, while SFT and “GRPO + Skills" fail to follow the oracle trend, suggesting the need for explicit granularity training.
Figure 2 : Overview of CIPO. Budgeted skill discovery mines recurring tool-use patterns and instantiates executable skills. Counterfactual imagination policy optimization creates paired branch rollouts by replacing the first eligible granularity choice with an atomic or skill alternative. The clipped reward difference provides a supplementary signal for adaptive tool granularity.
Method
TOOLATHLON
TOUCAN
TRAJECT-Bench
Avg. Comp./Raw ↓
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
General tool-use agents
ReAct
33.00
22.73
43.79
25.41
42.74
32.68
44.00
34.22
43.64
30.96
33.71
23.58
–
ToolLLM
21.10
11.62
35.43
20.43
26.28
17.59
33.34
27.11
50.10
31.50
54.68
32.00
–
Memory and workflow reuse
Table 1 : Main results across three benchmarks with Qwen2.5-7B and Llama3.1-8B. We report Tool F1 (%) and Task Success Rate (TSR, %). TSR is averaged over five independent judge evaluations of the same 100 evaluation episodes in each setting. Avg.Comp./Raw is the average compressed decision ratio for composite-action methods; lower values indicate stronger compression.
Figure 3 : Skill usage during training. CIPO approaches the coverage-derived reference, while GRPO + Skills remains near its initial level.
Method
Variant role
F1 ↑
Δ F1
TSR ↑
Δ TSR
C/R ↓
Decision red. ↑
SFT
skill-enabled imitation
36.05
0.00
22.91
0.00
0.958
4.2%
Standard GRPO
atomic-only RL
38.92
+2.87
26.74
+3.83
1.000
0.0%
GRPO + Skills
skill-enabled RL
41.20
+5.15
29.30
+6.39
0.912
8.8%
w/o Budgeted Mining
removes mining budget
42.82
+6.77
31.75
+8.84
0.872
12.8%
CIPO-random
random eligible branch
46.72
+10.67
34.90
+11.99
0.858
14.2%
CIPO-first
first eligible branch
48.23
+12.18
36.48
+13.57
0.845
15.5%
Table 2 : Ablation results averaged over all benchmarks and backbones. Δ F1 and Δ TSR are computed relative to the reported skill-enabled SFT baseline. C/R denotes Comp./Raw, and Decision red. denotes 1−C/R .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Meaning
Default value
K
maximum merge rounds
50
fmin
minimum adjacent-pair frequency
3
Lmax
maximum atomic length of a candidate
6
Lmin
minimum atomic length of a candidate
2
ρmin
minimum usage ratio for pruning
0.01
Appendix
Table 3: Default hyperparameters for budget-constrained skill mining.
Figure 4 : Training reward curves of CIPO across three benchmarks and two backbone models. The reward generally increases during training across TOOLATHLON, TOUCAN, and TRAJECT-Bench, suggesting that the counterfactual branch signal can be optimized together with the task reward.
Method
TOOLATHLON
TOUCAN
TRAJECT-Bench
Avg. F1
Avg. TSR
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
SFT
25.88 / 15.24
39.42 / 23.58
34.56 / 19.72
36.71 / 22.84
37.04 / 24.30
42.69 / 31.78
36.05
22.91
Standard GRPO
28.90 / 18.42
42.10 / 27.96
38.25 / 24.70
40.68 / 27.65
40.20 / 27.80
43.39 / 33.91
38.92
26.74
GRPO + Skills
31.10 / 22.20
44.35 / 32.55
40.20 / 30.10
42.30 / 31.20
42.85 / 29.10
46.40 / 30.65
41.20
29.30
w/o Budgeted Mining
31.40 / 24.20
45.10 / 34.20
42.90 / 31.50
44.60 / 33.20
44.80 / 30.10
48.10 / 37.30
42.82
31.75
w/o Executable Instantiation
31.50 / 23.70
44.80 / 33.60
41.80 / 30.70
43.70 / 32.50
43.60 / 29.80
47.20 / 36.00
42.10
31.05
Appendix
Table 4 : Detailed ablation results across all benchmark-backbone settings. Each entry reports F1 / TSR. Avg. reports the arithmetic mean over the six settings.
Figure 5 : Sensitivity analysis of the counterfactual reward. We vary one hyperparameter at a time around the default setting, including the reward scale α and the clipping bound ϵcf . Default CIPO-first achieves the best overall balance across benchmarks and models.
Hyperparameter
Value
Input trajectories
successful trajectories only
Token abstraction
tool name + parameter-key signature
Maximum merge rounds K
50
Minimum pair frequency fmin
3
Minimum skill length
2 atomic tools
Maximum skill length Lmax
6 atomic tools
Appendix
Table 6: Hyperparameters for budget-constrained skill mining.
Table 9: CIPO policy-optimization hyperparameters for Qwen-7B and Llama-8B.
Hyperparameter
Value
Skill reward weight λskill
0.30
Complete skill reward
1.00
Passed internal-step reward
0.20
Failed internal-step penalty
−0.50
Efficiency reward gate
0.40
Under-exploration penalty
0.05
Appendix
Table 10: Reward hyperparameters.
Hyperparameter
Value
Evaluation episodes per setting
100
Maximum turns
30
Maximum expanded atomic calls
50
Inference temperature
0.0
Top-p
0.95
Batch size
1
Appendix
Table 11: CIPO evaluation hyperparameters.
Benchmark
Backbone
Prec.
Rec.
NTAcc
Prog.
Skill Use
Skill Cov.
Comp.
Raw
TOOLATHLON
Qwen-7B
45.30
28.75
19.72
83.00
16.03
29.17
6.30
7.37
TOOLATHLON
Llama-8B
59.50
40.70
23.67
95.00
29.24
31.25
8.55
11.07
TOUCAN
Qwen-7B
47.80
49.33
47.80
78.00
9.50
27.50
7.30
8.49
TOUCAN
Llama-8B
49.70
51.35
50.20
84.00
22.80
42.30
8.95
11.02
TRAJECT-Bench
Qwen-7B
53.40
48.80
52.15
78.00
9.10
19.15
4.55
5.00
TRAJECT-Bench
Llama-8B
60.05
52.22
46.24
96.00
2.04
21.28
12.25
14.23
Appendix
Table 12 : Additional evaluation metrics of CIPO across benchmarks and backbones.
Benchmark
Backbone
F1
TSR
Comp./Raw
TOOLATHLON
Qwen-7B
35.17±1.42
30.78±1.65
0.855±0.024
TOOLATHLON
Llama-8B
48.32±1.87
40.55±2.01
0.772±0.028
TOUCAN
Qwen-7B
48.55±1.62
35.50±1.70
0.860±0.024
TOUCAN
Llama-8B
50.51±1.70
41.32±1.85
0.812±0.027
TRAJECT-Bench
Qwen-7B
50.99±1.45
32.00±1.63
0.910±0.020
TRAJECT-Bench
Llama-8B
55.86±1.44
38.75±1.81
0.861±0.018
Appendix
Table 13: Uncertainty estimates for CIPO. Tool F1 and Comp./Raw are reported with episode-level bootstrap 95% confidence interval half-widths. TSR is reported as mean ± standard deviation across five independent judge evaluations of the same 100 evaluation episodes.
Benchmark
Avg. tool invocations / episode
# Unique tools
Main difficulty
TOOLATHLON
∼13.5
1,250
Long-horizon dependency
TOUCAN
∼3.2
15,294
Large-scale tool selection
TRAJECT-Bench
∼6.4
715
Trajectory coherence
Appendix
Table 14: Statistics of the processed benchmark splits used in our experiments.
Method
yes -only strict success All six settings
Execution-based success TOOLATHLON
SFT
16.32±0.42
11.67±0.58
Standard GRPO
19.63±0.51
14.83±0.76
GRPO + Skills
21.40±0.47
17.67±0.58
CIPO
27.68±0.54
25.17±0.76
Appendix
Table 15: Strict and execution-based task success (%). The yes -only column averages all six benchmark–backbone settings and reports mean ± standard deviation across five judge runs. The execution-based column covers TOOLATHLON and reports mean ± standard deviation across three training seeds after averaging the two backbones.
Figure 6 : Effect of the BPE merge budget K . TSR improves as more reusable routines are mined, but declines when the skill library becomes too large and redundant. The default value K=50 gives the best average TSR.
Figure 7 : Task-level decision compression. Each point represents one evaluation episode. The diagonal line indicates no compression. Compared with GRPO + Skills, CIPO places more episodes above the diagonal, showing that composite skills allow fewer policy-level decisions to execute more atomic operations.
Method group
Built from training trajectories
Used at inference as
Comp./Raw reported
General tool-use agents
no
atomic tool policy
no
Memory and workflow reuse
yes
prompt guidance
no
Tool composition
yes
composite action or compiled interface
yes
Skill construction
yes
prompt skill or composite action
yes
CIPO
yes
executable composite skills + trained policy
yes
Appendix
Table 16: Summary of baseline adaptation. “Prompt” means that the method injects memory, workflow, or skill descriptions into the prompt. “Action” means that the method exposes composite tools or skills as callable actions.
Skill libraries have become a practical way for LLM agents to reuse procedural experience across tasks. However, existing systems typically treat skills as flat, single-resolution prompt blocks. This creates a tension between relevance and cost: injecting coarse skills can introduce irrelevant or misleading context, while rewriting entire skills is expensive and often unnecessary. We propose SkillLens, a hierarchical skill-evolution framework that organizes skills into a four-layer graph of policies, strategies, procedures, and primitives, and retrieves them at mixed granularity. Given a task, SkillLens first retrieves semantically relevant skill seeds, expands them through degree-corrected random walk over the skill graph, and then uses a verifier to decide whether each visited unit should be accepted, decomposed, rewritten, or skipped. This enables the agent to reuse compatible subskills directly while adapting only locally mismatched components. To improve the system over time, SkillLens further refines multi-granularity skills and verifier in order to improve its routing decisions. We provide theoretical analysis showing that mixed-granularity adaptation incurs sublinear cost under sparse mismatch assumptions and that the evolutionary update rule monotonically improves the validation objective until a local optimum. Across MuLocbench and ALFWorld, SkillLens consistently improves over strong skill-based baselines, achieving up to a 6.31 percentage-point Acc@1 gain for bug localization and raising agent success rate from 45.00% to 51.31%.
Agent skills provide a lightweight way to adapt LLM agents to specialized domains by storing reusable procedural knowledge in structured files. However, whether downloaded from third parties or self-generated, these skills are often unreliable, incomplete, or outdated. Existing skill-evolution methods often address these deficiencies through heuristic reflections without an explicit optimization formulation. In this paper, we propose SkillGrad, a gradient-descent-inspired framework for optimizing agent skills. SkillGrad treats the skill package as a structured parameter to optimize in a gradient descent fashion: task executions provide trajectory-level loss evidence, automatic diagnoses then provide text-based gradients that indicate the correction directions. To stabilize optimization across iterations, a momentum agent accumulates recurring diagnostic patterns into a persistent memory overlay. Finally, an LLM-based patcher executes the parameter update by applying layer-aware edits to the skill package. Evaluated on SpreadsheetBench Verified and WikiTableQuestions, SkillGrad consistently outperforms training-based skill evolution baselines across two backbone LLMs, improving over the strongest training-based baseline by 6.7 percentage points on average. Ablations further show that momentum and contrastive diagnosis both contribute to the final skill quality.
Hanyu Wang, Yifan Lan, Bochuan Cao +2
College of Information Sciences and Technology The Pennsylvania State University University Park, PA, USA
Large language model (LLM) agents increasingly interleave natural language reasoning with external tools such as web search and code execution. These tool-use policies are often optimized via reinforcement learning (RL), which can amplify spurious correlations in the training data. In this work, we study when and why RL-trained agents learn shortcut tool-selection policies: invoking tools based on superficial prompt cues rather than genuine task requirements. We construct controlled synthetic environments combining factual question answering and mathematical reasoning tasks, and inject cues that are strongly correlated with specific tools during training but causally irrelevant to tool necessity. Across counterfactual evaluations where cues are present but the associated tools are not required, agents exhibit substantial shortcut behavior, with spurious tool invocation rates increasing by up to 39 percent. However, shortcut formation is not universal: across the conditions we test, it arises only when the agent has already learned to use the target tool reliably, suggesting that task competence, rather than dataset imbalance alone, is a key factor in shortcut learning. A swapped-cue analysis further shows that semantic alignment between cues and tools substantially amplifies this effect. To mitigate these failures, we introduce a dense, decision-level reward in which an LLM judge evaluates the necessity of each tool call. This tool-necessity reward effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.
Yiwei Yang, Haoxiang Zhang, Bingbing Wen +6
University of Washington · University of California San Diego · Stanford University