CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection
Authors: Yu Li, Yunlu Wan, Zijian Zhu, Han Luo, Chao Ren, Long-Fei Li, Lei Feng
Organizations: School of Computer Science and Engineering, Southeast University, Nanjing, China · KTH Royal Institute of Technology · Huawei Noah’s Ark Lab, China
Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools and composite skills coexist, skill use becomes a policy problem: the agent must decide whether the current state requires atomic fine control or skill-level abstraction. In this paper, we argue that effective skill use should be studied as adaptive tool granularity selection. The most direct training signal for this problem is to compare the consequences of atomic and skill choices available from the same state. Based on this view, we propose CIPO, a Counterfactual Imagination Policy Optimization framework for adaptive tool granularity. CIPO constructs executable skills through budget-constrained mining of successful tool-use trajectories and trains granularity decisions with counterfactual branch rollouts. For each base rollout, CIPO branches at the first eligible granularity decision and replaces the chosen action with a feasible atomic or skill alternative. The paired outcome difference serves as a supplementary reward for policy optimization. Experiments across multiple benchmarks and model backbones show that CIPO improves task success and decision efficiency over baselines. Further analyses show that CIPO learns effective skill use by improving the choice between atomic tools and composite skills based on the current state, without simply increasing skill frequency.
Figures & tables
Figure 1 : Pilot analyses on tool-use granularity. (a) Successful TOOLATHLON trajectories contain recurrent tool patterns. (b) Looser mining constraints increase both the skill library size and candidate overlap, motivating budget-constrained skill construction. (c) Frequent subsequences often require runtime execution support, including parameter binding, candidate selection, dataflow handling, and failure recovery. (d) Granularity preference varies across states, while SFT and “GRPO + Skills" fail to follow the oracle trend, suggesting the need for explicit granularity training.
Figure 2 : Overview of CIPO. Budgeted skill discovery mines recurring tool-use patterns and instantiates executable skills. Counterfactual imagination policy optimization creates paired branch rollouts by replacing the first eligible granularity choice with an atomic or skill alternative. The clipped reward difference provides a supplementary signal for adaptive tool granularity.
Method
TOOLATHLON
TOUCAN
TRAJECT-Bench
Avg. Comp./Raw ↓
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
F1 ↑
TSR ↑
General tool-use agents
ReAct
33.00
22.73
43.79
25.41
42.74
32.68
44.00
34.22
43.64
30.96
33.71
23.58
–
ToolLLM
21.10
11.62
35.43
20.43
26.28
17.59
33.34
27.11
50.10
31.50
54.68
32.00
–
Memory and workflow reuse
Table 1 : Main results across three benchmarks with Qwen2.5-7B and Llama3.1-8B. We report Tool F1 (%) and Task Success Rate (TSR, %). TSR is averaged over five independent judge evaluations of the same 100 evaluation episodes in each setting. Avg.Comp./Raw is the average compressed decision ratio for composite-action methods; lower values indicate stronger compression.
Figure 3 : Skill usage during training. CIPO approaches the coverage-derived reference, while GRPO + Skills remains near its initial level.
Method
Variant role
F1 ↑
Δ F1
TSR ↑
Δ TSR
C/R ↓
Decision red. ↑
SFT
skill-enabled imitation
36.05
0.00
22.91
0.00
0.958
4.2%
Standard GRPO
atomic-only RL
38.92
+2.87
26.74
+3.83
1.000
0.0%
GRPO + Skills
skill-enabled RL
41.20
+5.15
29.30
+6.39
0.912
8.8%
w/o Budgeted Mining
removes mining budget
42.82
+6.77
31.75
+8.84
0.872
12.8%
CIPO-random
random eligible branch
46.72
+10.67
34.90
+11.99
0.858
14.2%
CIPO-first
first eligible branch
48.23
+12.18
36.48
+13.57
0.845
15.5%
Table 2 : Ablation results averaged over all benchmarks and backbones. Δ F1 and Δ TSR are computed relative to the reported skill-enabled SFT baseline. C/R denotes Comp./Raw, and Decision red. denotes 1−C/R .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Meaning
Default value
K
maximum merge rounds
50
fmin
minimum adjacent-pair frequency
3
Lmax
maximum atomic length of a candidate
6
Lmin
minimum atomic length of a candidate
2
ρmin
minimum usage ratio for pruning
0.01
Appendix
Table 3: Default hyperparameters for budget-constrained skill mining.
Figure 4 : Training reward curves of CIPO across three benchmarks and two backbone models. The reward generally increases during training across TOOLATHLON, TOUCAN, and TRAJECT-Bench, suggesting that the counterfactual branch signal can be optimized together with the task reward.
Method
TOOLATHLON
TOUCAN
TRAJECT-Bench
Avg. F1
Avg. TSR
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
Qwen-7B
Llama-8B
SFT
25.88 / 15.24
39.42 / 23.58
34.56 / 19.72
36.71 / 22.84
37.04 / 24.30
42.69 / 31.78
36.05
22.91
Standard GRPO
28.90 / 18.42
42.10 / 27.96
38.25 / 24.70
40.68 / 27.65
40.20 / 27.80
43.39 / 33.91
38.92
26.74
GRPO + Skills
31.10 / 22.20
44.35 / 32.55
40.20 / 30.10
42.30 / 31.20
42.85 / 29.10
46.40 / 30.65
41.20
29.30
w/o Budgeted Mining
31.40 / 24.20
45.10 / 34.20
42.90 / 31.50
44.60 / 33.20
44.80 / 30.10
48.10 / 37.30
42.82
31.75
w/o Executable Instantiation
31.50 / 23.70
44.80 / 33.60
41.80 / 30.70
43.70 / 32.50
43.60 / 29.80
47.20 / 36.00
42.10
31.05
Appendix
Table 4 : Detailed ablation results across all benchmark-backbone settings. Each entry reports F1 / TSR. Avg. reports the arithmetic mean over the six settings.
Figure 5 : Sensitivity analysis of the counterfactual reward. We vary one hyperparameter at a time around the default setting, including the reward scale α and the clipping bound ϵcf . Default CIPO-first achieves the best overall balance across benchmarks and models.
Hyperparameter
Value
Input trajectories
successful trajectories only
Token abstraction
tool name + parameter-key signature
Maximum merge rounds K
50
Minimum pair frequency fmin
3
Minimum skill length
2 atomic tools
Maximum skill length Lmax
6 atomic tools
Appendix
Table 6: Hyperparameters for budget-constrained skill mining.
Table 9: CIPO policy-optimization hyperparameters for Qwen-7B and Llama-8B.
Hyperparameter
Value
Skill reward weight λskill
0.30
Complete skill reward
1.00
Passed internal-step reward
0.20
Failed internal-step penalty
−0.50
Efficiency reward gate
0.40
Under-exploration penalty
0.05
Appendix
Table 10: Reward hyperparameters.
Hyperparameter
Value
Evaluation episodes per setting
100
Maximum turns
30
Maximum expanded atomic calls
50
Inference temperature
0.0
Top-p
0.95
Batch size
1
Appendix
Table 11: CIPO evaluation hyperparameters.
Benchmark
Backbone
Prec.
Rec.
NTAcc
Prog.
Skill Use
Skill Cov.
Comp.
Raw
TOOLATHLON
Qwen-7B
45.30
28.75
19.72
83.00
16.03
29.17
6.30
7.37
TOOLATHLON
Llama-8B
59.50
40.70
23.67
95.00
29.24
31.25
8.55
11.07
TOUCAN
Qwen-7B
47.80
49.33
47.80
78.00
9.50
27.50
7.30
8.49
TOUCAN
Llama-8B
49.70
51.35
50.20
84.00
22.80
42.30
8.95
11.02
TRAJECT-Bench
Qwen-7B
53.40
48.80
52.15
78.00
9.10
19.15
4.55
5.00
TRAJECT-Bench
Llama-8B
60.05
52.22
46.24
96.00
2.04
21.28
12.25
14.23
Appendix
Table 12 : Additional evaluation metrics of CIPO across benchmarks and backbones.
Benchmark
Backbone
F1
TSR
Comp./Raw
TOOLATHLON
Qwen-7B
35.17±1.42
30.78±1.65
0.855±0.024
TOOLATHLON
Llama-8B
48.32±1.87
40.55±2.01
0.772±0.028
TOUCAN
Qwen-7B
48.55±1.62
35.50±1.70
0.860±0.024
TOUCAN
Llama-8B
50.51±1.70
41.32±1.85
0.812±0.027
TRAJECT-Bench
Qwen-7B
50.99±1.45
32.00±1.63
0.910±0.020
TRAJECT-Bench
Llama-8B
55.86±1.44
38.75±1.81
0.861±0.018
Appendix
Table 13: Uncertainty estimates for CIPO. Tool F1 and Comp./Raw are reported with episode-level bootstrap 95% confidence interval half-widths. TSR is reported as mean ± standard deviation across five independent judge evaluations of the same 100 evaluation episodes.
Benchmark
Avg. tool invocations / episode
# Unique tools
Main difficulty
TOOLATHLON
∼13.5
1,250
Long-horizon dependency
TOUCAN
∼3.2
15,294
Large-scale tool selection
TRAJECT-Bench
∼6.4
715
Trajectory coherence
Appendix
Table 14: Statistics of the processed benchmark splits used in our experiments.
Method
yes -only strict success All six settings
Execution-based success TOOLATHLON
SFT
16.32±0.42
11.67±0.58
Standard GRPO
19.63±0.51
14.83±0.76
GRPO + Skills
21.40±0.47
17.67±0.58
CIPO
27.68±0.54
25.17±0.76
Appendix
Table 15: Strict and execution-based task success (%). The yes -only column averages all six benchmark–backbone settings and reports mean ± standard deviation across five judge runs. The execution-based column covers TOOLATHLON and reports mean ± standard deviation across three training seeds after averaging the two backbones.
Figure 6 : Effect of the BPE merge budget K . TSR improves as more reusable routines are mined, but declines when the skill library becomes too large and redundant. The default value K=50 gives the best average TSR.
Figure 7 : Task-level decision compression. Each point represents one evaluation episode. The diagonal line indicates no compression. Compared with GRPO + Skills, CIPO places more episodes above the diagonal, showing that composite skills allow fewer policy-level decisions to execute more atomic operations.
Method group
Built from training trajectories
Used at inference as
Comp./Raw reported
General tool-use agents
no
atomic tool policy
no
Memory and workflow reuse
yes
prompt guidance
no
Tool composition
yes
composite action or compiled interface
yes
Skill construction
yes
prompt skill or composite action
yes
CIPO
yes
executable composite skills + trained policy
yes
Appendix
Table 16: Summary of baseline adaptation. “Prompt” means that the method injects memory, workflow, or skill descriptions into the prompt. “Action” means that the method exposes composite tools or skills as callable actions.