Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, empirical results show existing methods proposed for question answering tasks severely degrade task performance when applied to agentic models. We trace this failure to two decisions: what to prune and how to recover. For pruning, one-shot importance estimates fail to track how the pruned model adapts. For recovery, offline distillation covers only teacher prefixes, while full-trajectory on-policy distillation causes student errors to compound across turns. In this work, we propose Trajectory-Anchored Pruning (TAP), the first structural pruning framework for reinforcement learning (RL)-trained agents. TAP couples structural pruning with efficient on-policy recovery, anchoring interactions to teacher trajectories while allowing the student to generate each reasoning-action response. A frozen dense teacher supervises the student's response prefixes, addressing within-response training-inference mismatch while preventing student-induced deviations from propagating across training turns. Instead of one-shot pruning, TAP re-scores channels using gradients of the recovery objective on the recovered student, connecting iterative channel selection to the evolving policy. With 60% of FFN channels removed, TAP retains 99.2% and 88.0% of the dense 7B agents' task success rates on ALFWorld and WebShop, respectively, while reducing GPU time per successful task by approximately 22% and 17%. These results demonstrate effective structural compression of long-horizon agents under a limited recovery budget.
Figures & tables
Figure 1: RL-trained agents are highly compressible, but only with suitable pruning and recovery. Share of the dense agent’s task success retained by 7B agents as FFN channels are removed. Pruning without recovery collapses beyond 40% removal, whereas TAP retains about 80 – 99% of dense success up to 80% removal.
Figure 2: TAP couples how to recover and what to prune through one objective. Top: teacher trajectories provide a fixed pool of anchor contexts. Left (how to recover): within anchor contexts, the student generates reasoning and action tokens, the frozen teacher supervises every generated token, and minimizing LAOPD updates the student. Right (what to prune): gradients of the same objective on the last recovery batch score channels at the recovered student, and remove the lowest-scoring channels. The prune–recover–rescore cycle repeats until the target removal ratio is reached.
ρ
Method
ALFWorld
WebShop
In SR (%)
Out SR (%)
Retention (%)
SR (%)
Retention (%)
Qwen2.5-1.5B
0%
Dense RL
87.86±2.86
83.83±1.55
100
77.93±0.12
100
RL + Pruning
36.19±2.30
35.82±2.69
41.96
4.07±1.22
5.22
LLM-Pruner + KD
73.81±4.65
61.94±4.54
78.95
14.00±0.72
17.96
CFSP + KD
82.62±3.67
75.12±0.86
91.82
39.73±16.01
50.98
Table 1: TAP retains the most task success, and its lead grows with the removal ratio. We test whether TAP preserves the success of RL-trained agents under aggressive FFN pruning, compared with pipelines that recover by online RL (RL + Pruning), offline logits KD (+ KD), or local reconstruction alone (SlimGPT). Entries are success rates over three seeds and retention relative to the dense agent; ALFWorld retention averages the In and Out splits. Bold / underline : best/second best per setting. At the milder ratio of each model, the strongest baselines stay within 3.5 retention points of TAP; at the aggressive ratio, TAP’s retention exceeds the best baseline’s by 8 – 60 points. The redundancy of task-specific agents is therefore exploitable, but only with suitable channel selection and recovery.
60% FFN removal
80% FFN removal
Variant
Schedule
Ranking
ALF In
ALF Out
WebShop
ALF In
ALF Out
WebShop
A
One-shot
Fixed
92.38±2.51
86.32±2.40
66.87±0.23
73.81±2.06
61.69±2.28
57.40±1.22
B
Cubic
Fixed
90.48±1.65
87.81±0.86
63.27±0.12
77.38±2.70
63.93±1.14
57.93±1.01
C (TAP)
Cubic
Recalibrated
92.62±1.65
89.80±1.55
68.67±0.95
76.19±1.80
72.14±3.11
62.33±1.21
U
Uniform
Recalibrated
90.71±1.24
89.55±0.75
65.80±1.11
50.00±1.89
43.03±0.86
51.93±1.92
Table 2: Iterative pruning helps when channels are re-scored with the recovery objective. To isolate what to prune as the student adapts, we fix the recovery objective and vary how channels are removed on Qwen2.5-7B. Cubic and uniform schedules use three pruning stages. A and B delete the same final channels, so B vs. A isolates interleaving recovery between cuts; B vs. C isolates recalibration; U vs. C compares uniform and cubic schedules. Bold / underline : best/second best. Iterating with a fixed ranking gives mixed changes, whereas recalibration improves five of six metrics, and C improves on one-shot pruning in all six. The gain therefore comes from coupling channel selection with recovery.
Source
60% FFN removal
80% FFN removal
Recovery
Context
Resp.
ALF In
ALF Out
WebShop
ALF In
ALF Out
WebShop
D: Logits KD
Teacher
Teacher
93.33±1.65
88.06±1.29
65.60±0.72
75.95±2.18
70.90±5.38
59.80±0.80
F: SeqKD
Teacher
Teacher
92.62±1.09
85.32±0.86
61.87±0.58
70.95±4.86
65.42±3.76
55.40±1.60
E: Traj. OPD
Student
Student
92.14±0.71
86.07±1.55
64.07±0.23
49.76±2.89
41.54±5.80
62.67±0.61
C: Anchored
Teacher
Student
92.62±1.65
89.80±1.55
68.67±0.95
76.19±1.80
72.14±3.11
62.33±1.21
Table 3: Recovery works best when the student generates its own responses within teacher interaction contexts. To isolate how a pruned agent should be recovered, we fix the pruning path to TAP’s and vary only the recovery objective on Qwen2.5-7B. Context and Resp. (response) indicate whose interaction contexts the student conditions on and whose responses it is supervised on: logits KD and SeqKD use teacher responses with soft and hard targets, full-trajectory OPD uses the student’s own rollouts, and anchored OPD uses student responses within teacher contexts. Entries are SR; bold / underline : best/second best. Anchored OPD is the best or within one standard deviation of the best in every column, whereas full-trajectory OPD collapses on ALFWorld at 80% removal. Comparing rows isolates the two design choices: supervising the student’s own responses improves on KD, and anchoring its contexts avoids the collapse of full-trajectory OPD.
Speedup
ALFWorld
WebShop
Model
Params
Prefill ↑
Decode ↑
Steps ↓
E2E ↑
CPS ↓
Steps ↓
E2E ↑
CPS ↓
Dense
7.62B
1.00×
1.00×
11.8
1.00×
1.00×
6.2
1.00×
1.00×
TAP, ρ=60%
4.19B
1.71×
1.73×
13.6
1.30×
0.78×
7.0
1.38×
0.83×
TAP, ρ=80%
3.05B
1.94×
2.26×
21.6
0.77×
1.61×
6.1
1.76×
0.71×
Table 4: Compression lowers the cost per successful task when success is retained. Agent cost depends on speed, steps per task, and success, so we report all three for dense and TAP-pruned Qwen2.5-7B agents. Steps: agent steps per task; E2E: end-to-end speedup per task; CPS: GPU time per success, including failed attempts, relative to dense (Appendix C.8 ). CPS falls by 17 – 29% except on ALFWorld at 80% removal, where 1.8× more steps per task outweigh faster inference.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Pruning during GiGPO continuation on Qwen2.5-1.5B agents, which removes one additional percentage point of FFN channels per optimizer step. Solid lines show periodic validation and dashed lines training-batch success. ALFWorld validation pools the 140 In and 134 Out tasks; WebShop uses 128 fixed tasks. Validation uses temperature 0.4 . Validation success falls as removal increases and drops sharply beyond about 30% removal on WebShop, so continued RL does not keep the pruned agent functional under this schedule.
ρ
FFN width
Split
Mean ± std
20%
8960→7168
In
87.62±2.51
20%
8960→7168
Out
79.85±2.59
Appendix
Table 5: SlimGPT on SPEAR ALFWorld 1.5B at 20% FFN channel removal. Success rates are percentages.
Environment
Removal ρ
Jaccard JD
Deleted RD
Retained RK
ALFWorld
60%
82.66
9.49
14.24
WebShop
60%
82.03
9.87
14.81
ALFWorld
80%
89.44
5.58
22.31
WebShop
80%
90.13
5.19
20.76
Appendix
Table 6: Final channel masks of TAP (C) versus one-shot selection from the same initial ranking (A) on Qwen2.5-7B. JD is the Jaccard similarity of the deleted sets; RD and RK are the fractions of A’s deleted and retained channels that C replaces. All values are percentages. Although the deleted sets largely overlap, recalibration replaces 14 – 22% of the retained channels.
Calibration objective
In SR (%)
Out SR (%)
KL (TAP)
92.62±1.65
89.80±1.55
NLL
91.90±1.80
84.33±2.99
Appendix
Table 7: Effect of the calibration objective on Qwen2.5-7B at 60% FFN channel removal on ALFWorld. NLL replaces KL for channel-importance estimation while the rest of TAP is unchanged. KL yields similar in-distribution but higher out-of-distribution success.
Steps (%)
Episodes (%)
Method
Missing/broken action block
Invalid action
Repeated action
15-step limit
CFSP+KD
87
95
58
84
LLM-Pruner+KD
93
93
82
98
TAP
<1
<1
4
4
Appendix
Table 8: Free-generation failure diagnostics for the same WebShop 1.5B setting. Percentages are computed on the 500 WebShop test goals. Action-block and invalid-action rates are computed over steps; repeated-action and step-limit rates are computed over episodes. Repeated action denotes at least three consecutive identical executed actions within an episode. Flags can overlap.
Model
Environment
Removal
Episodes
Response tokens Q (M)
Total tokens (M)
1.5B
ALFWorld
40% , 60%
500
0.86
5.65
1.5B
WebShop
40%
500
0.18
3.10
1.5B
WebShop
60%
1,000
0.35
6.18
7B
ALFWorld
60% , 80%
500
0.48
3.72
7B
WebShop
60%
500
0.60
3.96
7B
WebShop
80%
1,000
1.21
7.96
Appendix
Table 9: Recovery budget for each main-table setting. All distillation-based methods within a setting share the supervised response-token budget Q ; total tokens also count the unsupervised prompts.
Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successful trajectories are structurally compressible, decomposed into a small set of reusable abstract patterns. To formalize this, we introduce ReuseRL, which grounds agentic RL in the Minimum Description Length (MDL) principle. ReuseRL extracts a shared skill dictionary from successful trajectories and augments the RL objective with a segmentation cost, explicitly penalizing idiosyncratic behaviors that encode poorly. We prove a PAC-Bayes bound guaranteeing that a dictionary extracted from successful trajectories has bounded expected description length on future successful behavior. Across ALFWorld, TextWorld-Cooking, and Countdown-Stepwise, ReuseRL improves in- and out-of-distribution success over vanilla GRPO and strong round-length baselines.
Zhikun Xu, Yu Feng, Jacob Dineen +3
Arizona State University · University of Pennsylvania · University of Southern California
Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear. We introduce TRIAL, a trajectory-relative hindsight distillation framework with a unified turn-aligned scoring protocol. For each decision turn, TRIAL extracts an outcome view of that decision's realized consequence and evaluates the same response under ordinary and hindsight-conditioned contexts. The signed log-probability gap determines the direction and local strength of token-level supervision, while turn-level magnitudes are normalized jointly over the realized trajectory. The resulting allocation multipliers have an eligible-token-weighted mean of one, redistributing dense supervision across turns while fixing its average multiplier. Experiments on WebShop and ALFWorld with different backbones show that TRIAL outperforms GRPO across all eight combinations of backbone, environment, and evaluation metric, while achieving the best or tied-best performance among six methods on six of them. On WebShop with Qwen3-1.7B, TRIAL improves the success rate from 56.4% to 75.2% and the task score from 78.7% to 85.7%. Controlled ablations further show that trajectory-relative turn allocation provides substantial gains beyond those of dense hindsight distillation alone.
Haoyu Zheng, Yun Zhu, Qing Wang +1
Zhejiang University · Tencent · Shanghai AI Laboratory
Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR improves average reward from 0.699 to 0.718, while reducing input, output, and cache read tokens by 25.5%, 14.4%, and 33.3%, respectively. Ablations reveal trajectory amplification, where local reasoning deletion produces nonlinear changes in total computation by altering subsequent interaction. Representation probing, activation patching, and controlled trajectory analyses further suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback. These results characterize agent reasoning as dynamic working state rather than permanent interaction history.
Mingxuan Wang, Fei Luo, Bo Wang +6
TierFlow Team · Gaoling School of Artificial Intelligence, Renmin University of China · Tsinghua University