Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focus on studying \emph{how to better balance the trade-off between acquiring new capabilities and preserving existing ones during agent trace SFT}. By comparing several baselines in our setup, standard SFT improves the target benchmark while lowering several non-target benchmark scores; meanwhile, simply constraining distributional drift using KL penalty or limiting the update magnitude did not avoid this regression trend. Motivated by recent token-wise adaptive learning objectives, this work proposes \textbf{Privilege-Guided SFT (PG-SFT)} to leverage turn-level information gain of agent trajectories as an indicator to adjust supervision strength. PG-SFT yields a more favorable observed trade-off on the evaluated benchmarks, substantially reducing distributional drift and broad capability degradation at the cost of slight degradation in target-task performance. Our findings suggest that balancing the acquisition--retention trade-off depends not only on whether the model is anchored to its base behavior, but also on where and how strongly supervision should depart from that behavior.}
Figures & tables
Figure 1 : A typical case of DFT’s supervision allocation. The value above each token is its target-token probability pt=pθ,t(yt) . DFT weights each token’s cross-entropy by the stop-gradient of this value, wt=sg[pt] ; the per-token gradient-norm ratio then equals this weight, ∥∇θℓtDFT∥2/∥∇θℓtSFT∥2=wt .
Figure 2 : Turn-level perplexity distribution of the base model on the training data.
Figure 3 : The same frozen base model scores aligned assistant tokens with and without an action hint. The resulting positive gain determines a shared turn-level coefficient αk .
Target Adaptation
Cross-Benchmark Capability
Overall
Backbone
Objective
SWE-bench ↑
GPQA ↑
BFCL ↑
LCB ↑
Macro Avg ↑
Δ vs. Base ↑
/89
/198
/200
/120
(%)
(pp)
Qwen3.5-4B
Base
54
151
111
60
60.61
0.00
SFT
59
133
100
48
55.87
−4.74
SFT+Base-KL
57
130
105
44
54.72
−5.89
PSFT
53
142
112
63
59.94
−0.67
Table 1: Target adaptation and cross-benchmark capability retention. Macro Avg averages the four per-benchmark success rates with equal weight; Δ vs. Base is its change over the Base model in accuracy percentage points (pp), computed within each backbone.
Qwen3.5-4B
Qwen3-4B-Thinking
Metric
Base
SFT
SFT+ Base-KL
PSFT
PG-SFT
Base
SFT
SFT+ Base-KL
PSFT
PG-SFT
NLL ↓
0.4302
0.3756
0.3767
0.4217
0.3968
0.7672
0.4903
0.5309
0.7423
0.5748
Base KL ↓
-
0.0732
0.0431
0.0026
0.0080
-
0.2454
0.1105
0.0014
0.0342
Table 2 : Teacher-forced analysis on held-out offline trajectories for both backbones.
Objective
NLL
Base KL
Fixed intensity
0.4006
0.0057
PG-SFT
0.3968
0.0080
Table 3: Offline statistics for PG-SFT v.s. the matched fixed-intensity control.
Objective
SWE-bench /89
GPQA /198
BFCL /200
LCB /120
Base
54
151
111
60
Fixed intensity
50
140
114
72
PG-SFT
59
150
111
61
Table 4: Downstream performance of PG-SFT vs. the matched fixed-intensity control.
Figure 4 : PG-SFT learning intensity α .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Shared setting
Value
Backbone
Qwen3.5-4B
Training trajectories
1,000 offline agent trajectories
Supervised region
complete assistant generation mask
Holdout panel
120 fixed trajectories
Seed
42
Precision
BF16
Appendix
Table 5: Training contract of the compared objectives on the Qwen3.5-4B backbone. The SFT, SFT+Base-KL, PG-SFT, and fixed-intensity arms share their adaptation and optimization configuration; the PSFT arm follows its own schedule and is therefore reported as a reference point rather than as a budget-matched comparison (Section 4.1 ).
Shared setting
Value
Backbone
Qwen3-4B-Thinking-2507
Training trajectories
4,992 offline agent trajectories
Supervised region
complete assistant generation mask
Holdout panel
120 fixed trajectories
Seed
42
Precision
BF16
Appendix
Table 6: Training contract of the compared objectives on the Qwen3-4B-Thinking-2507 backbone. Adaptation, optimization, seed, and data processing match Table 5 ; only the training-set size and the per-objective optimization budgets differ. In particular, the SFT+Base-KL arm executes 32 steps rather than a full epoch, so it is not budget-matched to the SFT and PG-SFT arms on this backbone.
Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream approaches for addressing these failures typically either fine-tune directly on target environments or generate synthetic data that is not targeted to the model's actual capability deficits, resulting in low sample efficiency and limited generalization. We introduce TRACE (Turning Recurrent Agent failures into Capability-targeted training Environments), an end-to-end system for environment-specific agent self-improvement. TRACE contrasts successful and failed trajectories to automatically identify missing capabilities, synthesizes a targeted training environment for each capability that rewards whether the capability is exercised, trains a LoRA adapter via reinforcement learning on each synthetic environment, and then trains a mixture-of-experts model over the capability adapters. TRACE can be effectively applied across different environments, improving over the base agent by +15.3 points on τ2-Bench, a customer-service agent benchmark, and by +15.0 points Pass@1 on SWE-Bench Verified, a software-engineering benchmark. TRACE outperforms the strongest external baselines, GEPA and SWE-RL, by +8.6 points and +8.4 points, respectively. In addition, TRACE is more sample-efficient than strong fine-tuning baselines: using fewer than one-fourth the number of rollouts, TRACE outperforms the best-performing baselines, GRPO and GEPA, and achieves higher final accuracy by +10.4 and +8.6 points on τ2-Bench.
Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents. We evaluate supervised fine-tuning (SFT) with LoRA, reinforcement learning (RL) via Group Relative Policy Optimization (GRPO), and SFT followed by GRPO across six Qwen3 models from 0.6B to 32B parameters, covering both in-distribution performance and cross-dataset transfer. SFT with LoRA is the strongest in-distribution method throughout the 0.6B-32B range and best in 15 out of 18 experimental settings. On cross-dataset transfer, the methods are closer: GRPO wins 29 out of 54 settings where training and test datasets differ, but its margin over SFT averages under one point, and SFT->GRPO is rarely strongest in either comparison. Dataset mixing gives consistently strong transfer while staying close to specialized in-distribution training, regardless of method. Additional analysis further confirms that LoRA outperforms full-parameter fine-tuning, demonstrating that LoRA better preserves pretrained agentic behavior.
Supervised fine-tuning (SFT) is an efficient approach for downstream task adaptation and often serves as the initialization stage for reinforcement learning (RL), but it can show weaker generalization than RL. A key limitation is its off-policy objective: SFT fits fixed demonstrations token by token, including targets poorly aligned with the model's pretrained distribution, which can lead to overfitting. A recent line of work addresses this issue by assigning larger training weights to tokens better aligned with the current model's predictive distribution, with the intuition that fitting these tokens are less distortive to the model's pretrained knowledge and representations. However, computing the token weights from the model that is currently fine-tuned entangles token weights with the optimization trajectory, inducing a self-reinforcing dynamics as the distribution rapidly departs from the pretrained model. To address this, we propose PriFT (Prior-support guided Fine-Tuning), which derives token weights from a frozen pretrained reference to obtain a stable reweighting signal unaffected by fine-tuning. This signal estimates prior support: the extent to which each target token is supported by the pretrained distribution. Across multiple existing token-reweighting rules, replacing the reweighting signal from the online model to pretrained model consistently improves performance. We introduce two instantiations: PriFT-prob uses pretrained token probability, while PriFT-mass selects tokens by cumulative probability mass under the pretrained distribution. Extensive experiments on mathematical reasoning, code generation, and medical question answering show that PriFT achieves state-of-the-art results among SFT baselines and provides a better initialization for subsequent RL training.