Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
Authors: Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye, Qingyu Yang, Peng Ye
Organizations: Shanghai AI Laboratory · University of Science and Technology of China · Fudan University · KTH Royal Institute of Technology · The Chinese University of Hong Kong
Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-reweighting methods and show that they assign nonnegative coefficients to demonstrated tokens. Consequently, they can suppress or amplify supervised updates, but cannot reverse harmful features once learned. Moreover, larger training weights do not amount to feature extrapolation, since they change the optimization trajectory rather than scale a fixed SFT direction. We argue that reversal and extrapolation require a stable reference frame defined by a fixed SFT delta. Motivated by this, we propose SCALE (Selective Control of Adaptation via Local Entropy), an entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded token- and module-specific gates by minimizing predictive entropy alone. These gates suppress, reverse, or extrapolate frozen SFT features according to their alignment with entropy reduction. Across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base, SCALE achieves mathematical-reasoning averages of 37.84, 43.60, and 36.57, exceeding the strongest corresponding baselines while remaining competitive on general-retention benchmarks. It also attains the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three models. These results suggest that effective SFT correction can benefit from controlling how already learned residuals are used, rather than only modifying how they are learned.
Figures & tables
Mathematical reasoning
General Retention
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg
MMLU-P
BBH
OBQA
Avg
Qwen2.5-Math-1.5B ( Yang et al., 2024a )
Base
21.87
5.24
9.18
1.66
12.81
10.15
20.24
26.32
30.40
25.65
LoRA ( Hu et al., 2022 )
39.60
9.49
10.80
0.41
17.03
15.47
27.89
31.76
33.80
31.15
DoRA ( Liu et al., 2024 )
40.10
9.39
10.69
1.24
16.88
15.66
27.58
32.20
33.60
31.13
ProFit ( Liu et al., 2026b )
56.12
21.93
22.69
2.29
31.56
26.92
28.02
29.78
33.80
30.53
Table 1: Mathematical reasoning and general-capability retention results (%). Mathematical-reasoning scores are Avg@16. Boldface and underlining mark the best and second-best adapted scores in each column.
Figure 3: Gate-range sensitivity across backbones. We evaluate all endpoint pairs to characterize the intervention space. The selected maxima are reported as oracle sweep results, while fixed-range results are used for controlled comparisons.
Qwen2.5-Math-1.5B
Qwen2.5-Math-7B
Qwen3-4B-Base
Method
HE
HE+
MBPP
Avg
HE
HE+
MBPP
Avg
HE
HE+
MBPP
Avg
Base
28.66
17.68
32.80
26.38
42.28
39.02
34.20
38.50
43.09
39.02
37.00
39.70
LoRA ( Hu et al., 2022 )
34.15
25.61
30.80
30.19
48.37
43.29
38.00
43.22
51.22
45.73
40.20
45.72
DoRA ( Liu et al., 2024 )
31.91
28.05
31.20
30.39
46.14
43.29
36.20
41.88
50.81
45.12
40.00
45.31
DFT ( Wu et al., 2026 )
34.96
29.88
33.60
32.81
48.58
42.07
36.00
42.22
44.51
38.41
37.20
40.04
ProFit ( Liu et al., 2026b )
35.57
30.49
33.80
33.29
47.56
43.29
35.60
42.15
48.78
43.90
40.00
44.23
Table 2: Code-generation results (%). Boldface and underlining mark the best and second-best adapted scores in each column.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Definition
x
Input prompt
y=(y1,…,yT)
Demonstrated response tokens
ct=(x,y<t)
Prefix context at token position t
a∈V
Next-token action from vocabulary V
πθ(a∣ct)
Model next-token distribution parameterized by θ
zt
Output logits at position t
Appendix
Table 3: Notation used throughout the paper.
Method
CE coefficient wt
Signal / qualification
SFT
1
Uniform demonstration weighting.
DFT
sg[pt]
Probability-weighted CE ( Wu et al., 2026 ) .
ProFit
1[sg[pt]>τp]
Probability mask ( Liu et al., 2026b ) .
EAFT
Httop-20/log20
Normalized top- 20 entropy; frozen-coefficient CE interpretation ( Diao et al., 2026 ) .
TALR
max{wmin,sg[e−ℓt/τℓ]}
Loss-dependent weight, dynamic temperature, and floor ( Lin et al., 2026 ) .
Offline token mask
mt∈{0,1}
Fixed retained positions, e.g., token cleaning ( Pang et al., 2025 ) .
Appendix
Table 4: Nonnegative coefficients of demonstration-action CE components. Entries describe CE components. The EAFT row reports its published entropy weight without specifying a stop-gradient convention.
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg@16
Negative
30.85
6.54
8.85
1.24
15.47
12.59
LoRA
39.60
9.49
10.80
0.41
17.03
15.47
Positive
47.99
15.24
20.08
4.16
25.47
22.59
Ours
71.60
31.44
31.34
8.55
46.25
37.84
Appendix
Table 5: Zero-threshold entropy-sign ablation on Qwen2.5-Math -1.5B. Negative and Positive suppress deltas where ΔH<0 and ΔH>0 , respectively. Scores are Avg@16. Boldface and underlining mark the best and second-best scores in each column.
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg@16
LoRA
39.60
9.49
10.80
0.41
17.03
15.47
LoRA +20k
40.12
9.72
11.58
1.04
18.91
16.27
SCALE w/ CE
41.97
10.31
11.36
0.82
16.41
16.17
SCALE w/ Entropy
71.60
31.44
31.34
8.55
46.25
37.84
Appendix
Table 6: Calibration-loss ablation on Qwen2.5-Math -1.5B (%). Scores are Avg@16; boldface and underlining mark the best and second-best values.
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg@16
LoRA
39.52±0.07
9.60±0.33
10.59±0.39
0.62±0.21
16.72±1.43
15.41±0.44
DFT
65.30±1.03
21.42±1.89
27.22±0.37
7.64±1.78
39.95±0.65
32.31±0.92
SCALE
71.73±0.57
27.18±2.47
31.56±0.31
10.21±1.08
48.44±1.18
37.83±0.20
Appendix
Table 7: Three-seed mathematical-reasoning results on Qwen2.5-Math -1.5B (%). Scores are Avg@16, reported as mean ± sample standard deviation.
Backbone
λ<0
0≤λ≤1
λ>1
Qwen2.5-Math -1.5B
39.36
17.81
42.82
Qwen2.5-Math -7B
44.34
19.11
36.55
Appendix
Table 8: Distribution of gate multipliers on generated mathematical-reasoning tokens (%).
Gate applied at inference
Qwen2.5-Math -1.5B
Qwen2.5-Math -7B
Original λ
37.05
42.52
No reversal, max(λ,0)
20.10 ( −16.95 )
33.37 ( −9.15 )
No extrapolation, min(λ,1)
30.06 ( −6.99 )
30.06 ( −12.46 )
Unit interval, clip(λ,0,1)
19.44 ( −17.61 )
24.07 ( −18.45 )
Appendix
Table 9: Mathematical-reasoning gate interventions. Scores are Math Avg@16; parentheses show the change from the original gate in percentage points.
Gate range
Gate applied at inference
pass@1
Change
[0,6]
Original λ
57.11
—
Unit interval, clip(λ,0,1)
47.15
−9.96
[−3,3]
Original λ
49.80
—
No reversal, max(λ,0)
52.64
+2.85
No extrapolation, min(λ,1)
45.12
−4.67
Unit interval, clip(λ,0,1)
47.97
−1.83
Appendix
Table 10: Gate interventions for code generation on Qwen2.5-Math -7B. Scores are HumanEval pass@1 (%).
Figure 4: Single-GPU training time (minutes) for mathematical reasoning and code generation on Qwen2.5-Math-1.5B and 7B.
Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pressure on uncertain, low-confidence tokens, leading to suboptimal training dynamics. We propose Trimmed Logit-Gap SFT (TrimSFT), a simple token-level reweighting method that scales the SFT loss according to the logit gap between the gold token and its strongest competitor. TrimSFT trims supervision away from both extremes: tokens already mastered (large logit gap) and tokens weakly supported by the current model (small or negative logit gap), concentrating learning within an intermediate logit-gap region between them. We instantiate this principle with a Gaussian weight centered at margin m with bandwidth τ, requiring no reference model or additional forward pass. We evaluate TrimSFT on six base models from the Llama, Qwen, and DeepMath families across five mathematical reasoning benchmarks. TrimSFT consistently improves over standard SFT, achieving the best average performance on five out of six models, with gains of up to +26.9 points over SFT on MATH500. Further analyses show that the bandwidth τ matters more than the exact margin location, and that half-trim variants that remove supervision pressure from only one side yield inferior trade-offs. A token-level logit-gap distribution analysis suggests that TrimSFT reshapes model confidence in a more balanced way than uniform SFT or monotonic reweighting methods. These results suggest that reasoning SFT can benefit from trimming both extremes rather than treating all tokens uniformly.
Supervised fine-tuning (SFT) is an efficient approach for downstream task adaptation and often serves as the initialization stage for reinforcement learning (RL), but it can show weaker generalization than RL. A key limitation is its off-policy objective: SFT fits fixed demonstrations token by token, including targets poorly aligned with the model's pretrained distribution, which can lead to overfitting. A recent line of work addresses this issue by assigning larger training weights to tokens better aligned with the current model's predictive distribution, with the intuition that fitting these tokens are less distortive to the model's pretrained knowledge and representations. However, computing the token weights from the model that is currently fine-tuned entangles token weights with the optimization trajectory, inducing a self-reinforcing dynamics as the distribution rapidly departs from the pretrained model. To address this, we propose PriFT (Prior-support guided Fine-Tuning), which derives token weights from a frozen pretrained reference to obtain a stable reweighting signal unaffected by fine-tuning. This signal estimates prior support: the extent to which each target token is supported by the pretrained distribution. Across multiple existing token-reweighting rules, replacing the reweighting signal from the online model to pretrained model consistently improves performance. We introduce two instantiations: PriFT-prob uses pretrained token probability, while PriFT-mass selects tokens by cumulative probability mass under the pretrained distribution. Extensive experiments on mathematical reasoning, code generation, and medical question answering show that PriFT achieves state-of-the-art results among SFT baselines and provides a better initialization for subsequent RL training.
Supervised fine-tuning (SFT) provides the standard approach for teaching LLMs new behaviors from offline expert demonstrations. However, standard SFT uniformly fits all samples -- including those with low likelihood under the base model -- which can disproportionately drive training updates toward overfitting specific samples rather than learning the target behavior. Moreover, adapting to these unlikely samples induces substantial policy shifts that degrade prior capabilities. Existing methods mitigate this by filtering, regenerating, or down-weighting low-likelihood data. In doing so, they often suppress precisely the novel behaviors the base model has yet to learn. We propose InfoSFT, a principled weighting scheme for the SFT objective that concentrates learning signals on maximally informative, medium-confidence tokens -- those neither overly familiar to the base model nor too unlikely to cause instability. Requiring only a one-line modification to the standard token-wise loss, InfoSFT demonstrably improves generalization over vanilla SFT and likelihood-weighted baselines across math, code, and chain-of-thought tasks with diverse model families, while better preserving pre-existing capabilities.
Mahdi Sabbaghi, George Pappas, Adel Javanmard +1
University of Pennsylvania · University of Southern California