Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
Authors: Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye, Qingyu Yang, Peng Ye
Organizations: Shanghai AI Laboratory · University of Science and Technology of China · Fudan University · KTH Royal Institute of Technology · The Chinese University of Hong Kong
Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-reweighting methods and show that they assign nonnegative coefficients to demonstrated tokens. Consequently, they can suppress or amplify supervised updates, but cannot reverse harmful features once learned. Moreover, larger training weights do not amount to feature extrapolation, since they change the optimization trajectory rather than scale a fixed SFT direction. We argue that reversal and extrapolation require a stable reference frame defined by a fixed SFT delta. Motivated by this, we propose SCALE (Selective Control of Adaptation via Local Entropy), an entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded token- and module-specific gates by minimizing predictive entropy alone. These gates suppress, reverse, or extrapolate frozen SFT features according to their alignment with entropy reduction. Across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base, SCALE achieves mathematical-reasoning averages of 37.84, 43.60, and 36.57, exceeding the strongest corresponding baselines while remaining competitive on general-retention benchmarks. It also attains the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three models. These results suggest that effective SFT correction can benefit from controlling how already learned residuals are used, rather than only modifying how they are learned.
Figures & tables
Mathematical reasoning
General Retention
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg
MMLU-P
BBH
OBQA
Avg
Qwen2.5-Math-1.5B ( Yang et al., 2024a )
Base
21.87
5.24
9.18
1.66
12.81
10.15
20.24
26.32
30.40
25.65
LoRA ( Hu et al., 2022 )
39.60
9.49
10.80
0.41
17.03
15.47
27.89
31.76
33.80
31.15
DoRA ( Liu et al., 2024 )
40.10
9.39
10.69
1.24
16.88
15.66
27.58
32.20
33.60
31.13
ProFit ( Liu et al., 2026b )
56.12
21.93
22.69
2.29
31.56
26.92
28.02
29.78
33.80
30.53
Table 1: Mathematical reasoning and general-capability retention results (%). Mathematical-reasoning scores are Avg@16. Boldface and underlining mark the best and second-best adapted scores in each column.
Figure 3: Gate-range sensitivity across backbones. We evaluate all endpoint pairs to characterize the intervention space. The selected maxima are reported as oracle sweep results, while fixed-range results are used for controlled comparisons.
Qwen2.5-Math-1.5B
Qwen2.5-Math-7B
Qwen3-4B-Base
Method
HE
HE+
MBPP
Avg
HE
HE+
MBPP
Avg
HE
HE+
MBPP
Avg
Base
28.66
17.68
32.80
26.38
42.28
39.02
34.20
38.50
43.09
39.02
37.00
39.70
LoRA ( Hu et al., 2022 )
34.15
25.61
30.80
30.19
48.37
43.29
38.00
43.22
51.22
45.73
40.20
45.72
DoRA ( Liu et al., 2024 )
31.91
28.05
31.20
30.39
46.14
43.29
36.20
41.88
50.81
45.12
40.00
45.31
DFT ( Wu et al., 2026 )
34.96
29.88
33.60
32.81
48.58
42.07
36.00
42.22
44.51
38.41
37.20
40.04
ProFit ( Liu et al., 2026b )
35.57
30.49
33.80
33.29
47.56
43.29
35.60
42.15
48.78
43.90
40.00
44.23
Table 2: Code-generation results (%). Boldface and underlining mark the best and second-best adapted scores in each column.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Definition
x
Input prompt
y=(y1,…,yT)
Demonstrated response tokens
ct=(x,y<t)
Prefix context at token position t
a∈V
Next-token action from vocabulary V
πθ(a∣ct)
Model next-token distribution parameterized by θ
zt
Output logits at position t
Appendix
Table 3: Notation used throughout the paper.
Method
CE coefficient wt
Signal / qualification
SFT
1
Uniform demonstration weighting.
DFT
sg[pt]
Probability-weighted CE ( Wu et al., 2026 ) .
ProFit
1[sg[pt]>τp]
Probability mask ( Liu et al., 2026b ) .
EAFT
Httop-20/log20
Normalized top- 20 entropy; frozen-coefficient CE interpretation ( Diao et al., 2026 ) .
TALR
max{wmin,sg[e−ℓt/τℓ]}
Loss-dependent weight, dynamic temperature, and floor ( Lin et al., 2026 ) .
Offline token mask
mt∈{0,1}
Fixed retained positions, e.g., token cleaning ( Pang et al., 2025 ) .
Appendix
Table 4: Nonnegative coefficients of demonstration-action CE components. Entries describe CE components. The EAFT row reports its published entropy weight without specifying a stop-gradient convention.
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg@16
Negative
30.85
6.54
8.85
1.24
15.47
12.59
LoRA
39.60
9.49
10.80
0.41
17.03
15.47
Positive
47.99
15.24
20.08
4.16
25.47
22.59
Ours
71.60
31.44
31.34
8.55
46.25
37.84
Appendix
Table 5: Zero-threshold entropy-sign ablation on Qwen2.5-Math -1.5B. Negative and Positive suppress deltas where ΔH<0 and ΔH>0 , respectively. Scores are Avg@16. Boldface and underlining mark the best and second-best scores in each column.
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg@16
LoRA
39.60
9.49
10.80
0.41
17.03
15.47
LoRA +20k
40.12
9.72
11.58
1.04
18.91
16.27
SCALE w/ CE
41.97
10.31
11.36
0.82
16.41
16.17
SCALE w/ Entropy
71.60
31.44
31.34
8.55
46.25
37.84
Appendix
Table 6: Calibration-loss ablation on Qwen2.5-Math -1.5B (%). Scores are Avg@16; boldface and underlining mark the best and second-best values.
Method
Math500
Minerva
Olympiad
AIME24
AMC23
Avg@16
LoRA
39.52±0.07
9.60±0.33
10.59±0.39
0.62±0.21
16.72±1.43
15.41±0.44
DFT
65.30±1.03
21.42±1.89
27.22±0.37
7.64±1.78
39.95±0.65
32.31±0.92
SCALE
71.73±0.57
27.18±2.47
31.56±0.31
10.21±1.08
48.44±1.18
37.83±0.20
Appendix
Table 7: Three-seed mathematical-reasoning results on Qwen2.5-Math -1.5B (%). Scores are Avg@16, reported as mean ± sample standard deviation.
Backbone
λ<0
0≤λ≤1
λ>1
Qwen2.5-Math -1.5B
39.36
17.81
42.82
Qwen2.5-Math -7B
44.34
19.11
36.55
Appendix
Table 8: Distribution of gate multipliers on generated mathematical-reasoning tokens (%).
Gate applied at inference
Qwen2.5-Math -1.5B
Qwen2.5-Math -7B
Original λ
37.05
42.52
No reversal, max(λ,0)
20.10 ( −16.95 )
33.37 ( −9.15 )
No extrapolation, min(λ,1)
30.06 ( −6.99 )
30.06 ( −12.46 )
Unit interval, clip(λ,0,1)
19.44 ( −17.61 )
24.07 ( −18.45 )
Appendix
Table 9: Mathematical-reasoning gate interventions. Scores are Math Avg@16; parentheses show the change from the original gate in percentage points.
Gate range
Gate applied at inference
pass@1
Change
[0,6]
Original λ
57.11
—
Unit interval, clip(λ,0,1)
47.15
−9.96
[−3,3]
Original λ
49.80
—
No reversal, max(λ,0)
52.64
+2.85
No extrapolation, min(λ,1)
45.12
−4.67
Unit interval, clip(λ,0,1)
47.97
−1.83
Appendix
Table 10: Gate interventions for code generation on Qwen2.5-Math -7B. Scores are HumanEval pass@1 (%).
Figure 4: Single-GPU training time (minutes) for mathematical reasoning and code generation on Qwen2.5-Math-1.5B and 7B.