Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter β decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.
Figures & tables
Figure 1: Standard DPO penalizes every token of the rejected response, including the arithmetic; our method concentrates the penalty on the token that conflicts with the preferred response.
Benchmark
SFT
DPO
IPO
KTO
SimPO
TIS-DPO
TI-DPO
AAO
GAW-PO (Ours)
AIME 2024
5.73 ± 0.55
6.28 ± 0.80
6.25 ± 0.21
5.66 ± 0.97
6.63 ± 0.33
6.46 ± 0.81
6.53 ± 0.51
6.67 ± 0.38
7.36 ± 0.52
AIME 2025
6.88 ± 1.04
8.33 ± 1.00
7.57 ± 0.66
7.81 ± 0.65
7.95 ± 0.69
8.47 ± 1.16
8.44 ± 0.99
7.88 ± 0.99
7.95 ± 0.42
MATH
67.70 ± 0.27
68.69 ± 0.38
68.07 ± 0.73
67.77 ± 0.47
68.61 ± 0.81
69.41 ± 0.36
69.30 ± 0.47
68.59 ± 0.37
70.35 ± 0.24
Omega 500
11.53 ± 0.76
12.27 ± 1.03
12.47 ± 0.95
11.87 ± 1.85
11.93 ± 0.95
12.33 ± 0.92
12.13 ± 0.76
11.93 ± 0.99
11.60 ± 1.31
AGI Eval
58.51 ± 0.58
60.31 ± 1.08
59.45 ± 0.62
59.46 ± 0.59
58.47 ± 1.40
60.12 ± 0.62
59.75 ± 0.34
59.14 ± 0.29
61.15 ± 0.74
BBH (CoT)
44.40 ± 0.77
45.45 ± 0.55
45.37 ± 0.85
45.12 ± 0.09
45.50 ± 0.64
45.33 ± 0.11
45.13 ± 0.82
44.87 ± 0.69
49.81 ± 0.21
Table 1: Performance comparison across evaluation benchmarks: mean performance (%) across 3 runs, standard deviation shown in gray. Best results per benchmark are bolded .
Figure 2: Average benchmark performance for different values of the DPO regularization parameter β . Standard DPO exhibits a narrow optimum around β=0.02 and rapidly degrades under more aggressive optimization, whereas our method remains stable and continues to improve as β decreases.
Metric
Mean
K =2
K =4
K =8
Average
47.82
46.60
47.94
47.89
Table 2: Average performance for paired+global and clustered preferred directions at β=0.002 . We report mean performance across 3 runs. Full per-benchmark results are in Appendix E.2 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
(a) Biomedical relation extraction
Prompt
Extract every drug combination in the sentence “Reduction of ASS1 expression by siRNA significantly sensitized mesothelioma spheroids to the pro-apoptotic effects of bortezomib and of cisplatin plus pemetrexed ” , copying drug names verbatim and labelling each pos , neg or comb . “In your output, return only the json array and no other text.”
‘‘‘ json [ [" ASS 1 ", " si RNA ", " NEG "], [" b ort ez om ib ", " cis pl atin ", " p emet rex ed ", " POS "] ] ‘‘‘ <|endoftext|>
Errors
The ‘‘‘ fence breaks the instruction to return only the array; ASS1 is a gene and siRNA a technique, so the first combination is fabricated; and the genuine combination is labelled pos rather than comb .
(b) Structured numerical answer
Prompt
Flights of 8, 4 and 6 hours with a 2-hour delay at each border. Report the total travel time as JSON.
Appendix
Table 3: Per-token weights assigned by our method on three rejected completions. Shading is monotone in the weight a token receives (pale = low, deep = high). The weight concentrates on the tokens listed under Errors , while correctly transcribed content receives near-uniform low weight. *The prompts are summarized and example (c) is truncated for readability.
Benchmark
GAW-PO
Uniform Weights
Random Weights
AIME 2024
16.77 ± 0.79
17.12 ± 0.79
16.74 ± 0.52
AIME 2025
17.53 ± 0.68
16.91 ± 0.37
17.01 ± 0.81
MATH
76.63 ± 0.39
76.71 ± 0.35
76.61 ± 0.36
Omega 500
12.33 ± 0.70
13.20 ± 0.92
13.53 ± 0.81
AGI Eval
65.49 ± 0.24
65.12 ± 0.41
64.26 ± 0.67
BBH (CoT)
67.79 ± 0.49
64.96 ± 0.40
63.96 ± 0.14
Appendix
Table 4: Comparison of gradient-based, uniform, and random rejected-token weights. We report mean performance (%) across 3 runs, with standard deviation shown in gray. Best results per benchmark are bolded .
Figure 3: Performance across β grouped by benchmark category. Standard DPO degrades sharply under aggressive preference optimization across all categories, while our method remains stable and generally improves as β decreases.
Benchmark
Mean
K =2
K =4
K =8
AIME 2024
16.77 ± 0.79
17.67 ± 0.51
17.29 ± 1.08
16.01 ± 0.61
AIME 2025
17.53 ± 0.68
17.36 ± 0.57
17.81 ± 0.28
17.08 ± 0.28
MATH
76.63 ± 0.39
76.81 ± 0.40
76.98 ± 0.17
76.48 ± 0.79
Omega 500
12.33 ± 0.70
14.20 ± 0.53
12.27 ± 0.64
12.93 ± 0.64
AGI Eval
65.49 ± 0.24
66.37 ± 0.71
65.52 ± 0.28
65.71 ± 0.48
BBH (CoT)
67.79 ± 0.49
67.24 ± 0.93
67.26 ± 0.59
67.57 ± 0.77
Appendix
Table 5: Full benchmark results for the single dataset-level mean and clustered preferred-direction variants, all trained at β=0.002 . Values are percentages, mean ± standard deviation over 3 seeds. Bold marks the highest value in each row.
Direct Preference Optimization (DPO) aligns Large Language Models with human preferences without the need for a separate reward model. However, DPO treats all tokens in responses equally, neglecting the differing importance of individual tokens. Existing token-level PO methods compute the token weights using either token-position-based heuristic functions or probability estimates given by a separately trained model, which lacks robustness and incurs extra training cost. In contrast, we propose Token-weighted DPO (TwDPO) -- a novel training objective grounded on token-weighted RL -- and AttentionPO -- an instantiation of TwDPO that uses attention from the LLM itself to estimate token weights. AttentionPO prompts the LLM to serve as a pairwise judge and check where the model attends when comparing the responses. This design makes AttentionPO content-aware, adjusting weights based on response content, and efficient, incurring only two extra forward passes per example. Experiment results show that AttentionPO significantly improves performance on AlpacaEval, MT-Bench, and ArenaHard, surpassing existing Preference Optimization methods.
Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences over full sequences even though generation is driven by per-token decisions. Existing token-level extensions typically decompose a sequence-level Bradley-Terry objective across timesteps, leaving per-prefix (state-wise) optimality implicit. We study how to recover token-level preference optimality using only standard sequence-level pairwise comparisons. We introduce Token-level Bregman Preference Optimization (TBPO), which posits a token-level Bradley-Terry preference model over next-token actions conditioned on the prefix, and derive a Bregman-divergence density-ratio matching objective that generalizes the logistic/DPO loss while preserving the optimal policy induced by the token-level model and maintaining DPO-like simplicity. We introduce two instantiations: TBPO-Q, which explicitly learns a lightweight state baseline, and TBPO-A, which removes the baseline through advantage normalization. Across instruction following, helpfulness/harmlessness, and summarization benchmarks, TBPO improves alignment quality and training stability and increases output diversity relative to strong sequence-level and token-level baselines.
Truong Nguyen, Tien-Phat Nguyen, Linh Ngo Van +3
Hanoi University of Science and Technology · University of Stuttgart · German Research Center for Artificial Intelligence +3
Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset. This leaves a useful signal unused: the response that the reference model itself would generate for the same prompt. We propose Direct Preference Optimization with Penalization (DPOP), a simple extension of DPO that augments the base preference loss with a gated penalty on reference-greedy responses. DPOP activates this penalty only when the current policy still assigns a lower likelihood to the preferred response than to the rejected response. On AlpacaEval 2.0, DPOP improves length-controlled win rate over DPO, SimPO, and AlphaDPO on both Llama-3-8b-it and Gemma-2-9b-it, achieving relative gains of 5.3% and 4.4% over baselines on the two models, respectively. Ablations further show that a SimNPO-style length-normalized penalty is stronger than NPO and token-level unlikelihood in this setting.