Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for alignment to uncertain pairwise preferences. Our formulation has a major learner seeking a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting robust objective of the game directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for it. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an O(1/T) average-iteration convergence for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.
Figures & tables
Figure 1 : Experiments under Tabular Games
Evaluation radius ρeval
ρtrain
0.05
0.10
0.20
0.30
0.50
0.01
0.986
0.967
0.927
0.894
0.844
0.02
0.985
0.965
0.925
0.891
0.841
0.05
0.984
0.961
0.918
0.883
0.832
0.10
0.978
0.949
0.901
0.866
0.814
0.15
0.934
0.894
0.844
0.810
0.766
Table 1 : Win rate against NashMD under random kernel perturbations, averaged over N=8000 IMDb prompts with m=8 candidates. Each entry reports WinRate(πR vs πN;ρeval) from ( 24 ). Values above 0.50 indicate the robust policy is preferred under the majority of sampled kernels.
Fine-Tuning Model
ρtrain
Win Rate under Evaluation Model
DeepSeek-R1
Gemma-4
Nemotron-3
Qwen-1.5B
0.01
0.708
0.577
0.492
0.05
0.625
0.615
0.732
Table 2 : Win rate against NashMD under external LLM judges, evaluated on 1000 TL;DR prompts, with a non-BT (PairRM) nominal kernel.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Value
Candidates per prompt m
8
Kernel dimension d
(28)=28
Number of prompts N
8,000
Generator model
flan-t5-large
KL temperature τ
{0.001,0.005,0.01,0.05,0.1}
NashMD step size ηNash
0.30
Appendix
Table 3: Hyperparameters for the IMDb continuation experiments.
ρtrain
ρeval=0.05
0.10
0.20
0.30
0.50
0.005
0.974
0.945
0.897
0.863
0.818
0.01
0.974
0.944
0.897
0.862
0.818
0.02
0.973
0.943
0.896
0.862
0.818
0.05
0.970
0.939
0.893
0.859
0.816
0.10
0.962
0.928
0.884
0.851
0.811
0.15
0.917
0.882
0.840
0.811
0.775
Appendix
Table 4: WinRate(πR vs πN;ρeval) under group-shift kernels at τ=0.01 , averaged over 8,000 prompts. Bold marks the highest value in each column.
τ
best ρtr
ρeval=0.05
0.10
0.20
0.30
0.50
1.0
0.001
0.005
0.972
0.944
0.896
0.859
0.810
0.735
0.005
0.005
0.980
0.956
0.911
0.877
0.826
0.750
0.01
0.005
0.986
0.967
0.929
0.896
0.847
0.768
0.05
0.10
0.996
0.995
0.988
0.977
0.948
0.879
0.1
0.10
0.945
0.931
0.903
0.880
0.841
0.775
Appendix
Table 5: WinRate(πR vs πN;ρeval) at the best ρtrain for each τ (KL-random kernels).
Figure 2 : WinRate(πR vs πN;ρeval) at τ=0.05 , for each training radius ρtrain∈{0.005,0.01,0.02,0.05,0.10,0.15} . Radii ρtrain∈[0.005,0.10] (gray band and mean) are statistically tied at every evaluation radius; ρtrain=0.15 (red) is significantly worse, and increasingly so as ρeval grows.
τ\ρtr
0.005
0.01
0.02
0.05
0.10
0.15
0.001
0.860
0.857
0.850
0.832
0.813
0.780
0.005
0.877
0.874
0.870
0.857
0.831
0.784
0.01
0.896
0.895
0.893
0.883
0.866
0.810
0.05
0.976
0.976
0.976
0.976
0.977
0.972
0.1
0.859
0.860
0.862
0.868
0.880
0.717
Appendix
Table 6: WinRate(πR vs πN;ρeval) at ρeval=0.30 , by (τ,ρtrain) (KL-random kernels). Bold marks the row-wise maximum.
Fine-Tuning Model
ρtrain
Win Rate under Evaluation Model
DeepSeek-R1
Gemma-4
Nemotron-3
Qwen-1.5B
0.01
0.496
0.506
0.764
0.05
0.574
0.593
0.568
Gemma-2B
0.01
0.542
0.518
0.462
0.05
0.640
0.720
0.583
Appendix
Table 7 : Win rate against NashMD under external LLM judges, evaluated on 1000 TL;DR prompts, with a Bradley-Terry nominal kernel (cf. Table 2 , which uses a non-BT kernel).