Multivariate ordinal data along with covariates are commonly collected in problems ranging from alignment of language models with human preferences, as well as in recommender systems. For example, data sets such as MovieLens contain several movies rated on a scale 1--5 by human users, along with their demographic information such as age or gender. Similarly, data sets such as HelpSteer collect human feedback on several attributes such as "helpfulness" or "verbosity" of LLM response on an ordinal scale, with covariates depending on the LLM prompt--response pairs. Unfortunately, the standard approaches for modeling these data (a) look at the attributes individually rather than jointly, and (b) often convert the data into pairwise or list-wise win--loss comparisons for fitting models such as Bradley--Terry and Plackett--Luce. Both of these lead to a coarsening of what is actually observed, which we address via a joint covariate-dependent consecutive ratio Markov random field model. We also show pairwise or listwise comparison models are obtained under restrictions of our joint model, and that joint modeling improves comparisons. We also develop a maximum likelihood inference procedure even in the presence of an intractable normalizer.
Figures & tables
M=4
M=1
RMSE
Brier
RMSE
Brier
CR–MRF
0.028(0.002)
0.1449(0.0001)
0.028(0.002)
0.1215(0.0001)
MPLE
0.044(0.004)
0.1462(0.0004)
0.034(0.003)
0.1219(0.0004)
PL
0.037(0.002)
0.1455(0.0001)
0.039(0.002)
0.1222(0.0001)
BTD
0.031(0.002)
0.1451(0.0001)
0.030(0.001)
0.1216(0.0001)
Oracle
—
0.1441
—
0.1207
Table 1: Mean (sd) of RMSE and Brier score of the fitted pairwise preference probabilities P^jk(xi) over 10 replicates. Oracle is based on the true parameters.
Figure 1: Observed and model-implied pairwise preference probabilities on the 64 held-out MovieLens reviewers, split by reviewer gender, with ordinal ratings ( M=4 ).
MSE
Brier score
Size of Dc
Full CR–MRF
Ind. CR–MRF
Full CR–MRF
BTD
Weak PL
n=500
∣Dc∣ = 1
0.92
1.28
0.57
0.61
0.59
∣Dc∣ = 2
0.96
1.55
0.62
0.66
0.64
∣Dc∣ = 3
1.04
1.92
0.72
0.78
0.75
n=1000
∣Dc∣ = 1
0.75
1.13
0.52
0.57
0.58
∣Dc∣ = 2
0.82
1.49
0.60
0.63
0.62
Table S1: Mean square error (MSE) in predicting missing values in test data where for each test data the covariate information and a subset of the ordinal variables are observed and the missing values are predicted conditional on these. Brier scores for out-of-sample pairwise preference predictions are also reported.
Homogeneous ( M=4 )
Covariate ( M=2 )
p
Method
Node
Edge
Total
Node
Edge
Total
2
MLE
0.09
0.03
0.04
0.39
0.18
0.28
2
VGAM
0.09
0.03
0.04
0.40
0.21
0.29
3
MLE
0.16
0.04
0.05
0.61
0.81
0.72
3
VGAM
0.16
0.04
0.05
0.69
1.21
0.96
5
MLE
0.18
0.19
0.21
0.73
0.88
0.84
Table S2: Normalized Frobenius losses over 10 replicates with n=500 . Left: homogeneous CR–MRF ( M=4 ). Right: covariate-dependent CR–MRF ( M=2 , q=2 ). Lower errors are boldened.
Loss
Support recovery
p
Method
Node
Edge
Total
Prec.
Rec.
MCC
20
MLE
0.57
0.30
0.33
0.76
0.77
0.74
20
ordinalNet
1.90
0.29
0.50
0.40
0.97
0.58
30
MLE
0.71
0.39
0.43
0.84
0.51
0.63
30
ordinalNet
2.62
0.38
0.61
0.54
0.83
0.63
40
MLE
0.74
0.36
0.40
0.66
0.73
0.67
Table S3: High-dimensional homogeneous CR–MRF: normalized Frobenius losses and edge-support recovery under stability selection. Means over 10 replicates, n=500 , M=4 .
Female
Age
Movie
MLE
MPLE
MLE
MPLE
When Harry Met Sally…
−0.58
−0.61
0.06
−0.36
Airplane!
0.66
0.69
0.04
−0.03
Desperately Seeking Susan
−0.27
−0.55
−0.36
−0.58
Ferris Bueller’s Day Off
−0.03
−1.24
0.22
0.31
The Breakfast Club
−0.29
−0.73
0.22
0.13
Table S4: Estimated node slopes βj for MovieLens 1M under the joint MLE and the node-wise MPLE. Age is standardised.
Figure S1: Fitted edge network θjk(x) under the joint MLE, for men and women at ages 21, 29.5 and 47. Red edges are positive and blue edges negative. Line width is proportional to ∣θjk(x)∣ on a scale shared by all six panels.
Figure S2: Fitted edge network θjk(x) under the joint MPLE, for men and women at ages 21, 29.5 and 47. Red edges are positive and blue edges negative. Line width is proportional to ∣θjk(x)∣ on a scale shared by all six panels.
Figure S3: Observed and model-implied pairwise preference probabilities on the 64 held-out MovieLens reviewers, split by reviewer gender, with ordinal ratings ( M=4 ) and ties resolved randomly for fitting strict PL.
Figure S4: Observed and model-implied pairwise preference probabilities on the 64 held-out MovieLens reviewers, split by reviewer gender, with binary ratings ( M=1 ).
Pairwise preference data is widely used in language-model evaluation and alignment, often for model ranking, reward modeling, or preference optimization. This note formulates a more basic measurement question: given a reference distribution of pairwise preferences, what model-level quantity is estimated when we test whether a model ranks preferred responses above rejected responses? We define pairwise reference alignment as an ordinal observable induced by a model scoring function. Given a reference pair distribution Ppair over triples (x,y+,y−), and a scalar model score SM(x,y), we define the alignment observable as the probability that the model-induced ordering agrees with the reference preference ordering. We further define a centered order-parameter-like statistic and discuss a margin-based extension. The resulting quantities admit simple finite-sample estimators and concentration bounds under independent sampling assumptions. This note does not introduce a new benchmark. It provides a conceptual and statistical formulation for pairwise reference alignment, clarifies the role of the reference pair distribution, and distinguishes the general ordinal observable from scoring choices such as normalized log-probability or energy-based scores. We also provide an initial empirical study on Qwen2.5 models and RewardBench, where the proposed statistics increase with model size and instruction tuning and vary across reference-pair subsets as predicted by the formulation.
In this paper, we study prediction problems for paired comparison data, for example, predicting the win probability between two unmatched players and ranking all the players according to the order of their strengths by using win probability data between two matched players. Paired comparison data are typically analyzed using Bradley-Terry and Thurstone-Mosteller models. These models predict the win probability by transforming the difference between learned rate parameters, which represent players';strengths, with a pre-specified inverse link function, and employ the order of learned rate parameters for player ranking. However, these models may suffer from model misspecification owing to the selection of a fixed inverse link function. Therefore, in this study, we propose to learn the rate parameters by a (sub-)gradient method and the inverse link function by an isotonic regression technique alternately. The proposed model guarantees monotonic improvement in training error, and is likely to yield an exact tie when the available data is insufficient to establish a strict ranking. We also verified that the proposed model could improve the win probability prediction and ranking performance through numerical experiments with synthetic data and real-world data of football Premier League, baseball MLB, and tennis ATP tour.
Ryoya Yamasaki
Hitotsubashi Institute for Advanced Study, Hitotsubashi University, 2-1 Naka, Kunitachi, 186-8601, Tokyo, Japan.
We study learning a mixture of k Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, estimating a mixture of multi-way ranking models can become theoretically unidentifiable when k exceeds m/2, where m is the ranking length. We design an efficient algorithm to address this issue by first augmenting the rankings to a larger size (e.g., generating comparisons from a base model), followed by a gradient-based estimation to reduce inference cost (in the input embedding space). With this procedure in mind, we then fit a mixture of Plackett-Luce (PL) models via an expectation-maximization-style iteration, or MoPLEx in short. We conduct extensive experiments to verify this algorithm. First, we find that the gradient-based approximation estimates true probabilities with less than 5% error on models with up to 34 billion parameters. Second, MoPLEx improves clustering and ranking accuracy by an average of 43.7% and 15.2% over baselines using a single PL model or a mixture of Bradley-Terry models, on UltraFeedback and PERSONA datasets. These results demonstrate the effectiveness of MoPLEx for tackling multi-way rankings following heterogeneous preferences through measuring alignment via gradients.
Dongyue Li, Ziniu Zhang, Lu Wang +1
Northeastern University, Boston, MA · University of Michigan, Ann Arbor, MI