MORE-PLR: multi-output regression employed for partial label ranking
Authors: Santo M. A. R. Thies, Juan C. Alfaro, Viktor Bengs
Organizations: Chair of Artificial Intelligence and Machine Learning, Ludwig-Maximilians-Universität München, Munich, 80799, Germany. · Departamento de Sistemas Informáticos, Universidad de Castilla-La Mancha, Albacete, 02071, Spain. · Laboratorio de Sistemas Inteligentes y Minería de Datos, Universidad de Castilla-La Mancha, Albacete, 02071, Spain. · Munich Center for Machine Learning, Munich, Germany.
The partial label ranking problem is a supervised learning scenario that aims to fit a preference model that predicts a bucket order defined over a set of labels for a given input instance. This problem generalizes the well-known label ranking problem, which, in practice, is limited to outputting total orders of labels. Existing partial label ranking methods have primarily extended label ranking approaches to handle ties in predictions. This paper proposes using multi-output regression to address the partial label ranking problem, introducing an encoder that, during the learning phase, transforms the (possibly incomplete) rankings with ties of labels to multivariate regression targets, an underexplored perspective in both label ranking and partial label ranking. Moreover, during the inference phase, we introduce several post-hoc layers that convert the multi-output regression results into the output bucket order to effectively implement this approach. This framework provides learning strategies that are competitive with the current state-of-the-art partial label ranking methods, as demonstrated through experimental evaluations.
Figures & tables
Figure 1 : Given a prompt, the LLMgenerates multiple responses, which are ranked by a human labeler from worst to best. This ranking is then used to fine-tune the LLM.
Figure 2 : In this example, the goal is to predict a bucket order of vacation destinations for a customer based on their demographic information.
Figure 3 : Impact of encoding variants on algorithms’ averaged accuracy across all datasets
Figure 4 : Impact of q in the accuracy of the algorithms
Figure 5 : Impact of ϵ on the accuracy of the algorithms
LR
Missing percentage: 0%
Friedman p-value: 5.136×10−7
Holm results
Method
Rank
P-value
Win
Tie
Loss
ST-PI
2.27
-
-
-
-
Native-PI
4.42
1.146×10−1
10
0
3
Table 1 : Friedman’s and Holm’s tests for accuracy with varying missing percentage
Problem
Dataset
RPC
ST-RR
ST-PI
ST-EPS
Chain-RR
Chain-PI
Chain-EPS
Native-RR
Native-PI
Native-EPS
LR
authorship
0.637 ± 0.141
0.825 ± 0.137
0.872 ± 0.299
0.915 ± 0.147
1.420 ± 0.008
1.396 ± 0.009
1.416 ± 0.007
0.537 ± 0.014
0.532 ± 0.017
0.534 ± 0.013
glass
0.604 ± 0.019
0.593 ± 0.015
0.578 ± 0.015
0.666 ± 0.018
1.207 ± 0.015
1.198 ± 0.010
1.218 ± 0.009
0.462 ± 0.014
0.480 ± 0.013
0.458 ± 0.016
iris
0.429 ± 0.045
0.538 ± 0.013
0.523 ± 0.012
0.536 ± 0.019
0.860 ± 0.008
0.847 ± 0.009
0.865 ± 0.007
0.483 ± 0.007
0.482 ± 0.012
0.475 ± 0.014
letter
7.808 ± 0.042
5.364 ± 0.068
5.203 ± 0.022
5.490 ± 0.040
6.503 ± 0.019
6.552 ± 0.019
6.625 ± 0.034
1.413 ± 0.013
1.262 ± 0.023
1.282 ± 0.011
libras
1.654 ± 0.021
1.568 ± 0.336
1.500 ± 0.009
1.537 ± 0.012
2.445 ± 0.012
2.382 ± 0.017
2.464 ± 0.013
0.570 ± 0.011
0.549 ± 0.009
0.559 ± 0.010
movies
1.573 ± 0.067
0.870 ± 0.010
0.864 ± 0.011
0.929 ± 0.019
2.067 ± 0.025
2.013 ± 0.014
2.060 ± 0.012
0.456 ± 0.012
0.466 ± 0.013
0.454 ± 0.011
Table 2 : Time results (in seconds) for complete rankings
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Problem
Name
Identifier
#of instances
#of features
#of labels
#of rankings
\mbox#ofbuckets
LR
authorship
42834
841
70
4
17
4.0
glass
42847
214
9
6
30
6.0
iris
42851
150
4
3
5
3.0
letter
45727
20000
16
26
17806
26.0
libras
45736
360
90
15
317
15.0
pendigits
42856
10992
16
10
2081
10.0
Appendix
Table 3 : Datasets used for the experiments. The algae, movies, and political datasets correspond to real-world problems.
Problem
Dataset
RPC
ST-RR
ST-PI
ST-EPS
Chain-RR
Chain-PI
Chain-EPS
Native-RR
Native-PI
Native-EPS
LR
authorship
0.887 ± 0.019
0.917 ± 0.018
0.925 ± 0.019
0.923 ± 0.019
0.916 ± 0.020
0.919 ± 0.019
0.919 ± 0.020
0.914 ± 0.019
0.918 ± 0.020
0.917 ± 0.020
glass
0.856 ± 0.047
0.893 ± 0.036
0.901 ± 0.037
0.898 ± 0.036
0.883 ± 0.041
0.890 ± 0.041
0.889 ± 0.042
0.902 ± 0.032
0.907 ± 0.031
0.904 ± 0.030
iris
0.948 ± 0.037
0.962 ± 0.033
0.965 ± 0.031
0.965 ± 0.032
0.963 ± 0.034
0.962 ± 0.036
0.962 ± 0.036
0.968 ± 0.032
0.969 ± 0.030
0.969 ± 0.030
letter
0.902 ± 0.002
0.940 ± 0.001
0.942 ± 0.001
0.703 ± 0.005
0.932 ± 0.001
0.934 ± 0.001
0.675 ± 0.005
0.919 ± 0.001
0.919 ± 0.002
0.686 ± 0.005
libras
0.856 ± 0.017
0.898 ± 0.013
0.905 ± 0.012
0.874 ± 0.012
0.898 ± 0.014
0.904 ± 0.014
0.876 ± 0.014
0.883 ± 0.016
0.889 ± 0.015
0.861 ± 0.014
movies
0.218 ± 0.027
0.359 ± 0.033
0.350 ± 0.032
0.354 ± 0.033
0.364 ± 0.042
0.359 ± 0.041
0.359 ± 0.040
0.370 ± 0.033
0.358 ± 0.033
0.366 ± 0.033
Appendix
Table 4 : Accuracy results for complete rankings
Problem
Dataset
RPC
ST-RR
ST-PI
ST-EPS
Chain-RR
Chain-PI
Chain-EPS
Native-RR
Native-PI
Native-EPS
LR
authorship
0.866 ± 0.025
0.909 ± 0.021
0.919 ± 0.020
0.917 ± 0.019
0.726 ± 0.029
0.807 ± 0.025
0.851 ± 0.025
0.728 ± 0.026
0.817 ± 0.021
0.854 ± 0.023
glass
0.823 ± 0.053
0.885 ± 0.035
0.893 ± 0.039
0.891 ± 0.038
0.602 ± 0.060
0.617 ± 0.060
0.643 ± 0.061
0.625 ± 0.056
0.635 ± 0.059
0.668 ± 0.060
iris
0.927 ± 0.042
0.955 ± 0.042
0.964 ± 0.038
0.963 ± 0.038
0.511 ± 0.113
0.589 ± 0.112
0.599 ± 0.117
0.517 ± 0.099
0.591 ± 0.115
0.616 ± 0.110
letter
0.879 ± 0.002
0.934 ± 0.001
0.936 ± 0.001
0.734 ± 0.005
0.802 ± 0.002
0.650 ± 0.004
0.718 ± 0.002
0.722 ± 0.003
0.611 ± 0.003
0.629 ± 0.003
libras
0.784 ± 0.022
0.878 ± 0.014
0.885 ± 0.015
0.855 ± 0.014
0.690 ± 0.022
0.637 ± 0.024
0.670 ± 0.021
0.690 ± 0.020
0.634 ± 0.025
0.671 ± 0.020
movies
0.206 ± 0.031
0.353 ± 0.037
0.345 ± 0.039
0.346 ± 0.038
0.308 ± 0.040
0.257 ± 0.039
0.303 ± 0.041
0.285 ± 0.034
0.224 ± 0.035
0.280 ± 0.033
Appendix
Table 5 : Accuracy results for incomplete rankings with 30% of missing labels
Problem
Dataset
RPC
ST-RR
ST-PI
ST-EPS
Chain-RR
Chain-PI
Chain-EPS
Native-RR
Native-PI
Native-EPS
LR
authorship
0.814 ± 0.028
0.893 ± 0.021
0.909 ± 0.020
0.909 ± 0.019
0.567 ± 0.034
0.629 ± 0.037
0.722 ± 0.033
0.609 ± 0.026
0.667 ± 0.034
0.777 ± 0.029
glass
0.743 ± 0.058
0.851 ± 0.042
0.863 ± 0.040
0.861 ± 0.043
0.394 ± 0.068
0.389 ± 0.073
0.436 ± 0.078
0.424 ± 0.065
0.416 ± 0.069
0.475 ± 0.068
iris
0.860 ± 0.056
0.938 ± 0.046
0.947 ± 0.048
0.946 ± 0.049
0.326 ± 0.139
0.341 ± 0.134
0.359 ± 0.139
0.321 ± 0.142
0.325 ± 0.135
0.335 ± 0.148
letter
0.834 ± 0.002
0.923 ± 0.001
0.924 ± 0.001
0.761 ± 0.004
0.687 ± 0.003
0.293 ± 0.006
0.573 ± 0.006
0.579 ± 0.003
0.305 ± 0.004
0.462 ± 0.003
libras
0.632 ± 0.027
0.846 ± 0.016
0.850 ± 0.016
0.824 ± 0.015
0.504 ± 0.029
0.358 ± 0.027
0.478 ± 0.029
0.529 ± 0.026
0.335 ± 0.035
0.505 ± 0.027
movies
0.204 ± 0.035
0.344 ± 0.032
0.337 ± 0.030
0.339 ± 0.029
0.165 ± 0.040
0.107 ± 0.030
0.161 ± 0.041
0.179 ± 0.035
0.122 ± 0.028
0.177 ± 0.036
Appendix
Table 6 : Accuracy results for incomplete rankings with 60% of missing labels
Rank estimation under label noise poses a fundamental challenge, as ordinal annotations often exhibit structured uncertainty rather than simple label corruption. In this paper, we reformulate rank estimation with noisy ordinal labels as a stochastic ordering problem, in which each instance is inherently associated with multiple plausible ranks instead of a single deterministic label. Based on this view, we propose stochastic order learning (SOL), a learning framework that captures ordinal label uncertainty and learns an embedding space through two complementary objectives: a discriminative loss that structures instance--centroid interactions and a stochastic order loss that enforces probabilistic ordering relations between instances. Extensive experiments across diverse datasets demonstrate that SOL enables reliable rank estimation under various types and levels of label noise. The source code is available at https://github.com/cwlee00/SOL.
Chaewon Lee, Seon-Ho Lee, Chang-Su Kim
School of Electrical Engineering, Korea University, Seoul, Korea · Amazon AGI, Seattle, USA
Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively studied for classification and regression, calibration has not been formally addressed for probabilistic label ranking, where the goal is to predict a distribution over orderings of a label set. Naively treating rankings as classes ignores their structure and fails to capture important modalities such as pairwise and top-k predictions. We formalize calibration for label ranking and develop a hierarchy of notions covering full rankings, sub-rankings, and top-k rankings. We prove that full-rank calibration implies the others but not conversely, and sub-ranking and top-k calibration are incomparable. Empirically, we find popular label ranking models are often poorly calibrated, with substantial differences between sub-ranking and top-k metrics. Applying our framework to RLHF reward models, we find that calibration correlates strongly but not perfectly with benchmark accuracy, suggesting it captures a meaningful quality dimension beyond top-1 accuracy. These findings motivate future work on understanding the downstream effects of miscalibration and developing methods to correct it.
Santo M. A. R. Thies, Viktor Bengs, Timo Kaufmann +2
Data Science and its Applications, Deutsches Forschungsinstitut für Künstliche Intelligenz (DFKI), Kaiserslautern, Germany · Munich Center for Machine Learning, Munich (MCML), Germany · Ludwig Maximilian Universität München, Munich, Germany
We study learning a mixture of k Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, estimating a mixture of multi-way ranking models can become theoretically unidentifiable when k exceeds m/2, where m is the ranking length. We design an efficient algorithm to address this issue by first augmenting the rankings to a larger size (e.g., generating comparisons from a base model), followed by a gradient-based estimation to reduce inference cost (in the input embedding space). With this procedure in mind, we then fit a mixture of Plackett-Luce (PL) models via an expectation-maximization-style iteration, or MoPLEx in short. We conduct extensive experiments to verify this algorithm. First, we find that the gradient-based approximation estimates true probabilities with less than 5% error on models with up to 34 billion parameters. Second, MoPLEx improves clustering and ranking accuracy by an average of 43.7% and 15.2% over baselines using a single PL model or a mixture of Bradley-Terry models, on UltraFeedback and PERSONA datasets. These results demonstrate the effectiveness of MoPLEx for tackling multi-way rankings following heterogeneous preferences through measuring alignment via gradients.
Dongyue Li, Ziniu Zhang, Lu Wang +1
Northeastern University, Boston, MA · University of Michigan, Ann Arbor, MI