We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference alignment. Casting the bilevel optimization problem as a Stackelberg game enables solving the original a posteriori MOO problem in a single forward-backward pass. As a result, MOSEL matches the time and memory efficiency of standard single-objective training while enabling scalable Pareto stationary front learning. Empirically, MOSEL uncovers diverse and optimal Pareto frontiers in strongly conflicting settings (e.g., fairness-accuracy). Remarkably, even in weakly conflicting regimes such as multi-task learning, it consistently converges to solutions closer to the utopia point, outperforming both standard single-objective training and specialized multi-task learning methods. These results highlight the broader potential of a posteriori MOO learning as a pathway to efficiently learn more diverse and robust representations, ultimately improving generalization.
Figures & tables
Figure 1: Overview of MOSEL. (a) MOSEL partitions the NN into a representation learning network, which steers z toward Pareto optimal solutions, and a preference alignment network conditioning on r∼Dir(α) to cover diverse trade-offs. (b) Illustration on ResNet-34 [ 20 ] : z is extracted at the network split point and modulated via a FiLM layer [ 50 ] , which maps r to a scaling and shift factor per feature channel, steering the final predictions y^(x;r) toward the desired objective trade-off.
Dataset
Method
HV Loss ↑
HV Error ↑
# PO sol. (Loss) ↑
# PO sol. (Error) ↑
Adult
MOSEL-LS
0.939 ± 0.007
0.849 ± 0.012
22.8 ± 1.2
21.2 ± 1.5
MOSEL-CS
0.819 ± 0.017
0.776 ± 0.015
23.8 ± 0.7
22.0 ± 1.4
COSMOS
0.926 ± 0.014
0.830 ± 0.010
20.2 ± 0.7
17.4 ± 1.9
Classifier
0.906 ± 0.025
0.618 ± 0.016
8.0 ± 0.6
6.2 ± 0.7
COMPAS
MOSEL-LS
0.920 ± 0.019
0.568 ± 0.008
15.6 ± 1.4
14.6 ± 1.4
MOSEL-CS
0.930 ± 0.007
0.560 ± 0.005
15.2 ± 1.6
14.2 ± 1.6
Table 1: Fairness-Performance Trade-offs. Relative hypervolume and number of Pareto optimal solutions on the fairness benchmarks, computed on training loss and test error. Mean ± std over five runs; best entries per column shaded.
Figure 2: Fairness-Performance Trade-offs. Test-time Pareto fronts, plotting classification error against fairness violation. The grey dotted lines indicate zero unfairness and the best accuracy achievable by an unconstrained classifier; their intersection defines the utopia point—an ideal, yet unattainable solution toward which all methods aim to converge.
Figure 3: Two-task setting. Test-time Pareto fronts on the error, showing the trade-offs between task-specific errors. Grey dotted lines indicate the best achievable single-task performance on each task; their intersection marks the single-task baseline. In this weakly conflicting setting, positive transfer allows methods to surpass this baseline, with MOSEL achieving the largest improvements.
Method
Δ Param.
Time ( × )
CPU ( × )
Equal weight
17,495,136
248.4±45.7 min
11.15±0.25 MB
MOSEL-LS
+672
×1.0±0.1
×0.97±0.02
MOSEL-CS
+672
×1.0±0.2
×0.98±0.02
COSMOS
+1,072
×1.1±0.1
×0.99±0.02
PCGrad
+0
×9.8±2.2
×1.06±0.01
MGDA
+0
×3.0±0.5
×0.99±0.03
Table 5
Figure 4: Scalability to number of objectives on CelebA . Per-task accuracy improvement over the single-task baseline (gray column) for the most correlated (left) and least correlated (right) task sets, as the number of jointly optimized tasks increases. Results are averaged over 5 seeds.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Hypervolume computation
Figure 6: Test-time Pareto fronts on the fairness use case, showing training objectives (BCE loss vs DEO). The grey dotted lines mark zero unfairness and the best BCE loss achievable by an unconstrained classifier; their intersection defines the utopia point, the ideal but infeasible solution, toward which all methods aim to converge.
Dataset
Method
Δ #Param.
Time (rel.)
CPU Mem (rel.)
GPU Mem (rel.)
Adult
MOSEL-LS
+50
1.2±0.2
1.10±0.07
–
MOSEL-CS
+50
1.2±0.1
1.06±0.06
–
COSMOS
+120
1.2±0.2
1.09±0.03
–
Classifier
6,891
44.4±4.4 s
10.30±0.34 MB
–
COMPAS
MOSEL-LS
+50
1.5±0.7
1.00±0.03
–
MOSEL-CS
+50
1.5±0.7
0.96±0.03
–
Appendix
Table 2: Computational overhead on fairness datasets relative to classifier (bottom); # Param. are absolute counts, time and CPU are multiplicative factors. Mean ± std over five runs.
Multi-MNIST
Multi-Fashion
Multi-Fashion+MNIST
CelebA
Method
Acc TL
Acc BR
Acc TL
Acc BR
Acc TL
Acc BR
Acc Goatee
Acc Mustache
MOSEL-LS
0.919 ± 0.001
0.901 ± 0.002
0.829 ± 0.003
0.821 ± 0.001
0.930 ± 0.001
0.851 ± 0.003
0.968 ± 0.002
0.964 ± 0.001
MOSEL-CS
0.910 ± 0.002
0.895 ± 0.002
0.817 ± 0.004
0.812 ± 0.003
0.916 ± 0.003
0.845 ± 0.003
0.970 ± 0.001
0.966 ± 0.001
COSMOS
0.907 ± 0.003
0.886 ± 0.002
0.824 ± 0.004
0.821 ± 0.004
0.915 ± 0.004
0.837 ± 0.005
0.960 ± 0.005
0.959 ± 0.004
PCGrad
0.914 ± 0.006
0.894 ± 0.004
0.814 ± 0.004
0.828 ± 0.001
0.928 ± 0.003
0.845 ± 0.007
0.954 ± 0.000
0.961 ± 0.000
MGDA
0.912 ± 0.009
0.897 ± 0.006
0.822 ± 0.004
0.820 ± 0.002
0.919 ± 0.002
0.852 ± 0.002
0.967 ± 0.001
0.961 ± 0.000
Appendix
Table 3: Per-task accuracy at the solution closest to the utopia point on the test Pareto front, across all MTL benchmarks. As joint training enables positive transfer, we take (0,0) as the utopia point and select the solution minimizing Euclidean distance to it, which can lie beyond the single-task baseline. TL and BR denote top-left and bottom-right tasks on the synthetic benchmarks. Mean ± std over five independent runs; best results per column shaded.
Figure 7: Test-time Pareto fronts on the MTL use case, showing the trade-offs between task-specific (Binary) Cross-Entropy losses. Grey dotted lines indicate the best achievable single-task performance on each task; their intersection marks the single-task baseline. In this weakly conflicting setting, positive transfer allows methods to surpass this baseline, with MOSEL achieving the largest improvements.
Dataset
Method
Δ #Param.
Time (rel.)
CPU Mem (rel.)
GPU Mem (rel.)
Multi-MNIST
MOSEL-LS
+4,320
1.1±0.2
1.15±0.05
1.01±0.00
MOSEL-CS
+4,320
1.2±0.2
1.09±0.09
1.00±0.00
COSMOS
+708
1.3±0.2
1.07±0.02
1.09±0.03
PCGrad ∗
+0
6.2±0.6
1.22±0.04
1.00±0.00
MGDA
+0
2.0±0.4
1.04±0.06
1.12±0.02
GradNorm
+0
1.9 ± 0.2
1.07 ± 0.05
1.00 ± 0.00
Appendix
Table 4: Computational overhead on MTL datasets relative to equal weight training (bottom); # Param. are absolute counts, time and CPU are multiplicative factors. Mean ± std over five runs.
Dataset
Method
HV Loss ↑
HV Error ↑
# PS sol. (Loss) ↑
# PS sol. (Error) ↑
Multi-MNIST
MOSEL-LS
0.330 ± 0.010
0.441 ± 0.008
10.6 ± 0.8
9.0 ± 0.6
MOSEL-CS
0.318 ± 0.008
0.426 ± 0.009
17.8 ± 0.4
16.8 ± 1.2
COSMOS
0.303 ± 0.008
0.418 ± 0.007
20.0 ± 1.1
19.0 ± 3.0
PCGrad
0.274 ± 0.017
0.425 ± 0.010
2.2 ± 0.7
2.8 ± 0.4
MGDA
0.267 ± 0.026
0.409 ± 0.034
1.0 ± 0.0
1.0 ± 0.0
GradNorm
0.281 ± 0.010
0.410 ± 0.016
1.0 ± 0.0
1.0 ± 0.0
Appendix
Table 5: Relative hypervolume and number of Pareto stationary solutions on the MTL benchmarks, computed on training loss and test error. Mean ± std over five runs; best entries per column shaded.
Dataset
Scalarization
α
HV ↑
HV MCR ↑
Obj 1 -best (mean)
Obj 2 -best (mean)
Adult
0.5
0.939 ± 0.007
0.849 ± 0.012
(0.143,0.106)
(0.211,0.009)
LS
0.8
0.912 ± 0.011
0.838 ± 0.019
(0.144,0.105)
(0.202,0.014)
1.2
0.816 ± 0.015
0.774 ± 0.009
(0.143,0.105)
(0.186,0.030)
0.5
0.819 ± 0.017
0.776 ± 0.015
(0.142,0.106)
(0.188,0.030)
CS
0.8
0.702 ± 0.019
0.683 ± 0.022
(0.142,0.105)
(0.172,0.049)
1.2
0.524 ± 0.023
0.525 ± 0.025
(0.142,0.106)
(0.153,0.079)
Appendix
Table 6: Ablation on Dirichlet parameter α . Relative hypervolume on two strongly and two weakly conflicting datasets computed on training loss (HV) and test error (HV MCR). Mean ± std over five runs. Mean extreme solutions (Obj 1 -best, Obj 2 -best) are reported for varying α and scalarization.
Dataset
RL / PA
HV ↑
HV MCR ↑
Obj 1 -best
Obj 2 -best
CelebA
3 / 35
0.213
0.396
(0.075,0.080)
(0.076,0.079)
7 / 31
0.238
0.406
(0.070,0.077)
(0.071,0.076)
11 / 27
0.232
0.401
(0.070,0.079)
(0.071,0.078)
17 / 21
0.242
0.415
(0.068,0.079)
(0.071,0.076)
23 / 15
0.242
0.415
(0.068,0.077)
(0.069,0.077)
31 / 7
0.238
0.422
(0.069,0.078)
(0.071,0.076)
Appendix
Table 7: Ablation on split point. The split is represented by the number of allocated blocks to representation learning (RL) and preference alignment (PA) network respectively. We report obtained HV in training and evaluation space as well as extreme points.
Figure 8: Ablation study on CelebA: runtime and peak memory for increasing numbers of objectives, for correlated and uncorrelated features.
Fairness extreme
Accuracy extreme
RL / PA
HV ↑
HV MCR ↑
Acc@MinDEO(%) ↑
MinDEO ↓
DEO@MaxAcc ↓
MaxAcc(%) ↑
5 / 1
0.423
0.479
88.68
0.0000
0.037
92.22
4 / 2
0.419
0.477
88.69
0.0001
0.035
92.17
3 / 3
0.419
0.475
88.70
0.0002
0.036
92.12
Appendix
Table 8: Results on DistilBERT text-classification. Results show HV and extreme points for the front for different split points (splits between transformer blocks)
Epoch
MOSEL-LS Max
Equal Weight Max
Δ Max
0
5.54 ± 0.86
15.89 ± 5.91
-10.35
10
0.52 ± 0.03
0.42 ± 0.04
+0.09
30
0.38 ± 0.02
0.30 ± 0.05
+0.08
50
0.36 ± 0.02
0.28 ± 0.05
+0.08
70
0.37 ± 0.01
0.28 ± 0.05
+0.09
99
0.36 ± 0.02
0.27 ± 0.04
+0.09
Appendix
Table 9: Per-epoch maximum objective loss (mean ± std over 5 seeds). Δ is defined as difference between MOSEL and Equal Weight baseline; negative values indicate lower (better) worst-case loss for MOSEL. After an initial transient, MOSEL maintains a small, constant gap with tighter cross-seed variance, confirming stable optimization throughout training.
Institute of Artificial Intelligence, Beihang University · Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University · Beijing Academy of Artificial Intelligence (BAAI) +1