Learning Pareto Stationary Fronts via Single-Pass Backpropagation
Organizations: Department of Computer Science Saarland University, Germany · Instituto de Computação Universidade Estadual de Campinas, Brazil
Abstract
We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference alignment. Casting the bilevel optimization problem as a Stackelberg game enables solving the original a posteriori MOO problem in a single forward-backward pass. As a result, MOSEL matches the time and memory efficiency of standard single-objective training while enabling scalable Pareto stationary front learning. Empirically, MOSEL uncovers diverse and optimal Pareto frontiers in strongly conflicting settings (e.g., fairness-accuracy). Remarkably, even in weakly conflicting regimes such as multi-task learning, it consistently converges to solutions closer to the utopia point, outperforming both standard single-objective training and specialized multi-task learning methods. These results highlight the broader potential of a posteriori MOO learning as a pathway to efficiently learn more diverse and robust representations, ultimately improving generalization.
Figures & tables
| Dataset | Method | HV Loss | HV Error | # PO sol. (Loss) | # PO sol. (Error) |
| Adult | MOSEL-LS | 0.939 0.007 | 0.849 0.012 | 22.8 1.2 | 21.2 1.5 |
| MOSEL-CS | 0.819 0.017 | 0.776 0.015 | 23.8 0.7 | 22.0 1.4 | |
| COSMOS | 0.926 0.014 | 0.830 0.010 | 20.2 0.7 | 17.4 1.9 | |
| Classifier | 0.906 0.025 | 0.618 0.016 | 8.0 0.6 | 6.2 0.7 | |
| COMPAS | MOSEL-LS | 0.920 0.019 | 0.568 0.008 | 15.6 1.4 | 14.6 1.4 |
| MOSEL-CS | 0.930 0.007 | 0.560 0.005 | 15.2 1.6 | 14.2 1.6 |
| Method | Param. | Time ( ) | CPU ( ) |
| Equal weight | 17,495,136 | min | MB |
| MOSEL-LS | +672 | ||
| MOSEL-CS | +672 | ||
| COSMOS | +1,072 | ||
| PCGrad | +0 | ||
| MGDA | +0 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Method | #Param. | Time (rel.) | CPU Mem (rel.) | GPU Mem (rel.) |
| Adult | MOSEL-LS | +50 | – | ||
| MOSEL-CS | +50 | – | |||
| COSMOS | +120 | – | |||
| Classifier | 6,891 | s | MB | – | |
| COMPAS | MOSEL-LS | +50 | – | ||
| MOSEL-CS | +50 | – |
| Multi-MNIST | Multi-Fashion | Multi-Fashion+MNIST | CelebA | |||||
| Method | Acc TL | Acc BR | Acc TL | Acc BR | Acc TL | Acc BR | Acc Goatee | Acc Mustache |
| MOSEL-LS | 0.919 0.001 | 0.901 0.002 | 0.829 0.003 | 0.821 0.001 | 0.930 0.001 | 0.851 0.003 | 0.968 0.002 | 0.964 0.001 |
| MOSEL-CS | 0.910 0.002 | 0.895 0.002 | 0.817 0.004 | 0.812 0.003 | 0.916 0.003 | 0.845 0.003 | 0.970 0.001 | 0.966 0.001 |
| COSMOS | 0.907 0.003 | 0.886 0.002 | 0.824 0.004 | 0.821 0.004 | 0.915 0.004 | 0.837 0.005 | 0.960 0.005 | 0.959 0.004 |
| PCGrad | 0.914 0.006 | 0.894 0.004 | 0.814 0.004 | 0.828 0.001 | 0.928 0.003 | 0.845 0.007 | 0.954 0.000 | 0.961 0.000 |
| MGDA | 0.912 0.009 | 0.897 0.006 | 0.822 0.004 | 0.820 0.002 | 0.919 0.002 | 0.852 0.002 | 0.967 0.001 | 0.961 0.000 |
| Dataset | Method | #Param. | Time (rel.) | CPU Mem (rel.) | GPU Mem (rel.) |
| Multi-MNIST | MOSEL-LS | +4,320 | |||
| MOSEL-CS | +4,320 | ||||
| COSMOS | +708 | ||||
| PCGrad ∗ | +0 | ||||
| MGDA | +0 | ||||
| GradNorm | +0 | 1.9 0.2 | 1.07 0.05 | 1.00 0.00 |
| Dataset | Method | HV Loss | HV Error | # PS sol. (Loss) | # PS sol. (Error) |
| Multi-MNIST | MOSEL-LS | 0.330 0.010 | 0.441 0.008 | 10.6 0.8 | 9.0 0.6 |
| MOSEL-CS | 0.318 0.008 | 0.426 0.009 | 17.8 0.4 | 16.8 1.2 | |
| COSMOS | 0.303 0.008 | 0.418 0.007 | 20.0 1.1 | 19.0 3.0 | |
| PCGrad | 0.274 0.017 | 0.425 0.010 | 2.2 0.7 | 2.8 0.4 | |
| MGDA | 0.267 0.026 | 0.409 0.034 | 1.0 0.0 | 1.0 0.0 | |
| GradNorm | 0.281 0.010 | 0.410 0.016 | 1.0 0.0 | 1.0 0.0 |
| Dataset | Scalarization | HV | HV MCR | Obj 1 -best (mean) | Obj 2 -best (mean) | |
| Adult | 0.5 | 0.939 0.007 | 0.849 0.012 | (0.143,0.106) | (0.211,0.009) | |
| LS | 0.8 | 0.912 0.011 | 0.838 0.019 | (0.144,0.105) | (0.202,0.014) | |
| 1.2 | 0.816 0.015 | 0.774 0.009 | (0.143,0.105) | (0.186,0.030) | ||
| 0.5 | 0.819 0.017 | 0.776 0.015 | (0.142,0.106) | (0.188,0.030) | ||
| CS | 0.8 | 0.702 0.019 | 0.683 0.022 | (0.142,0.105) | (0.172,0.049) | |
| 1.2 | 0.524 0.023 | 0.525 0.025 | (0.142,0.106) | (0.153,0.079) |
| Dataset | RL / PA | HV | HV MCR | Obj 1 -best | Obj 2 -best |
| CelebA | 3 / 35 | 0.213 | 0.396 | (0.075,0.080) | (0.076,0.079) |
| 7 / 31 | 0.238 | 0.406 | (0.070,0.077) | (0.071,0.076) | |
| 11 / 27 | 0.232 | 0.401 | (0.070,0.079) | (0.071,0.078) | |
| 17 / 21 | 0.242 | 0.415 | (0.068,0.079) | (0.071,0.076) | |
| 23 / 15 | 0.242 | 0.415 | (0.068,0.077) | (0.069,0.077) | |
| 31 / 7 | 0.238 | 0.422 | (0.069,0.078) | (0.071,0.076) |
| Fairness extreme | Accuracy extreme | ||||||
| RL / PA | HV | HV MCR | Acc@MinDEO(%) | MinDEO | DEO@MaxAcc | MaxAcc(%) | |
| 5 / 1 | 0.423 | 0.479 | 88.68 | 0.0000 | 0.037 | 92.22 | |
| 4 / 2 | 0.419 | 0.477 | 88.69 | 0.0001 | 0.035 | 92.17 | |
| 3 / 3 | 0.419 | 0.475 | 88.70 | 0.0002 | 0.036 | 92.12 | |
| Epoch | MOSEL-LS Max | Equal Weight Max | Max |
| 0 | 5.54 0.86 | 15.89 5.91 | -10.35 |
| 10 | 0.52 0.03 | 0.42 0.04 | +0.09 |
| 30 | 0.38 0.02 | 0.30 0.05 | +0.08 |
| 50 | 0.36 0.02 | 0.28 0.05 | +0.08 |
| 70 | 0.37 0.01 | 0.28 0.05 | +0.09 |
| 99 | 0.36 0.02 | 0.27 0.04 | +0.09 |