Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise
Organizations: Pohang University of Science and Technology (POSTECH)
Abstract
Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.
Figures & tables
| Algorithm | Group Labels (Tr/Val) | No Extra Training | CMNIST | Waterbirds | CelebA | CivilComments | ||||
| Worst Acc | Avg Acc | Worst Acc | Avg Acc | Worst Acc | Avg Acc | Worst Acc | Avg Acc | |||
| Group DRO | Tr/Val | ✓ | 73.1 ±0.3 | 74.8 ±0.2 | 90.6 ±0.2 | 92.7 ±0.1 | 89.3 ±1.3 | 92.6 ±0.3 | 69.0 ±0.9 | 89.9 ±0.2 |
| LISA | Tr/Val | ✓ | 73.3 ±0.2 | 74.0 ±0.1 | 89.2 ±0.6 | 91.8 ±0.3 | 89.3 ±1.1 | 92.4 ±0.4 | 72.6 ±0.1 | 89.2 ±0.9 |
| DFR Tr | Tr/Val | - | 59.8 ±0.4 | 62.1 ±0.2 | 90.2 ±0.8 | 97.0 ±0.3 | 80.7 ±2.4 | 90.6 ±0.7 | 58.0 ±1.3 | 92.0 ±0.1 |
| PDE | Tr/Val | ✓ | 72.6 ±0.7 | 73.0 ±0.4 | 90.3 ±0.3 | 92.4 ±0.8 | 91.0 ±0.4 | 92.0 ±0.6 | 71.5 ±0.5 | 86.3 ±1.7 |
| Dataset | Method | Group Labels (Tr/Val) | No Extra Training | Label Noise (%) | |||
| 0 | 10 | 20 | 30 | ||||
| Waterbirds | Group DRO | Tr/Val | ✓ | 90.6 ±0.2 | 72.9 ±1.6 | 54.3 ±1.0 | 52.2 ±3.6 |
| AFR | Val | - | 88.3 ±0.9 | 58.7 ±0.4 | 61.2 ±5.9 | 52.9 ±0.0 | |
| KNN-RAD | Val | - | 91.0 ±0.1 | 82.4 ±0.7 | 74.7 ±1.1 | 68.9 ±2.1 | |
| POTER (Ours) | Val | ✓ | 90.9 ±0.3 | 89.4 ±0.4 | 87.9 ±1.6 | 85.3 ±1.4 | |
| CelebA | Group DRO | Tr/Val | ✓ | 89.3 ±1.3 | 66.9 ±0.3 | 59.5 ±2.6 | 54.8 ±1.8 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Identity | Train | Validation | ||
| male | 25,373 | 4,437 | 4,050 | 715 |
| female | 31,282 | 4,962 | 5,120 | 771 |
| LGBTQ | 6,155 | 2,265 | 1,099 | 358 |
| Christian | 24,292 | 2,446 | 4,166 | 384 |
| Muslim | 10,829 | 3,125 | 1,598 | 512 |
| Dataset | Optimizer | Learning Rate | Weight Decay | Batch Size | Epochs |
| CMNIST | SGD | 128 | 300 | ||
| Waterbirds | SGD | 128 | 300 | ||
| CelebA | SGD | 128 | 50 | ||
| CivilComments | AdamW | 32 | 3 |
| Dataset | Method | Label Noise (%) | |||||
| 10 | 20 | 30 | 40 | 50 | 60 | ||
| Waterbirds | KNN-RAD | 88.1 ±1.1 | 77.5 ±4.3 | 72.6 ±3.8 | 52.9 ±5.8 | 17.9 ±2.2 | 7.5 ±0.2 |
| POTER (Ours) | 90.0 ±0.1 | 89.0 ±0.4 | 86.4 ±0.4 | 84.8 ±0.3 | 82.4 ±0.3 | 78.3 ±0.9 | |
| CelebA | KNN-RAD | 79.4 ±0.0 | 76.9 ±0.8 | 73.3 ±0.6 | 70.4 ±0.6 | 73.0 ±0.3 | 63.7 ±0.3 |
| POTER (Ours) | 90.6 ±0.5 | 88.0 ±0.9 | 87.6 ±1.2 | 87.2 ±1.5 | 87.2 ±1.5 | 84.8 ±1.4 | |
| Dataset | Method | Group Labels (Tr/Val) | No Extra Training | Label Noise (%) | |||
| 0 | 10 | 20 | 30 | ||||
| CivilComments | Group DRO | Tr/Val | ✓ | 69.0 ±0.9 | 57.6 ±3.0 | 57.2 ±3.6 | 50.1 ±6.8 |
| AFR | Val | - | 66.2 ±1.8 | 64.6 ±2.1 | 65.4 ±1.4 | 63.3 ±0.4 | |
| KNN-RAD | Val | - | 69.6 ±0.0 | 65.0 ±1.0 | 62.7 ±1.4 | 62.9 ±2.8 | |
| POTER (Ours) | Val | ✓ | 71.0 ±0.6 | 68.4 ±1.4 | 64.2 ±0.9 | 63.1 ±1.7 | |
| Dataset | Method | Group Labels (Tr/Val) | Extra Training | Label Noise (%) | |||
| 0 | 10 | 20 | 30 | ||||
| Waterbirds | Group DRO | Tr/Val | Not required | 90.6 ±0.2 | 72.9 ±1.6 | 54.3 ±1.0 | 52.2 ±3.6 |
| JTT | Val | Full | 84.6 ±3.0 | 56.5 ±8.0 | 6.0 ±3.0 | 2.7 ±1.0 | |
| AFR | Val | Last | 88.3 ±0.9 | 58.7 ±0.4 | 61.2 ±5.9 | 52.9 ±0.0 | |
| END | Val | Full | 82.8 ±1.0 | 84.2 ±1.0 | 83.2 ±1.0 | 81.8 ±1.0 | |
| KNN-RAD | Val | Last | 91.0 ±0.1 | 82.4 ±0.7 | 74.7 ±1.1 | 68.9 ±2.1 | |
| Bias-conflicting training ratio | 20% | 10% | 5% | 1% |
| POTER |
| Reference size | 100% | 25% | 10% |
| POTER |
| Reference-label noise rate | 0% | 10% | 20% | 30% |
| POTER |
| Variant | Reference | Selection metric | Worst-group (%) |
| Standard POTER | Group-selected subset | Worst-group accuracy | |
| Without group annotations | Entire validation set | Mean accuracy |
| Extractor | Worst-group (%) | Average (%) | ||||
| Pretrained ResNet-50 | 0.31 | 6.32 | 10.19 | 1.87 | ||
| ERM-trained ResNet-50 | 0.37 | 7.79 | 12.43 | 1.30 | ||
| SigLIP 2 ViT-B/16 | 0.25 | 6.89 | 11.24 | 1.92 |
| Weighting method | Worst-group (%) | ||||
| Class-conditioned NN weighting | 0.66 | 1.24 | 3.84 | 1.94 | |
| OT-based weighting (POTER) | 0.31 | 6.32 | 10.19 | 1.87 |
| Dataset | Dimension | Time (s) | ||
| CMNIST | 30,000 | 5,022 | 2,048 | |
| Waterbirds | 4,795 | 599 | 2,048 | |
| CelebA | 162,770 | 8,717 | 2,048 | |
| CivilComments | 269,038 | 7,111 | 768 |
| Method | Timed operation | Time (s) | Worst-group (%) |
| SELF | Last-layer retraining | 5.657 | |
| POTER | Complete weight construction | 0.075 |