Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.
Figures & tables
Algorithm
Group Labels (Tr/Val)
No Extra Training
CMNIST
Waterbirds
CelebA
CivilComments
Worst Acc
Avg Acc
Worst Acc
Avg Acc
Worst Acc
Avg Acc
Worst Acc
Avg Acc
Group DRO
Tr/Val
✓
73.1 ±0.3
74.8 ±0.2
90.6 ±0.2
92.7 ±0.1
89.3 ±1.3
92.6 ±0.3
69.0 ±0.9
89.9 ±0.2
LISA
Tr/Val
✓
73.3 ±0.2
74.0 ±0.1
89.2 ±0.6
91.8 ±0.3
89.3 ±1.1
92.4 ±0.4
72.6 ±0.1
89.2 ±0.9
DFR Tr
Tr/Val
-
59.8 ±0.4
62.1 ±0.2
90.2 ±0.8
97.0 ±0.3
80.7 ±2.4
90.6 ±0.7
58.0 ±1.3
92.0 ±0.1
PDE
Tr/Val
✓
72.6 ±0.7
73.0 ±0.4
90.3 ±0.3
92.4 ±0.8
91.0 ±0.4
92.0 ±0.6
71.5 ±0.5
86.3 ±1.7
Table 1: Worst-group and average accuracy on CMNIST, Waterbirds, CelebA, and CivilComments under the standard evaluation setting. Group Labels indicates whether the method uses subgroup annotations from the training data, the validation data, or both. No Extra Training indicates whether the method operates without additional retraining. Among methods that use only validation group labels, boldface indicates the best performance and underlining denotes the second-best.
Figure 1: Analysis of reference-based sample importance on Waterbirds. (a) PCA projection of feature embeddings for the reference set and noisy training set. The reference set consists of validation samples from the two bias-conflicting minority groups, while the noisy training set contains all four groups with mislabeled samples marked by red outlines. (b) Training samples with observed label “waterbird” are colored by the class-conditioned OT dual potential f∗ computed against the waterbird-on-land reference samples. Brighter colors indicate smaller f∗ values and therefore larger sample importance. The visual examples, from top to bottom, show a mislabeled sample, a strongly bias-aligned sample, a bias-aligned sample containing both land and water background regions, and a bias-conflicting sample.
Dataset
Method
Group Labels (Tr/Val)
No Extra Training
Label Noise (%)
0
10
20
30
Waterbirds
Group DRO
Tr/Val
✓
90.6 ±0.2
72.9 ±1.6
54.3 ±1.0
52.2 ±3.6
AFR
Val
-
88.3 ±0.9
58.7 ±0.4
61.2 ±5.9
52.9 ±0.0
KNN-RAD
Val
-
91.0 ±0.1
82.4 ±0.7
74.7 ±1.1
68.9 ±2.1
POTER (Ours)
Val
✓
90.9 ±0.3
89.4 ±0.4
87.9 ±1.6
85.3 ±1.4
CelebA
Group DRO
Tr/Val
✓
89.3 ±1.3
66.9 ±0.3
59.5 ±2.6
54.8 ±1.8
Table 2: Worst-group accuracy on Waterbirds and CelebA under different symmetric label noise rates.
Figure 2: Worst-group accuracy under subgroup-concentrated label noise on Waterbirds and CelebA. Label noise is concentrated on the minority subgroups: (waterbird,land background) and (landbird,water background) for Waterbirds, and (blond,male) for CelebA. The y-axis range is adjusted for each dataset to highlight within-dataset performance differences.
Figure 3: Distribution of class-standardized OT dual potentials f∗ by group on Waterbirds and CelebA in the standard setting. Within each class, lower potentials favor larger sample weights. Bias-conflicting groups (orange) have lower median potentials than their bias-aligned counterparts (light blue).
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Example images from the CMNIST dataset. The groups are g1={0,green} , g2={1,green} , g3={0,red} , and g4={1,red} .
Figure 5: Example images from the Waterbirds dataset. The groups are g1={landbird, land} , g2={landbird, water} , g3={waterbird, land} , and g4={waterbird, water} .
Figure 6: Example images from the CelebA dataset. The groups are g1={non-blond hair, female} , g2={non-blond hair, male} , g3={blond hair, female} , and g4={blond hair, male} .
Identity
Train
Validation
Y=0
Y=1
Y=0
Y=1
male
25,373
4,437
4,050
715
female
31,282
4,962
5,120
771
LGBTQ
6,155
2,265
1,099
358
Christian
24,292
2,446
4,166
384
Muslim
10,829
3,125
1,598
512
Appendix
Table 4: Per-group sample counts in CivilComments-WILDS for each of the eight identity attributes. Counts are obtained by thresholding both the toxicity score and the identity-attribute annotation at 0.5 . Groups are overlapping: a comment that mentions multiple identities is counted in multiple rows.
Dataset
Optimizer
Learning Rate
Weight Decay
Batch Size
Epochs
CMNIST
SGD
10−2
10−2
128
300
Waterbirds
SGD
10−5
1.0
128
300
CelebA
SGD
10−5
10−2
128
50
CivilComments
AdamW
10−5
10−2
32
3
Appendix
Table 5: Training configurations for the final reweighted ERM stage.
Dataset
Method
Label Noise (%)
10
20
30
40
50
60
Waterbirds
KNN-RAD
88.1 ±1.1
77.5 ±4.3
72.6 ±3.8
52.9 ±5.8
17.9 ±2.2
7.5 ±0.2
POTER (Ours)
90.0 ±0.1
89.0 ±0.4
86.4 ±0.4
84.8 ±0.3
82.4 ±0.3
78.3 ±0.9
CelebA
KNN-RAD
79.4 ±0.0
76.9 ±0.8
73.3 ±0.6
70.4 ±0.6
73.0 ±0.3
63.7 ±0.3
POTER (Ours)
90.6 ±0.5
88.0 ±0.9
87.6 ±1.2
87.2 ±1.5
87.2 ±1.5
84.8 ±1.4
Appendix
Table 6: Worst-group accuracy on Waterbirds and CelebA under different subgroup-concentrated label noise rates.
Dataset
Method
Group Labels (Tr/Val)
No Extra Training
Label Noise (%)
0
10
20
30
CivilComments
Group DRO
Tr/Val
✓
69.0 ±0.9
57.6 ±3.0
57.2 ±3.6
50.1 ±6.8
AFR
Val
-
66.2 ±1.8
64.6 ±2.1
65.4 ±1.4
63.3 ±0.4
KNN-RAD
Val
-
69.6 ±0.0
65.0 ±1.0
62.7 ±1.4
62.9 ±2.8
POTER (Ours)
Val
✓
71.0 ±0.6
68.4 ±1.4
64.2 ±0.9
63.1 ±1.7
Appendix
Table 7: Worst-group accuracy on CivilComments under different symmetric label noise rates.
Dataset
Method
Group Labels (Tr/Val)
Extra Training
Label Noise (%)
0
10
20
30
Waterbirds
Group DRO
Tr/Val
Not required
90.6 ±0.2
72.9 ±1.6
54.3 ±1.0
52.2 ±3.6
JTT
Val
Full
84.6 ±3.0
56.5 ±8.0
6.0 ±3.0
2.7 ±1.0
AFR
Val
Last
88.3 ±0.9
58.7 ±0.4
61.2 ±5.9
52.9 ±0.0
END
Val
Full
82.8 ±1.0
84.2 ±1.0
83.2 ±1.0
81.8 ±1.0
KNN-RAD
Val
Last
91.0 ±0.1
82.4 ±0.7
74.7 ±1.1
68.9 ±2.1
Appendix
Table 8: Extended comparison of worst-group accuracy on Waterbirds and CelebA under symmetric label noise. Extra Training indicates the additional retraining required after base model training: Last for last-layer retraining, Full for full-model retraining, and Not required for no additional retraining.
Bias-conflicting training ratio
20%
10%
5%
1%
POTER
98.4±0.2
98.0±0.2
97.3±0.3
95.1±0.2
Appendix
Table 9: Worst-group accuracy (%) on CMNIST under stronger spurious correlation. Label noise is removed throughout the dataset for this experiment.
Reference size
100%
25%
10%
POTER
90.9±0.3
91.1±0.1
89.8±0.1
Appendix
Table 10: Worst-group accuracy (%) on Waterbirds when subsampling the reference set. The full reference contains 599 examples; all other settings are unchanged.
Reference-label noise rate
0%
10%
20%
30%
POTER
90.9±0.3
87.9±1.2
85.7±0.3
82.5±1.9
Appendix
Table 11: Worst-group accuracy (%) on Waterbirds under symmetric corruption of reference class labels. At 30% nominal noise, the observed-waterbird reference has approximately 60% contamination due to class imbalance.
Variant
Reference
Selection metric
Worst-group (%)
Standard POTER
Group-selected subset
Worst-group accuracy
90.9±0.3
Without group annotations
Entire validation set
Mean accuracy
89.3±0.5
Appendix
Table 12: POTER with and without validation group annotations on Waterbirds.
Extractor
g1
g2
g3
g4
Worst-group (%)
Average (%)
Pretrained ResNet-50
0.31
6.32
10.19
1.87
90.9±0.3
92.7±0.1
ERM-trained ResNet-50
0.37
7.79
12.43
1.30
91.0±0.1
91.3±0.0
SigLIP 2 ViT-B/16
0.25
6.89
11.24
1.92
87.9±0.1
95.2±0.2
Appendix
Table 13: Feature-extractor sensitivity on Waterbirds without extractor-specific retuning. Columns g1 – g4 report the mean sample weight for each group defined in Appendix C ; g2 and g3 are bias-conflicting groups.
Weighting method
g1
g2
g3
g4
Worst-group (%)
Class-conditioned k NN weighting
0.66
1.24
3.84
1.94
76.7±0.6
OT-based weighting (POTER)
0.31
6.32
10.19
1.87
90.9±0.3
Appendix
Table 14: Comparison of OT-based weighting and class-conditioned k NN distance weighting on Waterbirds.
Dataset
ntrain
nref
Dimension
Time (s)
CMNIST
30,000
5,022
2,048
0.41±0.00
Waterbirds
4,795
599
2,048
0.07±0.00
CelebA
162,770
8,717
2,048
4.52±0.03
CivilComments
269,038
7,111
768
4.31±0.00
Appendix
Table 15: Runtime for computing OT dual potentials f∗ after feature embeddings have been extracted. Results are averaged over five runs.
Method
Timed operation
Time (s)
Worst-group (%)
SELF
Last-layer retraining
5.657
14.7±0.2
POTER
Complete weight construction
0.075
87.9±1.6
Appendix
Table 16: Comparison of POTER and SELF on Waterbirds: additional computation and robustness to label noise. Worst-group accuracy is measured under 20% symmetric label noise.