Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · School of Computing and Information Sciences, Saint Francis University, Hong Kong, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China
Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.
Figures & tables
Figure 1: (a) The overall test accuracy of the existing unsupervised method CPL ( Zhang et al., 2024 ) under different imbalance ratios. (b) Comparison of head- and tail-class accuracies across three settings on the RESISC45 dataset: unsupervised learning with balanced data (Unsup. Bal.), supervised learning with imbalanced data (Sup. Imb.), and unsupervised learning with imbalanced data (Unsup. Imb.). (c) Comparison of average logit margin changes ( Δm=marginFine-tuned−marginZero-shot CLIP ). Under the long-tailed setting, the head-class margin significantly decreases compared to the uniform setting, indicating severe boundary erosion. (d) Label counts across three cases: ground-truth (GT), zero-shot predictions (ZS Pred.), and fine-tuned predictions (FT Pred.). We group confusable pairs by color (e.g., blue: “beach” vs. “sea ice”) and observe that zero-shot CLIP inherently exhibits bias toward confusable classes, which is further exacerbated by fine-tuning.
Figure 2: Visualization of attention maps for misassigned samples on RESISC45. We select head-class samples where the baseline method incorrectly assigns tail-class pseudo-labels. (a) The model mainly focuses on the foreground object. (b) The model is distracted by the irrelevant background. (c) The model correctly captures the foreground object while reducing background attention.
Figure 3: (a) BPA preserves head-class accuracy on OxfordPets as the imbalance ratio increases. (b) The complete MARS pipeline maintains stable overall accuracy on OxfordPets across imbalance ratios. (c) Evolution of classification margins during training. Samples with correct pseudo-labels (Correct PL) consistently exhibit increasing margins, whereas mislabeled samples (Wrong PL) fluctuate around zero or remain negative.
Imbalance Ratio = 10
Method
CUB.
Res.
FA.
Cal.
ES.
Flw.
Food.
Pets.
Cars.
Avg.
ZS CLIP
51.82 0.00
54.48 0.00
17.58 0.00
90.30 0.00
32.88 0.00
63.67 0.00
78.80 0.00
84.36 0.00
58.16 0.00
59.12 0.00
FPL (NeurIPS’23)
51.59 0.23
62.58 0.86
19.05 0.37
90.67 0.39
55.49 1.35
61.30 1.72
78.10 0.15
86.44 0.33
58.10 0.28
62.59 0.63
GRIP (NeurIPS’23)
48.01 0.68
61.95 0.99
18.25 0.24
88.98 0.62
58.44 5.52
64.00 1.55
73.06 0.49
84.15 0.29
50.93 0.36
60.86 1.19
CPL (ICML’24)
49.27 0.69
61.42 1.22
17.58 0.32
82.77 0.37
63.49 6.24
61.91 1.04
72.11 0.23
86.28 0.30
49.31 0.59
60.46 1.22
UEO (ICML’24)
53.56 0.11
60.07 0.01
18.73 0.15
92.74 0.08
43.68 0.14
65.71 0.07
80.67 0.04
87.94 0.09
58.03 0.18
62.35 0.09
Table 1: Comparison of the test accuracy (%) across three seeds on nine benchmark datasets with different imbalance ratios (10, 20, 50). The results are reported as meanstd over three random seeds, with the best results presented in bold. Zero-shot CLIP is abbreviated as ZS CLIP.
Method
Cal.
Flw.
Pets.
Avg.
ZS CLIP
90.30
63.67
84.36
79.44
FPL
92.08
67.47
87.41
82.32
CPL
92.37
68.20
90.32
83.63
UEO
92.98
66.82
89.13
82.98
CAP
94.36
72.35
91.22
85.98
microCLIP
94.12
72.11
89.92
85.38
Table 2: Comparison of test accuracy (%) under the balanced setting.
Method
Res.
Flw.
ES.
Pets.
Avg.
ZS OCLIP
63.90
71.43
51.74
90.62
69.42
FPL
60.59
66.04
50.32
87.54
66.12
GRIP
63.25
62.08
61.24
80.84
66.85
CPL
62.68
63.41
64.44
84.60
68.78
UEO
59.57
70.35
52.06
88.88
67.72
TMP
41.07
13.60
10.00
47.75
28.11
Table 3: Test accuracy (%) using OpenCLIP with an imbalance ratio of 20. Zero-shot OpenCLIP is abbreviated as ZS OCLIP.
BPA
MSR
Imbalance Ratio = 10
Imbalance Ratio = 20
Imbalance Ratio = 50
Avg.
Res.
Flw.
ES.
Res.
Flw.
ES.
Res.
Flw.
ES.
✗
✗
54.48
63.67
32.88
54.48
63.67
32.88
54.48
63.67
32.88
50.34
✓
✗
67.05
65.23
57.26
64.69
63.98
61.92
63.34
64.20
50.68
62.04 ↑ 11.70
✗
✓
61.50
68.42
52.16
61.71
67.33
52.26
58.65
66.92
52.18
60.13 ↑ 0 9.79
✓
✓
69.43
69.41
74.40
68.52
67.73
74.22
68.66
67.98
73.68
70.45 ↑ 20.11
Table 4: Ablation results of Boundary-Preserving Alignment (BPA) and Margin-aware Self-Refinement (MSR) with different imbalance ratios. The improvement is indicated by ↑ .
Figure 4: Density curve of logit margin changes ( Δm ). (a) Existing methods cause a negative shift for head classes. (b) Our method maintains positive margin changes.
Figure 5: Evaluation of margin-aware self-refinement on RESISC45. (a) Accuracy across head, medium, and tail groups. (b) Frequency consistency on confusable classes.
Method
Places-LT
ImageNet-LT
iNaturalist2018
Avg.
ZS CLIP
38.33
59.71
69.33
55.79
FPL
39.31
60.05
69.50
56.29
GRIP
36.64
59.92
66.60
54.39
CPL
38.38
55.77
65.30
53.15
UEO
39.71
60.23
70.17
56.70
TMP
32.18
48.32
42.17
40.89
Table 5: Comparison of test accuracy (%) on large-scale long-tailed datasets.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Classes
Train
Test
Balanced
CUB ( Wah et al., 2011 )
200
5,994
5,794
✗
RESISC45 ( Cheng et al., 2017 )
45
6,300
25,200
✓
FGVCAircraft ( Maji et al., 2013 )
100
6,667
3,333
✓
Caltech101 ( Fei-Fei et al., 2004 )
100
5,777
2,465
✗
EuroSAT ( Helber et al., 2019 )
10
22,000
5,000
✓
Flowers102 ( Nilsback and Zisserman, 2008 )
102
2,040
6,149
✗
Appendix
Table 6: Statistics for benchmark datasets used in main experiments. The rightmost column indicates whether the testing set is balanced.
Figure 6: Sensitivity to β on Caltech101 (left) and RESISC45 (right).
Method
IR = 10
IR = 20
IR = 50
Total
FPL
8 / 1 / 0
9 / 0 / 0
9 / 0 / 0
26 / 1 / 0
GRIP
9 / 0 / 0
8 / 1 / 0
8 / 1 / 0
25 / 2 / 0
CPL
8 / 1 / 0
8 / 1 / 0
8 / 1 / 0
24 / 3 / 0
UEO
9 / 0 / 0
6 / 3 / 0
7 / 2 / 0
22 / 5 / 0
TMP
9 / 0 / 0
9 / 0 / 0
9 / 0 / 0
27 / 0 / 0
CAP
4 / 5 / 0
7 / 2 / 0
7 / 2 / 0
18 / 9 / 0
Appendix
Table 7: Statistical significance of performance differences assessed with a two-sided paired t-test at a 0.05 significance level over three corresponding seeds, reported as win/tie/loss counts. Imbalance Ratio is abbreviated as IR.
Method
Imbalance Ratio = 10
Imbalance Ratio = 50
Avg.
Res.
Flw.
ES.
Pets.
Res.
Flw.
ES.
Pets.
ZS OpenCLIP
63.90
71.43
51.74
90.62
63.90
71.43
51.74
90.62
69.42
FPL
64.05
66.43
50.10
88.03
59.27
65.18
49.98
86.92
66.25
GRIP
64.63
64.29
62.08
82.47
57.94
66.12
60.20
81.22
67.37
CPL
66.66
63.20
64.18
86.67
59.65
63.70
60.42
81.38
68.23
UEO
62.49
70.16
53.54
88.66
58.77
70.22
51.44
89.02
68.04
Appendix
Table 8: Comparison of test accuracy (%) using OpenCLIP with imbalance ratios of 10 and 50. Zero-shot OpenCLIP is abbreviated as ZS OpenCLIP.
Method
Imbalance Ratio = 10
Imbalance Ratio = 20
Imbalance Ratio = 50
Avg.
Res.
Flw.
ES.
Pets.
Res.
Flw.
ES.
Pets.
Res.
Flw.
ES.
Pets.
ZS SigLIP
59.23
83.46
35.74
93.10
59.23
83.46
35.74
93.10
59.23
83.46
35.74
93.10
67.88
FPL
61.14
82.96
37.12
92.80
60.04
83.36
35.14
92.45
60.23
83.48
35.62
92.83
68.10
GRIP
62.36
83.72
45.38
91.52
61.73
83.66
42.62
91.71
60.62
83.72
34.72
91.11
69.41
CPL
65.56
81.33
46.74
90.92
62.76
81.22
52.14
88.91
60.50
82.50
44.24
88.72
70.46
UEO
59.65
83.48
40.10
93.08
59.55
83.46
39.94
93.10
59.40
83.46
39.74
93.08
69.00
Appendix
Table 9: Comparison of test accuracy (%) using SigLIP with varying imbalance ratios (10, 20, 50). Zero-shot SigLIP is abbreviated as ZS SigLIP.
Method
Imbalance Ratio = 10
Imbalance Ratio = 20
Imbalance Ratio = 50
Avg.
Res.
Cub.
Flw.
Pets.
Res.
Cub.
Flw.
Pets.
Res.
Cub.
Flw.
Pets.
ZS CLIP
62.25
55.32
71.15
88.20
62.25
55.32
71.15
88.20
62.25
55.32
71.15
88.20
69.23
FPL
65.02
56.81
66.74
89.40
63.02
55.50
65.08
89.29
61.38
55.35
65.88
88.85
68.53
GRIP
65.47
54.17
69.18
86.26
61.22
53.86
69.52
85.88
58.41
52.12
67.67
84.08
67.32
CPL
66.72
54.98
67.23
88.80
62.83
53.27
66.08
86.24
58.16
53.33
67.18
84.06
67.41
UEO
61.62
57.18
70.92
89.62
60.75
56.59
70.73
89.59
60.15
56.61
70.60
89.45
69.48
Appendix
Table 10: Comparison of test accuracy (%) using ViT-B/16 as the visual backbone with imbalance ratios of 10, 20, and 50.
Method
Res.
ES.
Overall
Head
Medium
Tail
Overall
Head
Medium
Tail
FPL
57.61
46.59
58.52
59.72
55.51
23.40
60.20
58.52
CPL
56.73
38.69
57.40
60.57
58.82
13.00
65.53
63.10
UEO
56.97
51.90
54.43
59.83
42.42
6.80
44.67
47.23
CAP
58.78
37.98
59.88
63.05
50.38
26.40
60.47
49.33
MARS
68.70
70.55
67.01
69.11
75.68
78.60
76.60
74.73
Appendix
Table 11: Comparison of test accuracy (%) across head, medium, and tail groups on RESISC45 and EuroSAT with an imbalance ratio of 50.
Dataset
Method
Head Acc.
Tail Acc.
H-to-H
H-to-T
RESISC45
ZS CLIP
69.20
56.84
429
428
CPL
38.69
60.57
164
1,060
BPA
72.89
65.37
210
392
OxfordPets
ZS CLIP
86.75
91.47
28
18
CPL
80.12
92.14
17
56
BPA
88.14
92.24
30
16
Appendix
Table 12: Head/tail accuracy (%) and prediction flow of true head-class samples on RESISC45 and OxfordPets with an imbalance ratio of 50.
Dataset
Cases
CPL Δm
BPA Δm
No longer tail
Exact repairs
Caltech101
204
−3.88
+2.28
199 (97.55%)
192 (94.12%)
Food101
1,579
−4.41
+1.04
1,204 (76.25%)
965 (61.11%)
Appendix
Table 13: Average logit margin changes relative to zero-shot CLIP and repairs over all true head-class samples predicted as tail classes by CPL on Caltech101 and Food101 with an imbalance ratio of 50. “No longer tail” counts the samples that BPA no longer predicts as tail classes, and “Exact repairs” counts those that BPA predicts correctly.
Method
Res.
Pets.
Base
New
H
Base
New
H
ZS CLIP
55.36
53.56
54.45
84.77
83.95
84.36
FPL
69.05
51.39
58.92
83.68
86.63
85.13
GRIP
64.15
42.65
51.23
77.67
85.23
81.28
CPL
66.33
45.27
53.81
82.46
85.68
84.04
UEO
60.02
49.66
54.35
85.06
88.59
86.79
Appendix
Table 14: Base-to-new generalization on RESISC45 and OxfordPets (imbalance ratio=50).
Hangzhou Dianzi University, Hangzhou, China · Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · The Hong Kong Polytechnic University, Hong Kong, China +2
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China · College of Computer Science and Software Engineering, Shenzhen University, China · Institute of High Performance Computing, A*STAR, Singapore