Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · School of Computing and Information Sciences, Saint Francis University, Hong Kong, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China
Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.
Figures & tables
Figure 1: (a) The overall test accuracy of the existing unsupervised method CPL ( Zhang et al., 2024 ) under different imbalance ratios. (b) Comparison of head- and tail-class accuracies across three settings on the RESISC45 dataset: unsupervised learning with balanced data (Unsup. Bal.), supervised learning with imbalanced data (Sup. Imb.), and unsupervised learning with imbalanced data (Unsup. Imb.). (c) Comparison of average logit margin changes ( Δm=marginFine-tuned−marginZero-shot CLIP ). Under the long-tailed setting, the head-class margin significantly decreases compared to the uniform setting, indicating severe boundary erosion. (d) Label counts across three cases: ground-truth (GT), zero-shot predictions (ZS Pred.), and fine-tuned predictions (FT Pred.). We group confusable pairs by color (e.g., blue: “beach” vs. “sea ice”) and observe that zero-shot CLIP inherently exhibits bias toward confusable classes, which is further exacerbated by fine-tuning.
Figure 2: Visualization of attention maps for misassigned samples on RESISC45. We select head-class samples where the baseline method incorrectly assigns tail-class pseudo-labels. (a) The model mainly focuses on the foreground object. (b) The model is distracted by the irrelevant background. (c) The model correctly captures the foreground object while reducing background attention.
Figure 3: (a) BPA preserves head-class accuracy on OxfordPets as the imbalance ratio increases. (b) The complete MARS pipeline maintains stable overall accuracy on OxfordPets across imbalance ratios. (c) Evolution of classification margins during training. Samples with correct pseudo-labels (Correct PL) consistently exhibit increasing margins, whereas mislabeled samples (Wrong PL) fluctuate around zero or remain negative.
Imbalance Ratio = 10
Method
CUB.
Res.
FA.
Cal.
ES.
Flw.
Food.
Pets.
Cars.
Avg.
ZS CLIP
51.82 0.00
54.48 0.00
17.58 0.00
90.30 0.00
32.88 0.00
63.67 0.00
78.80 0.00
84.36 0.00
58.16 0.00
59.12 0.00
FPL (NeurIPS’23)
51.59 0.23
62.58 0.86
19.05 0.37
90.67 0.39
55.49 1.35
61.30 1.72
78.10 0.15
86.44 0.33
58.10 0.28
62.59 0.63
GRIP (NeurIPS’23)
48.01 0.68
61.95 0.99
18.25 0.24
88.98 0.62
58.44 5.52
64.00 1.55
73.06 0.49
84.15 0.29
50.93 0.36
60.86 1.19
CPL (ICML’24)
49.27 0.69
61.42 1.22
17.58 0.32
82.77 0.37
63.49 6.24
61.91 1.04
72.11 0.23
86.28 0.30
49.31 0.59
60.46 1.22
UEO (ICML’24)
53.56 0.11
60.07 0.01
18.73 0.15
92.74 0.08
43.68 0.14
65.71 0.07
80.67 0.04
87.94 0.09
58.03 0.18
62.35 0.09
Table 1: Comparison of the test accuracy (%) across three seeds on nine benchmark datasets with different imbalance ratios (10, 20, 50). The results are reported as meanstd over three random seeds, with the best results presented in bold. Zero-shot CLIP is abbreviated as ZS CLIP.
Method
Cal.
Flw.
Pets.
Avg.
ZS CLIP
90.30
63.67
84.36
79.44
FPL
92.08
67.47
87.41
82.32
CPL
92.37
68.20
90.32
83.63
UEO
92.98
66.82
89.13
82.98
CAP
94.36
72.35
91.22
85.98
microCLIP
94.12
72.11
89.92
85.38
Table 2: Comparison of test accuracy (%) under the balanced setting.
Method
Res.
Flw.
ES.
Pets.
Avg.
ZS OCLIP
63.90
71.43
51.74
90.62
69.42
FPL
60.59
66.04
50.32
87.54
66.12
GRIP
63.25
62.08
61.24
80.84
66.85
CPL
62.68
63.41
64.44
84.60
68.78
UEO
59.57
70.35
52.06
88.88
67.72
TMP
41.07
13.60
10.00
47.75
28.11
Table 3: Test accuracy (%) using OpenCLIP with an imbalance ratio of 20. Zero-shot OpenCLIP is abbreviated as ZS OCLIP.
BPA
MSR
Imbalance Ratio = 10
Imbalance Ratio = 20
Imbalance Ratio = 50
Avg.
Res.
Flw.
ES.
Res.
Flw.
ES.
Res.
Flw.
ES.
✗
✗
54.48
63.67
32.88
54.48
63.67
32.88
54.48
63.67
32.88
50.34
✓
✗
67.05
65.23
57.26
64.69
63.98
61.92
63.34
64.20
50.68
62.04 ↑ 11.70
✗
✓
61.50
68.42
52.16
61.71
67.33
52.26
58.65
66.92
52.18
60.13 ↑ 0 9.79
✓
✓
69.43
69.41
74.40
68.52
67.73
74.22
68.66
67.98
73.68
70.45 ↑ 20.11
Table 4: Ablation results of Boundary-Preserving Alignment (BPA) and Margin-aware Self-Refinement (MSR) with different imbalance ratios. The improvement is indicated by ↑ .
Figure 4: Density curve of logit margin changes ( Δm ). (a) Existing methods cause a negative shift for head classes. (b) Our method maintains positive margin changes.
Figure 5: Evaluation of margin-aware self-refinement on RESISC45. (a) Accuracy across head, medium, and tail groups. (b) Frequency consistency on confusable classes.
Method
Places-LT
ImageNet-LT
iNaturalist2018
Avg.
ZS CLIP
38.33
59.71
69.33
55.79
FPL
39.31
60.05
69.50
56.29
GRIP
36.64
59.92
66.60
54.39
CPL
38.38
55.77
65.30
53.15
UEO
39.71
60.23
70.17
56.70
TMP
32.18
48.32
42.17
40.89
Table 5: Comparison of test accuracy (%) on large-scale long-tailed datasets.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Classes
Train
Test
Balanced
CUB ( Wah et al., 2011 )
200
5,994
5,794
✗
RESISC45 ( Cheng et al., 2017 )
45
6,300
25,200
✓
FGVCAircraft ( Maji et al., 2013 )
100
6,667
3,333
✓
Caltech101 ( Fei-Fei et al., 2004 )
100
5,777
2,465
✗
EuroSAT ( Helber et al., 2019 )
10
22,000
5,000
✓
Flowers102 ( Nilsback and Zisserman, 2008 )
102
2,040
6,149
✗
Appendix
Table 6: Statistics for benchmark datasets used in main experiments. The rightmost column indicates whether the testing set is balanced.
Figure 6: Sensitivity to β on Caltech101 (left) and RESISC45 (right).
Method
IR = 10
IR = 20
IR = 50
Total
FPL
8 / 1 / 0
9 / 0 / 0
9 / 0 / 0
26 / 1 / 0
GRIP
9 / 0 / 0
8 / 1 / 0
8 / 1 / 0
25 / 2 / 0
CPL
8 / 1 / 0
8 / 1 / 0
8 / 1 / 0
24 / 3 / 0
UEO
9 / 0 / 0
6 / 3 / 0
7 / 2 / 0
22 / 5 / 0
TMP
9 / 0 / 0
9 / 0 / 0
9 / 0 / 0
27 / 0 / 0
CAP
4 / 5 / 0
7 / 2 / 0
7 / 2 / 0
18 / 9 / 0
Appendix
Table 7: Statistical significance of performance differences assessed with a two-sided paired t-test at a 0.05 significance level over three corresponding seeds, reported as win/tie/loss counts. Imbalance Ratio is abbreviated as IR.
Method
Imbalance Ratio = 10
Imbalance Ratio = 50
Avg.
Res.
Flw.
ES.
Pets.
Res.
Flw.
ES.
Pets.
ZS OpenCLIP
63.90
71.43
51.74
90.62
63.90
71.43
51.74
90.62
69.42
FPL
64.05
66.43
50.10
88.03
59.27
65.18
49.98
86.92
66.25
GRIP
64.63
64.29
62.08
82.47
57.94
66.12
60.20
81.22
67.37
CPL
66.66
63.20
64.18
86.67
59.65
63.70
60.42
81.38
68.23
UEO
62.49
70.16
53.54
88.66
58.77
70.22
51.44
89.02
68.04
Appendix
Table 8: Comparison of test accuracy (%) using OpenCLIP with imbalance ratios of 10 and 50. Zero-shot OpenCLIP is abbreviated as ZS OpenCLIP.
Method
Imbalance Ratio = 10
Imbalance Ratio = 20
Imbalance Ratio = 50
Avg.
Res.
Flw.
ES.
Pets.
Res.
Flw.
ES.
Pets.
Res.
Flw.
ES.
Pets.
ZS SigLIP
59.23
83.46
35.74
93.10
59.23
83.46
35.74
93.10
59.23
83.46
35.74
93.10
67.88
FPL
61.14
82.96
37.12
92.80
60.04
83.36
35.14
92.45
60.23
83.48
35.62
92.83
68.10
GRIP
62.36
83.72
45.38
91.52
61.73
83.66
42.62
91.71
60.62
83.72
34.72
91.11
69.41
CPL
65.56
81.33
46.74
90.92
62.76
81.22
52.14
88.91
60.50
82.50
44.24
88.72
70.46
UEO
59.65
83.48
40.10
93.08
59.55
83.46
39.94
93.10
59.40
83.46
39.74
93.08
69.00
Appendix
Table 9: Comparison of test accuracy (%) using SigLIP with varying imbalance ratios (10, 20, 50). Zero-shot SigLIP is abbreviated as ZS SigLIP.
Method
Imbalance Ratio = 10
Imbalance Ratio = 20
Imbalance Ratio = 50
Avg.
Res.
Cub.
Flw.
Pets.
Res.
Cub.
Flw.
Pets.
Res.
Cub.
Flw.
Pets.
ZS CLIP
62.25
55.32
71.15
88.20
62.25
55.32
71.15
88.20
62.25
55.32
71.15
88.20
69.23
FPL
65.02
56.81
66.74
89.40
63.02
55.50
65.08
89.29
61.38
55.35
65.88
88.85
68.53
GRIP
65.47
54.17
69.18
86.26
61.22
53.86
69.52
85.88
58.41
52.12
67.67
84.08
67.32
CPL
66.72
54.98
67.23
88.80
62.83
53.27
66.08
86.24
58.16
53.33
67.18
84.06
67.41
UEO
61.62
57.18
70.92
89.62
60.75
56.59
70.73
89.59
60.15
56.61
70.60
89.45
69.48
Appendix
Table 10: Comparison of test accuracy (%) using ViT-B/16 as the visual backbone with imbalance ratios of 10, 20, and 50.
Method
Res.
ES.
Overall
Head
Medium
Tail
Overall
Head
Medium
Tail
FPL
57.61
46.59
58.52
59.72
55.51
23.40
60.20
58.52
CPL
56.73
38.69
57.40
60.57
58.82
13.00
65.53
63.10
UEO
56.97
51.90
54.43
59.83
42.42
6.80
44.67
47.23
CAP
58.78
37.98
59.88
63.05
50.38
26.40
60.47
49.33
MARS
68.70
70.55
67.01
69.11
75.68
78.60
76.60
74.73
Appendix
Table 11: Comparison of test accuracy (%) across head, medium, and tail groups on RESISC45 and EuroSAT with an imbalance ratio of 50.
Dataset
Method
Head Acc.
Tail Acc.
H-to-H
H-to-T
RESISC45
ZS CLIP
69.20
56.84
429
428
CPL
38.69
60.57
164
1,060
BPA
72.89
65.37
210
392
OxfordPets
ZS CLIP
86.75
91.47
28
18
CPL
80.12
92.14
17
56
BPA
88.14
92.24
30
16
Appendix
Table 12: Head/tail accuracy (%) and prediction flow of true head-class samples on RESISC45 and OxfordPets with an imbalance ratio of 50.
Dataset
Cases
CPL Δm
BPA Δm
No longer tail
Exact repairs
Caltech101
204
−3.88
+2.28
199 (97.55%)
192 (94.12%)
Food101
1,579
−4.41
+1.04
1,204 (76.25%)
965 (61.11%)
Appendix
Table 13: Average logit margin changes relative to zero-shot CLIP and repairs over all true head-class samples predicted as tail classes by CPL on Caltech101 and Food101 with an imbalance ratio of 50. “No longer tail” counts the samples that BPA no longer predicts as tail classes, and “Exact repairs” counts those that BPA predicts correctly.
Method
Res.
Pets.
Base
New
H
Base
New
H
ZS CLIP
55.36
53.56
54.45
84.77
83.95
84.36
FPL
69.05
51.39
58.92
83.68
86.63
85.13
GRIP
64.15
42.65
51.23
77.67
85.23
81.28
CPL
66.33
45.27
53.81
82.46
85.68
84.04
UEO
60.02
49.66
54.35
85.06
88.59
86.79
Appendix
Table 14: Base-to-new generalization on RESISC45 and OxfordPets (imbalance ratio=50).
Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs). Despite its promise, current methods still struggle to maintain tail-class discriminability when adapting to class-imbalanced datasets. In this work, we propose cluster-aware neural collapse prompt tuning (CPT), which enhances the discriminability of tail classes in prompt-tuned VLMs without sacrificing their overall generalization. First, we design a cluster-invariant space by mining semantic assignments from the pre-trained VLM and mapping them to prompt-tuned features. This computes cluster-level boundaries and restricts the constraints to local neighborhoods, which reduces interference with the global semantic structure of the pre-trained VLM. Second, we introduce neural-collapse-driven discriminability optimization with three losses: textual Equiangular Tight Frame (ETF) separation loss, class-wise convergence loss, and rotation stabilization loss. These losses work together to shape intra-cluster geometry for better inter-class separation and intra-class alignment. Extensive experiments on 11 diverse datasets demonstrate that CPT outperforms SOTA methods, with stronger performance on long-tail classes and good generalization to unseen classes.
Boyang Guo, Liang Li, Lin Peng +3
Hangzhou Dianzi University, Hangzhou, China · Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · The Hong Kong Polytechnic University, Hong Kong, China +2
Learning from real-world data is frequently hindered by the compound challenge of long-tailed class distributions and noisy annotations. Existing methods partially address these issues but typically ignore the non-uniform impact of label noise across classes, resulting in ineffective correction for tail classes and over-regularization for head classes. To address this issue, we propose Class-Adaptive Rectification with Experts (CARE), a parameter-efficient framework that leverages three complementary supervision sources from vision-language models (VLM): observed noisy labels, VLM text embeddings, and visual features. CARE introduces a class-adaptive expert consensus mechanism that enforces stricter agreement for tail classes and more permissive agreement for head classes based on class frequency. By aggregating high-confidence predictions across these sources, CARE filters unreliable signals and recalibrates class distributions, yielding more reliable rectification under long-tailed distributions. Extensive experiments on both synthetic and real-world benchmarks demonstrate that CARE consistently outperforms state-of-the-art methods, achieving up to 3.0% performance gains. The source code is available at https://github.com/qwq123-study/CARE.
Mengke Li, Haiquan Ling, Lihao Chen +3
College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China
Long-tailed distributions are common in real-world recognition tasks, where a few head classes have many samples while most tail classes have very few. Recently, fine-tuning foundation models for long-tailed learning has gained attention due to their excellent performance. However, most existing methods focus solely on mitigating long-tailed distribution bias while overlooking concept confusion caused by the long-tailed distribution. In this paper, we study this problem and attribute it to the mutual exclusivity of single-label supervision under long-tailed distributions, which suppresses feature sharing among related classes and amplifies the dominance of head classes, leading to disrupted inter-class discriminability. To address this, we propose CUE, Concept-aware mUlti-label Expansion, which introduces multi-label concept signals to preserve disrupted inter-class relationships. Specifically, CUE constructs concept sets by (i) extracting instance-level visual cues from zero-shot CLIP and (ii) generating class-level semantic cues with LLM; the two cues are incorporated via separately weighted Binary Logit-Adjustment (BLA) auxiliary losses and jointly optimized with the baseline Logit-Adjustment (LA) loss. Experiments on several long-tailed benchmarks, CUE achieves balanced and strong performance, surpassing recent state-of-the-art methods. Code is available at: https://github.com/zhangruichi/CUE.
Ruichi Zhang, Chikai Shang, Jiacheng Yang +4
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China · College of Computer Science and Software Engineering, Shenzhen University, China · Institute of High Performance Computing, A*STAR, Singapore