cs.CVOct 6, 2026

Unsupervised Long-Tailed Adaptation of Vision-Language Models

Authors: Keliang Chen, Yaxin Hou, Hui Liu, Yuheng Jia

Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · School of Computing and Information Sciences, Saint Francis University, Hong Kong, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China

Abstract

Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models

    May 12, 2026Boyang Guo, Liang Li, Lin Peng +3Vision-Language Model AdaptationNeural Collapse

  2. CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy Labels

    May 22, 2026Mengke Li, Haiquan Ling, Lihao Chen +3Noisy LabelsVision-Language Model Adaptation

  3. CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning

    May 2, 2026Ruichi Zhang, Chikai Shang, Jiacheng Yang +4Multi-Label ClassificationLong-Tailed Distribution