cs.CVSep 29, 2026

Visual Branch is What You Need for CLIP-based Class-Incremental Learning

Authors: Tao Hu, Zhen-Hao Xie, Jingcai Guo, De-Chuan Zhan, Da-Wei zhou

Organizations: School of Artificial Intelligence, Nanjing University · State Key Laboratory for Novel Software Technology, Nanjing University · Hong Kong Polytechnic University

Abstract

Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual features. Motivated by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VIS uses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VIS employs a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VIS accumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VIS achieves state-of-the-art performance without a textual branch.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning

    Nov 14, 2025Lan Li, Tao Hu, Da-Wei Zhou +3Class-Incremental LearningRepresentation Learning

  2. Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning

    May 13, 2026Hao Sun, Zi-Jun Ding, Da-Wei ZhouClass-Incremental LearningVisual Embeddings

  3. Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning

    Sep 29, 2026Chiyuan He, Zihuan Qiu, Fanman Meng +5Class-Incremental LearningImage-Text Alignment