cs.CVMar 25, 2026

Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification

Authors: Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bartłomiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer

Organizations: Department of Computer Science, Universitat Autònoma de Barcelona, Spain · Computer Vision Center, Barcelona, Spain · Media Integration and Communication Center, University of Florence, Italy · Bernoulli Institute, University of Groningen, the Netherlands · IDEAS Research Institute, Poland · ESAT-PSI, KU Leuven, Belgium

Abstract

Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that exploiting few-shot image embeddings from a training set is effective for CLIP-based classification. In this work, we analyze mixing image and text prototypes from a bias-variance perspective and show that mixing prototypes acts like a variance shrinkage estimator. Naively mixing text and image prototypes combines two partially aligned spaces since the two modalities are not perfectly aligned. To address this, we project image prototypes onto the principal directions of the semantic text embedding space to obtain a task-semantic image subspace. Mixing the image prototypes with text embeddings in the task-semantic subspace improves few-shot classification. However, when the task-semantic subspace captures insufficient discriminative visual information, relying on this subspace alone can be suboptimal. On extensive experiments over several few-shot classification benchmarks, we show that combining a task-semantic mixed prototype classifier and an anisotropic image-specific classifier systematically outperforms existing methods.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight Recycling

    May 21, 2026David Méndez, Roberto Confalonieri, Natalia Díaz RodríguezVision-Language AlignmentCross-Modal

  2. Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners

    Jun 30, 2026Yunhan Wang, Eshika Khandelwal, Edson Araujo +3Few-Shot LearningImage Classification

  3. Concept-Constrained Prompt Learning for Few-Shot CLIP Adaptation

    Jun 21, 2026Na Sang, Ding Ma, Rui Sang +1Multiple Prompt Learning FrameworksContrastive Language-Image Pre-Training Model