Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification
Authors: Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bartłomiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
Organizations: Department of Computer Science, Universitat Autònoma de Barcelona, Spain · Computer Vision Center, Barcelona, Spain · Media Integration and Communication Center, University of Florence, Italy · Bernoulli Institute, University of Groningen, the Netherlands · IDEAS Research Institute, Poland · ESAT-PSI, KU Leuven, Belgium
Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that exploiting few-shot image embeddings from a training set is effective for CLIP-based classification. In this work, we analyze mixing image and text prototypes from a bias-variance perspective and show that mixing prototypes acts like a variance shrinkage estimator. Naively mixing text and image prototypes combines two partially aligned spaces since the two modalities are not perfectly aligned. To address this, we project image prototypes onto the principal directions of the semantic text embedding space to obtain a task-semantic image subspace. Mixing the image prototypes with text embeddings in the task-semantic subspace improves few-shot classification. However, when the task-semantic subspace captures insufficient discriminative visual information, relying on this subspace alone can be suboptimal. On extensive experiments over several few-shot classification benchmarks, we show that combining a task-semantic mixed prototype classifier and an anisotropic image-specific classifier systematically outperforms existing methods.
Figures & tables
Figure 1: The latent space spanned by text prototypes defines a task-semantic subspace . Task semantic regions denoted by attention maps are obtained by the proposed task-specific semantic projection P1 and P2.
Figure 2: MSE of image and mixed prototypes as a function of the number of shots using CLIP-B/16.
Figure 3: Impact of the task-semantic projection: (left) Attention maps reveal that visual features projected on the task-semantic subspace focus on the text information; (right) Using NCM classifier for standard datasets (Stanford Cars ( Krause et al., 2013 ) and UCF101 ( Soomro et al., 2012 ) ) on CLIP-ViT-B/16, we show that task-semantic subspace is more discriminative than the full image space while the task-orthogonal subspace performs poorly. For OOD datasets like Lego Bricks ( Garcia, 2019 ) , the full image space and the orthogonal subspace is more discriminative than the semantic subspace due to poor cross-modal alignment in CLIP.
Figure 4: Analysis of μ(SM) .
Figure 5: Impact of different classifiers across 3 datasets and varying shots. The proposed method CMP combines Proj+Mix (NCM) and Image (Maha).
Figure 6: Ablations of using NCM classifiers with image prototypes (Image), image prototypes in task-semantic subspace (Proj), mixed prototypes in the full image space (Mix) and mixed prototypes in the task-semantic subspace (Proj+Mix).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Analysis of the mean of principal angle cosines μ(SM) indicating cross-modal alignment across all datasets in the 4-shot setting using CLIP-ViT-B/16.
Figure 8: Analysis of the impact of varying subsets of classes on the semantic and orthogonal subspaces.
Figure 9: Impact of different classifiers across 3 datasets and varying shots on 12 datasets with CLIP-ViT-B/16.. The proposed method CMP combines Proj+Mix (NCM) and Image (Maha).
Figure 10: Ablations of using NCM classifiers with image prototypes (Image), image prototypes in task-semantic subspace (Proj), mixed prototypes in the full image space (Mix) and mixed prototypes in the task-semantic subspace (Proj+Mix) on 12 datasets with CLIP-ViT-B/16.
Figure 11: Performance comparison of training-free methods using CLIP-ViT-B/16 in the validation-free setting.
Figure 12: Analysis of the sensitivity of the hyper-parameters λ and γ to the validation set accuracy on 5 datasets with CLIP-ViT-B/16.
Vision-Language Models (VLMs) excel at tasks like zero-shot classification and cross-modal retrieval by mapping images and text to a shared space, but this requires expensive end-to-end training with massive paired datasets. Current post-hoc alignment methods reduce computational costs by connecting pretrained encoders through lightweight mappings, yet still demand substantial paired data. In this work, we investigate the potential of repurposing the classification heads of pretrained vision models as semantic prototypes. The recycling of these weights, typically discarded after pretraining, unlocks two distinct capabilities: it enables zero-shot alignment by using weights as semantic anchors, and serves as a robust data augmentation strategy by mixing these prototypes with real image-text pairs. We demonstrate that integrating our approach with several state-of-the-art post-hoc alignment techniques consistently boosts accuracy in cross-modal retrieval, zero- and few-shot classification tasks.
David Méndez, Roberto Confalonieri, Natalia Díaz Rodríguez
Department of Computer Science and Artificial Intelligence, DaSCI Institute, University of Granada, Granada, Spain · Department of Mathematics “Tullio Levi-Civita”, University of Padova, Padova, Italy.
Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot image classification remains challenging. To bridge this gap, we present DeCoDe, a simple yet effective technique that enables off-the-shelf MLLMs to act as strong few-shot classifiers without any additional training. Our approach builds on the idea of few-shot classification as a set of pairwise image comparisons, decomposing the task into a set of binary decisions. Given a query image and a support image from a candidate class, the MLLM is prompted to decide whether the two images depict the same class. The logit corresponding to an affirmative response is then used as a similarity score to assign the query image to the most likely class. While this already yields good results, we show that providing additional high-level information, such as the data domain, to the model further improves performance. Our evaluation provides an extensive analysis of various inference variants on a suite of twelve datasets, six established and six newly curated few-shot benchmarks spanning across diverse domains. The results show that the proposed simple decomposition technique can turn off-the-shelf MLLMs into powerful few-shot learners, significantly outperforming current state-of-the-art few-shot methods on both standard and novel domains. Code is available at https://github.com/yunhanwang1105/DeCoDe.
Yunhan Wang, Eshika Khandelwal, Edson Araujo +3
Tuebingen AI Center, University of Tuebingen, Germany
Few-shot prompt learning is an effective strategy for adapting CLIP to downstream tasks, but class-only prompt optimization can overfit base-class supervision and weaken transfer to unseen classes. We propose Concept-Constrained Prompt Learning (CCPL), a lightweight regularization framework that anchors learnable class prompts to frozen concept-level text prototypes without updating CLIP encoders. CCPL learns a set of shared context tokens, instantiates class prompts by appending class names, and constructs frozen concept prototypes from a class-level concept bank. During training, a text-space cosine consistency objective aligns learnable class-prompt embeddings with frozen concept prototypes; concept dropout provides additional regularization against over-reliance on fixed concept lists. At inference, CCPL optionally fuses class-prompt logits with concept-prototype logits using a controllable ensemble weight alpha. Our default configuration uses text-space concept regularization lambda = 0.5, concept dropout p = 0.3 and weak concept-guided fusion (alpha = 0.1), with no KL-based prediction consistency term. Experiments under identical automatically-generated fallback splits show that CCPL improves the base-to-new harmonic mean on DTD (+0.6) and EuroSAT (+2.9) compared with CoOp, while remaining near-neutral on OxfordPets (-0.1). Ablations indicate that text-space concept regularization is consistently beneficial, while the best concept-guided inference strength is dataset- and protocol-sensitive. These results suggest concept constraints are most effective when concept prototypes align naturally with dataset semantics, and identify fine-grained categories as a current boundary condition. The code is released at: https://github.com/richael-sang/concept-constrained-prompt-learning.
Na Sang, Ding Ma, Rui Sang +1
University of California, San Diego, La Jolla, CA, USA · Georgia Institute of Technology, Atlanta, GA, USA · Independent Researcher