Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification
Authors: Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bartłomiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
Organizations: Department of Computer Science, Universitat Autònoma de Barcelona, Spain · Computer Vision Center, Barcelona, Spain · Media Integration and Communication Center, University of Florence, Italy · Bernoulli Institute, University of Groningen, the Netherlands · IDEAS Research Institute, Poland · ESAT-PSI, KU Leuven, Belgium
Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that exploiting few-shot image embeddings from a training set is effective for CLIP-based classification. In this work, we analyze mixing image and text prototypes from a bias-variance perspective and show that mixing prototypes acts like a variance shrinkage estimator. Naively mixing text and image prototypes combines two partially aligned spaces since the two modalities are not perfectly aligned. To address this, we project image prototypes onto the principal directions of the semantic text embedding space to obtain a task-semantic image subspace. Mixing the image prototypes with text embeddings in the task-semantic subspace improves few-shot classification. However, when the task-semantic subspace captures insufficient discriminative visual information, relying on this subspace alone can be suboptimal. On extensive experiments over several few-shot classification benchmarks, we show that combining a task-semantic mixed prototype classifier and an anisotropic image-specific classifier systematically outperforms existing methods.
Figures & tables
Figure 1: The latent space spanned by text prototypes defines a task-semantic subspace . Task semantic regions denoted by attention maps are obtained by the proposed task-specific semantic projection P1 and P2.
Figure 2: MSE of image and mixed prototypes as a function of the number of shots using CLIP-B/16.
Figure 3: Impact of the task-semantic projection: (left) Attention maps reveal that visual features projected on the task-semantic subspace focus on the text information; (right) Using NCM classifier for standard datasets (Stanford Cars ( Krause et al., 2013 ) and UCF101 ( Soomro et al., 2012 ) ) on CLIP-ViT-B/16, we show that task-semantic subspace is more discriminative than the full image space while the task-orthogonal subspace performs poorly. For OOD datasets like Lego Bricks ( Garcia, 2019 ) , the full image space and the orthogonal subspace is more discriminative than the semantic subspace due to poor cross-modal alignment in CLIP.
Figure 4: Analysis of μ(SM) .
Figure 5: Impact of different classifiers across 3 datasets and varying shots. The proposed method CMP combines Proj+Mix (NCM) and Image (Maha).
Figure 6: Ablations of using NCM classifiers with image prototypes (Image), image prototypes in task-semantic subspace (Proj), mixed prototypes in the full image space (Mix) and mixed prototypes in the task-semantic subspace (Proj+Mix).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Analysis of the mean of principal angle cosines μ(SM) indicating cross-modal alignment across all datasets in the 4-shot setting using CLIP-ViT-B/16.
Figure 8: Analysis of the impact of varying subsets of classes on the semantic and orthogonal subspaces.
Figure 9: Impact of different classifiers across 3 datasets and varying shots on 12 datasets with CLIP-ViT-B/16.. The proposed method CMP combines Proj+Mix (NCM) and Image (Maha).
Figure 10: Ablations of using NCM classifiers with image prototypes (Image), image prototypes in task-semantic subspace (Proj), mixed prototypes in the full image space (Mix) and mixed prototypes in the task-semantic subspace (Proj+Mix) on 12 datasets with CLIP-ViT-B/16.
Figure 11: Performance comparison of training-free methods using CLIP-ViT-B/16 in the validation-free setting.
Figure 12: Analysis of the sensitivity of the hyper-parameters λ and γ to the validation set accuracy on 5 datasets with CLIP-ViT-B/16.
Department of Computer Science and Artificial Intelligence, DaSCI Institute, University of Granada, Granada, Spain · Department of Mathematics “Tullio Levi-Civita”, University of Padova, Padova, Italy.