cs.CVSep 29, 2026

Are In-Context Images Worth 10 Dimensions?

Authors: Adhemar de Senneville, Xavier Bou, Jérémy Anger, Rafael Grompone, Gabriele Facciolo

Organizations: Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli, Paris, France · Pôle recherche de l’AMIAD, Palaiseau, France · Institut Universitaire de France, Paris, France

Abstract

There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning

    May 4, 2026Haoyu Wang, Haonan Wang, Yuyan Chen +5Multimodal ReasoningIn-Context Learning