Are In-Context Images Worth 10 Dimensions?
Organizations: Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli, Paris, France · Pôle recherche de l’AMIAD, Palaiseau, France · Institut Universitaire de France, Paris, France
Abstract
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.
Figures & tables
| Ablation effect (3-shot) | |
| (1) 128D- Text SDG at | |
| (2) 128D- Vision SDG at | |
| (3) Attention from vision at | |
| to text at | |
| 128D- Text SDG at |
| Classification | Matching | ||||||
| Open MI | VLGuard | VizWiz | Matching MI | SugarCrepe | MHaluBench | NaturalBench | |
| Baseline | 84.0 | 63.0 | 72.0 | 81.0 | 57.5 | 75.0 | 61.0 |
| 128D-SDG Ablation | -11.0 | -17.0 | -4.0 | +6.5 | +5.5 | +3.0 | -2.0 |
| 128D-SDG Boost | +11.5 | +6.0 | +1.5 | -4.0 | -3.0 | -5.5 | -1.0 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Ministral 3 (8B) | 48.9 | 0.0 |
| Qwen 3.5 (9B) | 23.3 | 0.0 |
| Training setup | Support counts used for the loss | Effective dimension of |
| 3-way 3-shot | 3, 6, 9 | 9.69 |
| 3-way 3-shot | 9 | 9.68 |
| 4-way 4-shot | 4, 8, 12, 16 | 11.08 |
| 4-way 4-shot | 16 | 10.79 |
| Vision SDG | Effective dimension |
| training data | at ( ) |
| All 10 datasets | 9.69 |
| Textures only | 9.71 |
| Aircraft only | 6.57 |
| Intervention | 1-shot | 2-shot | 3-shot | 4-shot | 5-shot | |
| LVLM Accuracy | 34.4% | 77.2% | 82.4% | 84.2% | 85.4% | |
| All | 1D to 128D - Control Ablation | -0.4 | -1.2 | -1.9 | -2.5 | -2.2 |
| 1D to 15D - SDG Ablation | -7.1 | -10.8 | -12.9 | -12.0 | -10.6 | |
| 1D to 128D - SDG Ablation | -13.3 | -11.3 | -11.4 | -11.7 | -9.7 | |
| 16D to 128D - SDG Ablation | -5.7 | +1.8 | +1.6 | +1.4 | +0.5 | |
| 1D to 15D - SDG Boost | +7.5 | +4.7 | +3.9 | +3.0 | +2.9 | |
| Head 7 | Head 8 | Head 17 | |
| Flowers | |||
| Textures | |||
| Cars |
| Model | Attention | Layers | Heads | Total heads | Token dim. | Head dim. | MLP dim. |
| Ministral 3 (3B) | Full | 26 | 32 | 832 | 3072 | 128 | 9216 |
| Ministral 3 (8B) | Full | 34 | 32 | 1088 | 4096 | 128 | 14336 |
| Qwen 3.5 (4B) | Full/linear | 32 | 16/32 | 128/768 | 2560 | 256/128 | 9216 |
| Qwen 3.5 (9B) | Full/linear | 32 | 16/32 | 128/768 | 4096 | 256/128 | 12288 |
| Qwen 3 (4B) | Full | 36 | 32 | 1152 | 2560 | 128 | 9728 |
| Model | Effective dimension of | |
| Ministral 3 (3B) | 16 | 10.22 |
| Ministral 3 (8B) | 17 | 10.28 |
| Qwen 3.5 (9B) | 19 | 7.46 |
| Qwen 3.5 (4B) | 18 | 9.18 |
| Qwen 3 (4B) | 23 | 8.92 |
| Dataset | Question format | Answers |
| Open MI | This is a | Provided concept names |
| VLGuard | Original instruction field | harmful / unharmful |
| VizWiz | Original question field | answerable / unanswerable |
| Matching MI | Do the two images satisfy the induced relationship? | Yes / No |
| SugarCrepe | Does this image show ’<caption>’? | Yes / No |
| MHaluBench | Original claim field | hallucination / non-hallucination |