Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings
Organizations: University of Wisconsin–Madison · Northwestern University
Abstract
In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.
Figures & tables
| Idefics2-8B | Qwen-VL-7B | ||||||||||||
| Method | # Params (M) | VizWiz | OK-VQA | DTD | Flowers | CUB | Avg. | VizWiz | OK-VQA | DTD | Flowers | CUB | Avg. |
| Zero-shot | – | 37.74 ±0.00 | 49.44 ±0.00 | 90.23 ±0.00 | 87.59 ±0.00 | 87.84 ±0.00 | 70.57 ±0.00 | 45.49 ±0.00 | 57.51 ±0.00 | 83.77 ±0.00 | 73.88 ±0.00 | 91.10 ±0.00 | 70.35 ±0.00 |
| 4-shot ICL | – | 43.24 ±0.11 | 50.38 ±0.19 | 88.87 ±0.12 | 80.22 ±0.40 | 85.91 ±0.12 | 69.72 ±0.05 | 46.36 ±0.32 | 60.92 ±0.07 | 84.44 ±0.25 | 87.17 ±0.05 | 88.73 ±0.25 | 73.53 ±0.12 |
| PT-Pre | 0.082 ( 10.00) | 66.19 ±1.26 | 54.49 ±2.57 | 94.64 ±0.35 | 87.50 ±0.41 | 94.25 ±0.53 | 79.41 ±0.77 | 63.83 ±0.82 | 50.56 ±0.73 | 94.44 ±0.85 | 93.83 ±0.26 | 95.65 ±0.71 | 79.66 ±0.29 |
| PT-App | 0.082 ( 10.00) | 64.56 ±2.43 | 55.10 ±0.56 | 94.14 ±0.07 | 89.60 ±0.10 | 95.10 ±0.15 | 79.70 ±0.63 | 60.19 ±3.26 | 37.84 ±4.11 | 93.10 ±0.96 | 93.56 ±0.28 | 94.14 ±0.43 | 75.77 ±1.03 |
| PT-BL | 0.082 ( 10.00) | 66.35 ±0.21 | 56.58 ±0.98 | 94.13 ±0.19 | 88.97 ±0.52 | 94.55 ±1.54 | 80.12 ±0.22 | 60.30 ±4.64 | 44.35 ±3.61 | 93.44 ±0.28 | 93.49 ±0.29 | 94.73 ±0.50 | 77.26 ±0.48 |
| Vanilla ICL | PT | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Zero-shot | LoRA | Prefix | Pre | App | BL | FV | TV | SITE | STAVE | |||
| Pythia 2.8B | 13.1 ±0.0 | 60.4 ±0.6 | 85.6 ±0.4 | 90.6 ±0.2 | 79.1 ±5.1 | 61.8 ±1.6 | 70.4 ±5.7 | 79.4 ±6.7 | 78.4 ±5.2 | 40.7 ±2.1 | 75.6 ±0.1 | 89.9 ±0.2 | 91.2 ±0.3 |
| Pythia 6.9B | 10.9 ±0.0 | 65.0 ±0.9 | 86.6 ±0.5 | 90.6 ±0.1 | 78.1 ±2.4 | 65.7 ±2.6 | 70.5 ±6.2 | 83.5 ±1.8 | 76.1 ±8.2 | 42.4 ±3.0 | 75.5 ±0.2 | 90.2 ±0.1 | 91.4 ±0.2 |
| Pythia 12B | 11.4 ±0.0 | 64.5 ±0.3 | 87.0 ±0.4 | 91.0 ±0.1 | 81.5 ±4.6 | 55.2 ±3.8 | 71.8 ±4.1 | 82.9 ±1.0 | 73.9 ±3.9 | 42.2 ±0.8 | 75.0 ±0.4 | 91.1 ±0.4 | 91.2 ±0.5 |
| LLaMA 7B | 17.1 ±0.0 | 67.8 ±0.5 | 88.4 ±0.2 | 90.6 ±0.3 | 76.1 ±2.7 | 70.2 ±6.8 | 59.9 ±3.8 | 87.5 ±2.5 | 82.5 ±4.7 | 59.4 ±0.5 | 80.4 ±0.9 | 92.0 ±0.2 | 91.9 ±0.4 |
| GPT-J 6B | 10.2 ±0.0 | 63.6 ±0.4 | 86.3 ±0.3 | 91.0 ±0.2 | 73.7 ±4.0 | 66.5 ±6.1 | 77.8 ±3.6 | 86.0 ±0.8 | 84.9 ±3.2 | 56.3 ±1.7 | 72.4 ±0.6 | 91.8 ±0.2 | 91.9 ±0.4 |
| vs. ICL | Tokens | TTFT (ms) | Mem. (GiB) |
|---|---|---|---|
| Zero-shot | 128 | 98 | 16.5 |
| 8-shot ICL | 1,026 | 1,430 | 29.8 |
| 16-shot ICL | 1,922 | 2,854 | 43.0 |
| 32-shot ICL | 3,717 | 5,877 | 68.7 |
| STAVE | 128 | 98 | 16.5 |
| vs. trained | State (KiB) | TTFT (ms) | Decode (ms/tok) |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | State (KiB) | Prompt tok. | TTFT (ms) | Decode (ms/tok) | Fixed-20 (ms) | Peak mem. (GiB) |
|---|---|---|---|---|---|---|
| Zero-shot | 0 | 128 | 98 | 23.9 | 560 | 16.5 |
| 8-shot ICL | 0 | 1,026 | 1,430 | 24.0 | 2,021 | 29.8 |
| 16-shot ICL | 0 | 1,922 | 2,854 | 24.4 | 3,605 | 43.0 |
| 32-shot ICL | 0 | 3,717 | 5,877 | 26.1 | 6,933 | 68.7 |
| LoRA ( ) | 68,704 | 128 | 116 | 33.6 | 763 | 16.5 |
| LIVE | 512 | 128 | 101 | 27.4 | 629 | 16.5 |
| Method | Stored state | Parameters | Size |
|---|---|---|---|
| Zero-shot, ICL | none | 0 | 0 |
| Single vector (TV, FV) | one embedding vector | 4,096 | 16 KiB |
| ICV | one vector per layer | 131,072 | 512 KiB |
| MTV | one vector per head | 131,072 | 512 KiB |
| LIVE | one vector and one scale per layer | 131,104 | 512 KiB |
| I2CL | two vectors and four scales per layer | 262,272 | 1.0 MiB |
| Backbone | Language model | Vision encoder and connector | Img. | ||
|---|---|---|---|---|---|
| Idefics2-8B | Mistral-7B | SigLIP-SO400M, perceiver resampler | 4096 | 32 | 64 |
| LLaVA-Interleave-7B | Qwen1.5-7B | SigLIP-SO400M, MLP projector | 4096 | 32 | 729 |
| Qwen-VL-7B | Qwen-7B | OpenCLIP ViT-bigG, cross-attention | 4096 | 32 | 256 |
| InternVL3.5-8B | Qwen3-8B | InternViT-300M, pixel shuffle, MLP | 4096 | 36 | 256 |
| Qwen2.5-VL-7B | Qwen2.5-7B | Native-resolution ViT, patch merger | 3584 | 28 | 256 |
| Idefics3-8B | Llama-3.1-8B | SigLIP-SO400M, pixel shuffle, MLP | 4096 | 32 | 169 |
| InternVL3.5-8B | Qwen2.5-VL-7B | Idefics3-8B | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | # Params (M) | VizWiz | OK-VQA | # Params (M) | VizWiz | OK-VQA | # Params (M) | VizWiz | OK-VQA |
| Zero-shot | – | 27.7 | 19.4 | – | 48.1 | 54.5 | – | 29.9 | 45.2 |
| 4-shot ICL | – | 56.5 | 47.4 | – | 54.5 | 43.0 | – | 37.6 | 45.9 |
| LIVE | 0.15 ( 18.00) | 63.54 ±0.71 | 49.98 ±0.13 | 0.10 ( 14.00) | 52.97 ±9.08 | 55.31 ±0.33 | 0.13 ( 16.00) | 46.93 ±1.96 | 39.25 ±1.03 |
| MimIC | 0.30 ( 36.14) | 66.79 ±1.42 | 57.36 ±0.42 | 0.20 ( 28.11) | 73.44 ±0.87 | 64.55 ±0.47 | 0.26 ( 32.13) | 66.71 ±0.86 | 57.22 ±0.22 |
| HiFICL | 2.5 ( 306.00) | 68.09 ±0.41 | 58.74 ±0.17 | 1.7 ( 238.00) | 71.74 ±0.51 | 59.81 ±0.53 | 2.2 ( 272.00) | 65.60 ±0.77 | 56.53 ±0.64 |
| Method | 0 | 1 | 2 | 4 | 8 | mismatched | Avg. drop |
|---|---|---|---|---|---|---|---|
| LLaVA-Interleave, COCO | |||||||
| MimIC | 134.6 | 3.3 | 3.0 | 3.4 | 3.9 | 4.3 | -131.0 |
| HiFICL | 138.0 | 3.4 | 3.0 | 3.4 | 4.5 | 4.0 | -134.4 |
| STAVE (source-only) | 134.0 | 135.2 | 134.8 | 135.8 | 137.4 | 135.7 | +1.8 |
| STAVE (target-only) | 135.2 | 11.3 | 21.8 | 21.6 | 25.5 | 15.0 | -116.2 |
| STAVE (paired) | 138.8 | 133.9 | 134.7 | 135.5 | 134.4 | 133.0 | -4.5 |
| Objective | CHAIRs | CHAIRi | Recall |
|---|---|---|---|
| Paired | 2.62 | 1.76 | 44.81 |
| Target-only | 2.84 | 1.93 | 44.47 |
| Source-only | 2.36 | 1.64 | 43.08 |
| Pythia 6.9B | GPT-J 6B | LLaMA 7B | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | w/o demo | w/ 1 demo | w/o demo | w/ 1 demo | w/o demo | w/ 1 demo | |||
| SITE | 90.4 | 89.4 | 91.6 | 90.6 | 92.1 | 91.3 | |||
| STAVE-source | 66.6 | 92.0 | 82.8 | 92.3 | 89.6 | 92.0 | |||
| STAVE-target | 91.3 | 88.1 | 91.7 | 87.4 | 91.6 | 91.2 | |||
| STAVE-paired | 91.6 | 91.8 | 92.3 | 92.4 | 92.3 | 92.5 | |||
| COCO | VQAv2 | |||
| Configuration | Idefics2 | LLaVA | Idefics2 | LLaVA |
| One vector | ||||
| Context only | 132.65 | 129.57 | 72.84 | 75.51 |
| Readout only | 132.24 | 129.85 | 71.78 | 75.86 |
| All tokens | 131.95 | 129.18 | 72.75 | 75.74 |
| Two vectors | ||||
| Image tok. | Size (px) | One shared | Predicted | Readout / context | Ratio | Image share (%) |
|---|---|---|---|---|---|---|
| OK-VQA | ||||||
| 64 | 224 | 0.408 | 0.409 | 1.137 | 2.8 | 1.95 |
| 144 | 336 | 0.270 | 0.271 | 0.956 | 3.5 | 1.80 |
| 256 | 448 | 0.274 | 0.276 | 1.337 | 4.9 | 1.41 |
| 576 | 672 | 0.209 | 0.209 | 1.897 | 9.1 | 1.27 |
| 1,024 | 896 | 0.178 | 0.184 | 1.320 | 7.4 | 1.35 |
| Task | One shared vector (nats) | Readout / context | Random splits |
|---|---|---|---|
| VQAv2 | 0.113 | 7.52 | 0.99–1.00 |
| OK-VQA | 0.094 | 7.31 | 1.01–1.03 |
| COCO | 0.028 | 3.43 | 0.95–1.01 |
| VQAv2 | COCO | ||||
|---|---|---|---|---|---|
| Normalization | Scale | Idefics2 | LLaVA | Idefics2 | LLaVA |
| Vector count | 73.46 | 76.12 | 134.38 | 131.78 | |
| Group count | 72.36 | 75.73 | 133.14 | 129.10 | |
| Vector mean | 72.10 | 75.87 | 133.80 | 130.09 | |
| None | 67.61 | 74.96 | 132.98 | 130.23 | |
| Model | Color | Sport | Animal |
|---|---|---|---|
| Idefics2-8B-base | blue, white, black, brown | tennis, baseball, soccer, sk | cat, horse, gir, z |
| LLaVA | white, black, blue, red | base, ten, soc, sk | cat, g, dog, horse |
| Model | Food | Vehicle | Time or weather |
| Idefics2-8B-base | pizza, hot, sandwich, cake | bus, motor, truck, train | afternoon, winter, morning, sun |
| LLaVA | pizza, hot, grass, sand | bus, bike, motor, pickup | winter, day, summer, s |