cs.CLOct 6, 2026

Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

Authors: Xi Ding, Naichen Shi, Jiawei Zhang

Organizations: University of Wisconsin–Madison · Northwestern University

Abstract

In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning

    Jun 10, 2026Zhirui Chen, Ziwei Chen, Ling ShaoMultimodal Large Language ModelsIn-Context Learning

  2. Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning

    May 20, 2026Jihoon Kwon, Jiwon Choi, Jy-yong SohnIn-Context LearningDistributional Alignment