cs.CVSep 30, 2026

MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models

Authors: Yanshu Li, Jiaqian Li, Canran Xiao, Xi Xiao, Tianyang Wang, Yongtai Liu

Organizations: Brown University · University of Alabama at Birmingham · Hanyang University

Abstract

Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation

    Jul 1, 2026Jaehyun Jang, Eunseop Yoon, Hee Suk Yoon +3Knowledge DistillationFew-Step Distillation

  2. Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning

    Jun 3, 2026Niloufar Alipour Talemi, Hossein Kashiani, Fatemeh AfghahDataset Distillation

  3. Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning

    May 4, 2026Haoyu Wang, Haonan Wang, Yuyan Chen +5Multimodal ReasoningIn-Context Learning