Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.
Figures & tables
Figure 1: Overview of the proposed Multimodal Causal Distillation (MCD) framework.
Figure 2: Causal graph of multimodal ICL inference. The query and ICDs provide complementary mechanism evidence and jointly shape the model state.
Model Family
Model / Method
VQAv2
VizWiz
MMStar
MathVision
MDK12
MMIQ
LogicVista
Avg.
LLaVA-OneVision
Teacher (72B)
83.61
71.87
61.05
27.32
46.57
28.03
32.14
50.08
Student (7B)
78.00
63.32
51.17
17.52
38.78
23.21
25.74
42.53
+Vanilla KD
81.29
67.42
52.38
18.46
40.17
24.17
27.53
44.49
+LLaVA-KD
82.61
69.23
54.73
18.77
41.33
24.14
28.95
45.68
+CompoDistill
81.57
69.45
53.19
17.95
40.63
23.81
28.42
45.00
+Align-TI
82.74
69.89
55.26
20.47
41.73
26.42
28.86
46.48
Table 1: Four-shot performance of different models on seven multimodal benchmarks. Bold denotes the best student result within each model family. Avg. denotes the average score across all seven benchmarks.
Figure 3: Performance under (a) different ICD selection strategies and (b) numbers of ICDs.
Variant
VQAv2
MMStar
MathVision
LogicVista
Avg.
Vanilla KD
84.59
67.39
44.67
46.97
60.91
Full MCD
87.29
73.01
52.37
52.67
66.34
w/o ℓkeep
86.85
71.65
49.83
51.23
64.89
w/o ℓeffect
86.52
71.26
49.25
50.42
64.36
w/o Verification
85.87
70.49
48.02
49.49
63.47
Table 2: Ablation of the causal objectives and verification procedure.
Variant
VQAv2
MMStar
MathVision
LogicVista
Avg.
Random evidence
84.93
68.52
47.27
48.33
62.26
Attention-based evidence
86.51
71.08
49.89
50.86
64.59
Unrestricted replacement
86.34
70.49
48.56
50.00
63.85
Full MCD
87.29
73.01
52.37
52.67
66.34
Table 3: Ablation of causal evidence discovery.
Figure 4: Causal behavior transfer and composition of verified evidence. (a) Evidence-sufficiency error, causal-response error, and Pearson correlation for the original student, Vanilla KD, and MCD. Lower errors and higher correlation indicate closer behavior to the teacher. (b) Composition of verified evidence across four benchmarks.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Sensitivity of MCD to three hyperparameters: evidence retention ratio r , verification threshold ρ , and causal loss weight λ . Results are averaged across the seven evaluation benchmarks. Dashed lines denote Vanilla KD as reference.
Variant
VQAv2
MMStar
MathVision
LogicVista
Avg.
Text-only
86.79
70.62
51.18
51.20
64.95
Vision-only
85.72
69.85
50.83
50.54
64.24
ICD-only
85.94
70.50
50.79
50.76
64.50
Query-only
86.45
70.28
51.31
51.42
64.87
Full MCD
87.29
73.01
52.37
52.67
66.34
Appendix
Table 4: Ablation of the modality and ICL prompt component receiving causal supervision.
Variant
VQAv2
MMStar
MathVision
LogicVista
Avg.
No verification
85.87
70.49
48.02
49.49
63.47
Retain-only
86.35
71.33
49.58
51.64
64.73
Remove-only
86.51
71.67
50.25
52.20
65.16
Joint verification
87.29
73.01
52.37
52.67
66.34
Appendix
Table 5: Decomposition of the evidence verification criterion. All verified variants retain the same proportion of training examples.
Sampling strategy
VQAv2
MMStar
MathVision
LogicVista
Avg.
Fixed-1 (default)
87.29
73.01
52.37
52.67
66.34
Fixed-2
87.14
73.13
52.31
53.12
66.43
Fixed-4
87.42
72.79
52.72
53.20
66.53
Epoch-wise resampling
86.69
72.31
51.22
51.45
65.42
Appendix
Table 6: Robustness to replacement sampling. Fixed- K averages teacher signals from K independently sampled matched replacements.
Method
Teacher processing
Teacher in training
Student evals./epoch
Student evals./3 epochs
Cache (GiB)
Inference cost
Vanilla KD
FT once
No
1.00FS
3.00FS
0.34
1.00×
LLaVA-KD
FT /epoch
Yes
1.00FS
3.00FS
–
1.00×
Align-TI
FT /epoch
Yes
2.00FS
6.00FS
–
1.00×
MCD
4FT+BTin once
No
1.75FS
5.25FS
0.80†
1.00×
Appendix
Table 7: Algorithmic efficiency on the fixed 60K-example training set. Counts are per example; the three-epoch column includes only student evaluations. Cache values use Kcache=128 , Lˉ=8 , and aˉ=0.5 . Inference cost is normalized by the original student.
Case type
Concrete multimodal ICL scenario
Why the case is challenging
Potential extension
Coherence-sensitive intervention
OCR strings, diagram labels, or relational statements must remain consistent with nearby visual content.
A structurally matched replacement creates an implausible local image-text pair and induces a response unrelated to the target mechanism.
Use context-conditioned replacements that preserve local semantic consistency.
Teacher uncertainty
Low-quality images, rare mechanisms, or open-ended questions admit uncertain or multiple valid predictions.
Evidence rankings and removal effects change across valid teacher responses, or the prompt is removed by correctness screening.
Use multi-response or ensemble verification to marginalize over acceptable teacher predictions.
Appendix
Table 8: Representative challenging cases and potential extensions of MCD in multimodal ICL.
While knowledge distillation (KD) is widely adopted for training lightweight models by leveraging supervision from larger teacher models, relying solely on output token distributions has proven insufficient for compressing Multimodal Large Language Models (MLLMs). Since output tokens are a byproduct of the model attending to visual inputs, prior works have explored explicitly distilling attention to provide a direct supervisory signal. While promising, the precise utility of which attention signals to distill remains under-explored. In this work, we challenge the conventional reliance on prompt-to-vision attention by revealing that downstream performance correlates strongly with response-to-vision attention similarity to the teacher, but negligibly with that of prompt-conditioned attention. Furthermore, we observe that attention distributions exhibit significant variance across individual tokens, indicating that a uniform distillation objective is suboptimal. To this end, we introduce Token-level Response-visual Attention Guidance (TRAG), a distillation objective that 1) shifts the focus to response-to-vision signals and 2) employs token-specific objectives by adaptively weighting the Kullback-Leibler divergence based on attention entropy, effectively guiding the student to mirror the teacher's precise visual focus. Extensive experimental results on multiple benchmarks demonstrate that TRAG significantly outperforms prior distillation baselines.
Jaehyun Jang, Eunseop Yoon, Hee Suk Yoon +3
Korea Advanced Institute of Science and Technology, Daejeon, Republic of Korea · University of Illinois Urbana-Champaign, Champaign, USA
Multimodal In-Context Learning (ICL) has emerged as a practical inference paradigm for Multimodal Large Language Models, where a small set of interleaved image-text In-Context Demonstrations (ICDs) conditions the model to solve new tasks. Despite its flexibility, multimodal ICL incurs high inference latency and suffers from instability due to sensitivity to demonstration formatting, ordering, and content. To address these limitations, we propose Hyper-ICL, a lightweight, training-based framework for demonstration-free multimodal ICL that reconstructs demonstration effects directly without requiring ICDs at inference time. Hyper-ICL learns a parameter-efficient low-rank logit-level adapter that calibrates attention distributions to better match demonstration-induced attention redistribution. To capture how demonstration influence varies across queries, we introduce a query-adaptive modulation mechanism that adaptively controls intervention strength at token level across layers and heads based on the current query. Finally, we propose a layer-wise hyperbolic anchor distillation loss that aligns intermediate student features to a demonstration-conditioned teacher via Lorentz geodesic distance. This loss encourages the student to reconstruct the demonstration-query relationships induced by ICDs. Extensive experiments across six different multimodal benchmarks (including VQAv2, OK-VQA, and COCO Caption) demonstrate that Hyper-ICL consistently improves accuracy and stability over vanilla ICL and existing state-of-the-art methods.
In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile. Our analysis reveals that the fundamental limitation lies in an inductive gap, models often produce correct answers from flawed reasoning, while struggling to extract consistent rules across demonstrations. This gap is further exacerbated by two visual-level obstacles: an overwhelming proportion of redundant visual tokens that obscure textual cues, and a skewed attention distribution that favors the initial image at the expense of subsequent context. To address these issues, we introduce a framework that restructures multimodal ICL as a principled inductive-deductive process. The framework incorporates a similarity-based visual token compression module to filter out redundant patches, a dynamic attention rebalancing mechanism to distribute focus equitably across all images, and a chain-of-thought paradigm that explicitly guides the model to analyze individual examples, derive a generalizable rule, and then apply it to the query. An auxiliary learning pipeline combines supervised fine-tuning with reinforcement learning using verifiable rewards to reinforce faithful citation and noise filtering. Evaluations across eight benchmarks covering visual perception, logical reasoning, STEM problems, and sarcasm detection demonstrate consistent and significant improvements over standard ICL baselines for multiple open-source VLMs, highlighting the potential of equipping models with genuine inductive capabilities in multimodal settings.
Haoyu Wang, Haonan Wang, Yuyan Chen +5
Tencent QQ · Fudan University · Cornell University