MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models
Organizations: Brown University · University of Alabama at Birmingham · Hanyang University
Abstract
Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.
Figures & tables
| Model Family | Model / Method | VQAv2 | VizWiz | MMStar | MathVision | MDK12 | MMIQ | LogicVista | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| LLaVA-OneVision | Teacher (72B) | 83.61 | 71.87 | 61.05 | 27.32 | 46.57 | 28.03 | 32.14 | 50.08 |
| Student (7B) | 78.00 | 63.32 | 51.17 | 17.52 | 38.78 | 23.21 | 25.74 | 42.53 | |
| +Vanilla KD | 81.29 | 67.42 | 52.38 | 18.46 | 40.17 | 24.17 | 27.53 | 44.49 | |
| +LLaVA-KD | 82.61 | 69.23 | 54.73 | 18.77 | 41.33 | 24.14 | 28.95 | 45.68 | |
| +CompoDistill | 81.57 | 69.45 | 53.19 | 17.95 | 40.63 | 23.81 | 28.42 | 45.00 | |
| +Align-TI | 82.74 | 69.89 | 55.26 | 20.47 | 41.73 | 26.42 | 28.86 | 46.48 |
| Variant | VQAv2 | MMStar | MathVision | LogicVista | Avg. |
|---|---|---|---|---|---|
| Vanilla KD | 84.59 | 67.39 | 44.67 | 46.97 | 60.91 |
| Full MCD | 87.29 | 73.01 | 52.37 | 52.67 | 66.34 |
| w/o | 86.85 | 71.65 | 49.83 | 51.23 | 64.89 |
| w/o | 86.52 | 71.26 | 49.25 | 50.42 | 64.36 |
| w/o Verification | 85.87 | 70.49 | 48.02 | 49.49 | 63.47 |
| Variant | VQAv2 | MMStar | MathVision | LogicVista | Avg. |
|---|---|---|---|---|---|
| Random evidence | 84.93 | 68.52 | 47.27 | 48.33 | 62.26 |
| Attention-based evidence | 86.51 | 71.08 | 49.89 | 50.86 | 64.59 |
| Unrestricted replacement | 86.34 | 70.49 | 48.56 | 50.00 | 63.85 |
| Full MCD | 87.29 | 73.01 | 52.37 | 52.67 | 66.34 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Variant | VQAv2 | MMStar | MathVision | LogicVista | Avg. |
|---|---|---|---|---|---|
| Text-only | 86.79 | 70.62 | 51.18 | 51.20 | 64.95 |
| Vision-only | 85.72 | 69.85 | 50.83 | 50.54 | 64.24 |
| ICD-only | 85.94 | 70.50 | 50.79 | 50.76 | 64.50 |
| Query-only | 86.45 | 70.28 | 51.31 | 51.42 | 64.87 |
| Full MCD | 87.29 | 73.01 | 52.37 | 52.67 | 66.34 |
| Variant | VQAv2 | MMStar | MathVision | LogicVista | Avg. |
|---|---|---|---|---|---|
| No verification | 85.87 | 70.49 | 48.02 | 49.49 | 63.47 |
| Retain-only | 86.35 | 71.33 | 49.58 | 51.64 | 64.73 |
| Remove-only | 86.51 | 71.67 | 50.25 | 52.20 | 65.16 |
| Joint verification | 87.29 | 73.01 | 52.37 | 52.67 | 66.34 |
| Sampling strategy | VQAv2 | MMStar | MathVision | LogicVista | Avg. |
|---|---|---|---|---|---|
| Fixed-1 (default) | 87.29 | 73.01 | 52.37 | 52.67 | 66.34 |
| Fixed-2 | 87.14 | 73.13 | 52.31 | 53.12 | 66.43 |
| Fixed-4 | 87.42 | 72.79 | 52.72 | 53.20 | 66.53 |
| Epoch-wise resampling | 86.69 | 72.31 | 51.22 | 51.45 | 65.42 |
| Method | Teacher processing | Teacher in training | Student evals./epoch | Student evals./3 epochs | Cache (GiB) | Inference cost |
|---|---|---|---|---|---|---|
| Vanilla KD | once | No | ||||
| LLaVA-KD | /epoch | Yes | – | |||
| Align-TI | /epoch | Yes | – | |||
| MCD | once | No |
| Case type | Concrete multimodal ICL scenario | Why the case is challenging | Potential extension |
|---|---|---|---|
| Coherence-sensitive intervention | OCR strings, diagram labels, or relational statements must remain consistent with nearby visual content. | A structurally matched replacement creates an implausible local image-text pair and induces a response unrelated to the target mechanism. | Use context-conditioned replacements that preserve local semantic consistency. |
| Teacher uncertainty | Low-quality images, rare mechanisms, or open-ended questions admit uncertain or multiple valid predictions. | Evidence rankings and removal effects change across valid teacher responses, or the prompt is removed by correctness screening. | Use multi-response or ensemble verification to marginalize over acceptable teacher predictions. |