Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While this degradation has been observed in prior work, its nature remains poorly understood. We empirically observe visual elements from different images become entangled in the model's representations and responses. We refer to this phenomenon as cross-image information leakage. To address this issue, we propose FOCUS, a training-free and architecture-agnostic method. FOCUS masks all but one image with random noise, guiding the model to focus on the single clean image. This process is applied across the target images to obtain logits under partially masked contexts. These logits are aggregated and then refined using a noise-only reference input, which suppresses the leakage and yields more accurate outputs. FOCUS consistently improves performance on diverse multi-image benchmarks. We further show that FOCUS generalizes to video understanding, extending its applicability beyond static multi-image inputs. This demonstrates that FOCUS offers a general solution for enhancing multi-image reasoning without additional training or architectural modifications.
Figures & tables
Figure 2: Illustration of cross-image information leakage across different image pairs. Given the same target image (red border), the baseline LVLM consistently selects the incorrect merged caption that combines content from both images (Failure Cases), regardless of which non-target image is present. FOCUS mitigates this confusion and recovers the correct answer in both cases (Corrected Results).
Figure 3: Estimation of probability density function of Qwen2.5-VL 3B predictions over three answer options. We evaluated on 500 image pairs from VisMin awal2024vismin . In the single-image setting, the correct option is favored. In the multi-image setting, all distributions located their center near random chance (the dotted line).
Figure 4: Overview of FOCUS. It consists of three main steps: (a) visual masking, (b) image-wise focused inference, and (c) contrastive aggregation.
Table 1: Comparison on Winoground across model families and sizes with and without FOCUS . Metrics: T (Text), I (Image), G (Group). Bold indicates improvement over the baseline.
Table 2: Comparison on VisMin-Easy across model families and sizes with and without FOCUS . Metrics: T (Text), I (Image), G (Group). Bold indicates improvement over the baseline.
Table 3: Comparison on VisMin-Hard across model families and sizes with and without FOCUS . Metrics: T (Text), I (Image), G (Group). Bold indicates improvement over the baseline.
Table 4: Mantis-Eval, MuirBench and MIRB accuracy across model families and sizes, with and without FOCUS . Bold indicates improvement over the baseline.
Table 5: Results on Vinoground with and without FOCUS . Metrics: T (Text), V (Video), G (Group). Bold indicates improvement over the baseline.
w/o Noise ( λ )
w/o Reference ( α )
FOCUS
Acc.
71.43
61.90
76.19
Table 6: Ablation study on each component of FOCUS . We use Qwen2.5-VL-7B on Mantis-Eval validation set.
Figure 5: Average pairwise cosine similarity distribution of image token representations after the LLM decoder under three conditions. We use Qwen2.5-VL-3B on samples with 4 images each from Mantis-Eval.
Method
InternVL3
Qwen2.5-VL
LLaVA-OV
2B
8B
3B
7B
0.5B
7B
Process Independently
Per-Img
19.61
35.71
34.57
38.18
10.53
32.82
Process Jointly
Baseline
31.32
41.86
49.27
50.53
22.55
48.02
Conv-CD
31.68
41.07
49.64
51.29
20.43
44.79
Table 7: Comparison of FOCUS with alternative approaches on MIRB. Per-Img: per-image processing. Conv-CD: conventional contrastive decoding. SoFA: tian2025identifying .
Figure 6: Qualitative results on Mantis-Eval. We compare the Qwen2.5-VL 3B model with and without FOCUS. These samples require the model to interpret each image independently while also integrating information across all images to reach the correct answer. The baseline fails under this setting, however FOCUS preserves per-image understanding and produces the correct responses.
N
Qwen3-VL-4B
Qwen3-VL-8B
Baseline
+ FOCUS
Baseline
+ FOCUS
2
1.38 ± 2.82
2.12 ± 2.72
1.62 ± 2.91
2.20 ± 3.05
3
1.00 ± 1.48
2.64 ± 1.22
1.26 ± 1.78
2.67 ± 1.53
4
0.43 ± 0.34
1.35 ± 1.15
1.90 ± 4.26
2.47 ± 3.15
5
0.43 ± 0.00
1.03 ± 0.00
0.48 ± 0.00
1.04 ± 0.00
Acc.
66.82
68.66
65.44
68.66
Table 8: Latency (mean ± std, seconds) and accuracy on Mantis-Eval across different numbers of input images.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Var.
InternVL3
Qwen2.5-VL
LLaVA-OV
2B
8B
3B
7B
0.5B
7B
Winog.
λ
1
1
0.6
1
1
1
α
0.6
0.8
0.15
0.2
0.1
0.3
VisMin
λ
1
0.7
0.5
0.7
1
0.9
α
0.4
0.7
0.5
0.15
0.1
0.2
Mantis
λ
0.8
0.3
0.8
0.3
0.1
0.2
Appendix
Table C1: Hyperparameters used for evaluation. Search space: λ∈(0,1) and α∈(−1,1) .
Latency
Speedup
N
Baseline
FOCUS
vs. (N+1)×
2
1.87
1.96
2.86
4
1.93
2.58
3.73
8
2.11
4.97
3.81
16
2.47
14.26
2.94
Appendix
Table C2: Inference cost of FOCUS versus the number of input images N . Per-sample latency in seconds, measured on synthetic (512,512,3) tensors and averaged over three runs. The last column is the speedup of our batched implementation over the naive (N+1)× cost.
Figure D1: Estimation of probability density function of model predictions over three answer options, replicating the analysis of Figure 3 at larger model scales: Qwen2.5-VL-7B (top) and Qwen3-VL-8B (bottom). Accuracy consistently degrades under multi-image conditions across both scales, confirming that cross-image information leakage is not an artifact of smaller models.
Figure D2: Average pairwise cosine similarity distribution of image token representations under three conditions (single-image, multi-image, and + FOCUS ), replicating the analysis of Figure 5 at larger model scales: Qwen2.5-VL-7B (left) and Qwen3-VL-8B (right).
Benchmark
Metric
Baseline
+ FOCUS
VisMin
T
91.06
91.06
I
70.12
73.51
G
67.88
70.98
Mantis-Eval
Acc.
64.06
66.36
Appendix
Table D1: Results on Qwen3-VL-30B-A3B (MoE) on VisMin and Mantis-Eval. FOCUS consistently improves Image Score (+3.39 pp), Group Score (+3.10 pp), and Mantis-Eval accuracy (+2.30 pp), while Text Score remains unchanged.
Benchmark
Metric
Baseline
+ FOCUS
Winoground
T
81.50
82.50
I
54.75
70.50
G
49.75
65.50
VisMin
T
91.42
91.42
I
88.08
90.12
G
83.89
85.47
Appendix
Table D2: Results with Qwen3-VL-32B. T, I, and G denote Text, Image, and Group score. VisMin results are averaged over the Object, Attribute, and Counting categories.
Category
Dataset
Metric
Baseline
+ FOCUS
ICL
Cars
Acc
20.75
26.42
Food
Acc
24.00
34.00
iNat
Acc
18.00
22.00
DocVQA
LongDocURL
doc_qa
58.94
60.60
MMLongDoc
doc_qa
51.05
53.29
SlideVQA
doc_qa
65.75
68.68
Appendix
Table D3: Results on MMLongBench wang2026mmlongbench using Qwen2.5-VL-7B. FOCUS shows consistent gains on ICL and DocVQA tasks requiring cross-image comparison.
(a) Noise Type
(b) Noise Scale
(c) Weight
Variant
Acc.
λ
Acc.
α
Acc.
Gaussian
71.43
0.1
71.43
0.1
66.67
Impulse
66.67
0.3
76.19
0.4
76.19
Uniform
76.19
1.0
52.38
1.0
61.90
Appendix
Table D4: Impact of FOCUS design choices: noise type, masking strength λ , and contrastive weight α . Mantis validation accuracies using Qwen2.5-VL-7B. Bold indicates best per group.
Reference Type
Mantis-Eval (Test)
w/o Reference
69.12
Clean Image as Reference
67.28
FOCUS (Noise Reference)
70.05
Appendix
Table D5: Ablation on the choice of reference type, using Qwen2.5-VL-7B on the Mantis-Eval test set. The per-image noise term is held fixed and only the reference type varies.
Category
Object
Attribute
Relation
Counting
T
I
G
T
I
G
T
I
G
T
I
G
InternVL3-2B
95.68
36.10
35.92
90.82
23.81
22.79
60.77
1.29
0.96
70.29
12.39
11.38
+ FOCUS
95.85
78.58
76.86
90.82
70.07
65.65
60.13
10.93
8.36
69.61
42.61
39.56
InternVL3-8B
96.03
86.01
83.94
93.20
83.67
79.25
90.68
47.43
46.14
75.04
47.88
43.46
+ FOCUS
95.85
85.49
83.77
93.54
80.61
77.89
89.87
50.64
48.71
75.21
54.67
49.07
QwenVL2.5-3B
94.82
48.19
47.84
89.46
47.28
44.56
83.44
24.12
20.90
71.31
30.05
24.96
Appendix
Table D6: Performance across categories (Object, Attribute, Relation, Counting) with and without the proposed FOCUS decoding method. T: Text Score, I: Image Score, G: Group Score. Bold indicates improvement over the baseline. Underline highlights the largest gain within each model group for a given category.
Figure D3: Qualitative samples using the Qwen2.5-VL 3B model. In both multi-image settings, the baseline decoding strategy often produces a mixed information of the other image not indicated by the question. On the other hand, FOCUS disentangles cross-image information effectively, resulting in better multi-image understanding. These examples illustrate that FOCUS helps generate image-specific responses while suprressing cross-image information leakage.
Figure D4: More qualitative samples using the Qwen2.5-VL 3B model. Examples are from the Mantis benchmark. Figure shows that baseline decoding confuses cross-image differences, while FOCUS correctly identifies image-specific information.
OLIVES at the Center for Signal and Information Processing CSIP, School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA, USA