Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocates more resolution to query-relevant content while compressing less informative areas, all while preserving global context. At test time, the approach uses an MLLM's cross-modal attention to perform rectilinear warping of the input image, reallocating spatial resolution toward regions the model deems important, without changing model weights or architecture. This attention-guided warping preserves all original image information but redistributes it non-uniformly, so small objects and subtle relationships become easier for the same model to read while the global layout remains intact. Across five benchmarks (TextVQA, GQA, DocVQA, POPE, MMMU) and four MLLMs (LLaVA, Qwen-VL, InternVL, and InstructBLIP), AttWarp consistently improves accuracy, strengthens compositional reasoning, and reduces hallucinations, outperforming four competitive baselines that manipulate raw images at test time. Together, these results show that attention-guided warping prioritizes information relevant to the query while preserving context, and that the same MLLMs perform better when given such warped inputs.
Figures & tables
Figure 1: AttWarp overview. Given a query, our method extracts cross-modal attention maps from the MLLM’s language decoder, aggregates them into marginal attention profiles, and uses rectilinear warping to expand high-attention regions while compressing low-attention areas. The warped image is then processed by the same MLLM, which now produces the correct answer.
Figure 2: AttWarp-Distill architecture: CLIP vision tokens are FiLM - modulated by text and projected to 1D marginal predictors.
Figure 3: AttWarp improves compositional & spatial reasoning e.g . from GQA dataset (a) correctly identifying zebra behind the path, (b) television below the artwork; text understanding in documents e.g . from DocVQA (d) consumer focus groups, fine-grained recognition of small/occluded objects e.g . from POPE (f) detecting a bottle
Figure 4: AttWarp and prior works of image manipulation on the running example. While plausible, prior works are unable to answer the question correctly.
#
Methods
Key Technique
TextVQA
GQA
MMMU
POPE
DocVQA
LLaVA ( Liu et al., 2024a ) (MLP vision-language connector & open data)
1
Base MLLM
49.3
60.5
36.9
85.3
18.1
2
FGVP-mask ( Yang et al., 2023b )
Green mask overlay
39.4
59.2
36.1
85.3
19.0
3
FGVP-blur ( Yang et al., 2023b )
Blur background
33.9
59.5
35.0
83.1
18.6
4
SoM ( Yang et al., 2023a )
Grounded segments
18.8
54.5
35.6
78.5
15.8
5
API ( Yu et al., 2024 )
Alpha channel fade
49.9
60.6
36.9
85.9
17.4
Table 1: Main results on TextVQA, GQA, MMMU, POPE, and DocVQA datasets in accuracy (%). The Δ Accuracy row reports the absolute improvement of AttWarp-Chain over the base MLLM.
LLaVA
MMVP
BLINK
RealWorldQA
MIA
Base MLLM
48.3
38.3
49.3
65.9
AttWarp-Distilled (ours)
49.3
39.7
51.1
66.4
AttWarp (ours)
50.7
40.4
52.1
67.8
AttWarp-Chains (ours)
51.0
41.2
53.1
68.8
Δ Accuracy
+2.7
+2.9
+3.8
+2.9
Table 2: AttWarp performance (in %) on visual-centric benchmarks. We report individual accuracy for MMVP, as it aligns with the format of other evaluations. We use LLaVA as the base MLLM. For BLINK, we took the single-image subset.
Figure 5: AttWarp-Chain improves on AttWarp
TFLOPs ↓
Peak VRAM ↓
MLLM passes ↓
ViCrop
24.2
22
3
AttWarp-Distill
8.7 (0.4 × )
15
1
Base MLLM
8.5
15
1
Table 3: Comparison of computational overhead. Base MLLM used is LLaVA. Metrics are TFLOPs, peak VRAM (in GB), and number of MLLM passes. Values in brackets show relative cost compared to ViCrop.
Figure 6: ( 6(a) ) Mahalanobis distance to the train Gaussian. ( 6(b) ) Train → Test, AttWarp , Non-Rect. FID/KID summary. ( 6(c) – 6(d) ) Attention–redistribution GT alignment. See Appx. I for setup and additional plots.
Figure 7: Error Characterization and Failure Cases.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Layer-wise localisation quality of LLaVA-1.5-7B cross-attention maps on TextVQA Zhang et al. (2025a) images. Curves report Top-1 Hit Rate ( Zhang et al., 2016 ) (left) and AM@all ( Wang et al., 2020 ) (right) for absolute (blue) and relative (orange) attention.
Figure 9: Below (left to right): (a) Original image, zoomed in at yamaha (answer), (b) to (d) attention maps captured from layers 10, 16, 20 respectively. As can be seen, the attention localization improves drastically from layer 10 to 20, indicating improved query-specific spatial understanding.
LLaVA
Method
TextVQA
GQA
MMMU
POPE
DocVQA
Base MLLM
49.3
60.5
36.9
85.3
18.1
AttWarp (Layer 20)
58.1
63.7
40.4
87.5
25.5
AttWarp (Layer 22)
58.4
62.8
39.1
87.1
24.9
Qwen
Method
TextVQA
GQA
MMMU
POPE
DocVQA
Appendix
Table 4: Performance of AttWarp across different attention layers on LLaVA and Qwen backbones.
Aggregation Methodology
TextVQA Accuracy
Mean over all 32 heads (Ours)
58.1
Max-pooling across 32 heads (token-wise)
55.3
Mean over 8 randomly selected heads
54.6
Random single head (re-sampled per run)
51.9
Appendix
Table 5: Effect of attention head aggregation strategy on TextVQA accuracy (%). Our method of averaging all heads is the most effective and robust.
Table 15
Corruption
LLaVA
LLaVA + AttWarp
Impulse noise
36.8
40.4
Gaussian noise
37.6
41.0
Shot noise
36.0
39.8
Appendix
Table 8: Accuracy (%) of LLaVA and LLaVA+ AttWarp under different image corruptions, following the ImageNet-C protocol.
Figure 10: (top row) the original image, the adversarial perturbation designed to corrupt attention, and the resulting perturbed image in which the task - relevant region is shrunk; (bottom row) outputs of AttWarp at Depth 1–3, illustrating how our method progressively overcomes the interference to refocus on the relevant region.
Task Category
Dataset
LLaVA
AttWarp
Δ
Fine-grained perception
TextVQA (fine-grained)
49.9
59.6
+9.3
DocVQA (table/list/form/handwritten)
13.6
19.5
+5.9
Spatial reasoning
TextVQA (spatial)
54.5
64.9
+10.4
DocVQA (layout)
29.4
37.6
+8.2
MMMU (spatial)
38.2
44.8
+6.6
GQA (relation)
51.5
56.4
+4.9
Appendix
Table 9: Category-level improvements of AttWarp . Accuracy in %. Best results in bold .
Category
LLaVA
AttWarp
Relation
51.5
56.4 (+4.9)
Attribute
67.8
69.3 (+1.5)
Category
51.7
55.1 (+3.4)
Object
86.1
89.4 (+3.3)
Global
62.5
65.5 (+3.0)
Appendix
Table 10: Performance on GQA semantic categories. Accuracy in %.
Dataset
Category
LLaVA
AttWarp
Δ
TextVQA
Positional reasoning
54.5
64.9
+10.4
DocVQA
Layout (positional)
29.4
37.6
+8.2
MMMU
Object shape
36.6
40.7
+4.1
DocVQA
Diagram (object shape)
18.8
22.2
+3.4
Appendix
Table 11: Cross-dataset improvements on positional reasoning and object shape. Accuracy in %.
Category
E[Δlog∣detJ∣]
Attribute
0.19
Category
0.18
Global
0.16
Object
0.19
Relation
0.22
Fine-grained
0.26
Appendix
Table 12: Warping intensity by question type. Values are mean log-change of Jacobian determinants.
Category
LLaVA
ViCrop
AttWarp
Relation
51.5
51.9
56.4
Attribute
67.8
68.2
69.3
Category
51.7
52.0
55.1
Object
86.1
86.3
89.4
Global
62.5
63.4
65.5
Appendix
Table 13: Comparison with ViCrop on GQA semantic categories. Accuracy in %. Best results in bold .
Method
InstructBLIP
InternVL-3 8B
TextVQA
GQA
TextVQA
GQA
Baseline
35.2
49.4
80.2
61.4
ViCrop
46.6
49.7
82.7
63.9
AttWarp
47.8
51.3
84.6
65.9
Appendix
Table 14: Generalization of AttWarp across InstructBLIP and InternVL-3 . Evaluation metric: accuracy (%). Best results in bold .
Method
Overall
Free Text
Table
Layout
Form
Hand-written
Diagram
Others
Image
Y/N
Qwen
77.8
78.1
76.3
85.6
75.8
63.2
79.6
80.0
71.4
82.1
API
68.4
68.9
60.8
79.5
68.9
60.9
72.1
80.0
66.1
64.3
ViCrop
82.3
83.0
78.8
87.2
77.9
68.3
80.8
84.8
76.5
96.4
AttWarp
84.1
86.2
79.1
89.9
83.5
71.4
81.5
86.7
82.1
92.9
Appendix
Table 15: Category-wise accuracy (%) on DocVQA using Qwen2.5-VL-7B as the base MLLM. All results are reported on images resized to 512 × 512 to ensure computational feasibility across methods. AttWarp consistently outperforms all baselines across every document structure category except Yes/No, where ViCrop performs best. Gains span both fine-grained (Form, Hand-written, Table) and global (Layout, Diagram) reasoning categories.
Figure 11: Impact of iterative warping depth on accuracy for TextVQA and GQA . Fixed-length warping initially improves performance but degrades with excessive iterations due to recursive instability. The adaptive AttWarp-Chain (dashed lines), guided by the KL divergence stopping criterion, consistently achieves the best accuracy while avoiding over-warping.
Method (Source of Attention)
TextVQA
GQA
Base LLaVA
49.3
60.5
+ AttWarp (Internal: LLaVA)
58.1
63.7
+ AttWarp (Stable Diffusion Rombach et al. (2022) )
56.0
62.7
+ AttWarp (Qwen-VL Yang et al. (2024a) )
59.3
63.9
Appendix
Table 16: Performance of AttWarp using attention from internal vs. external models. Evaluation metric: accuracy (%). Best results in bold .
Dataset
% Boxes Expanded
Mean Area Increase
TextVQA (single-region)
94.0
+76%
gRef (multi-region)
88.6
+39%
Appendix
Table 17: Expansion of query-relevant bounding boxes under AttWarp .
Model
Original Images
Warped Images
Δ
LISA-LLaVA-7B
54.0
61.0
+7.0
Appendix
Table 18: OVOD performance with LISA-LLaVA-7B. Accuracy in %.
Figure 12: Qualitative OVOD results with LISA-LLaVA-7B. Each example shows the original, warped, and inverse-warped images. Predicted bounding boxes are in blue , ground-truth annotations in red .
Model
Without Warping
With Warping (7B maps)
Δ
LLaVA-1.6-34B
72.6
74.1
+1.5
Appendix
Table 19: TextVQA accuracy (%) of LLaVA-1.6-34B with and without image warping using attention maps from LLaVA-1.5-7B.
Figure 13: The marginal distributions predicted by the student networks closely align with the ground truth marginals, demonstrating robust knowledge transfer.
Figure 14: Comparison of Train → {Test, AttWarp , Non-Rectilinear} FID–KID across two ViT backbones.
When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast, current multimodal large language models (MLLMs) attend to all visual tokens at each generation step, leading to diluted focus and unnecessary computational overhead. In this work, we introduce Gaze Attention, a novel mechanism that enables MLLMs to selectively attend to task-relevant visual regions during generation. Specifically, we spatially group visual embeddings-stored as key-value caches-into compact gaze regions, each represented by a lightweight descriptor. At each decoding step, the model dynamically selects the most relevant regions and restricts attention to them, reducing redundant computation while enhancing focus. To mitigate the loss of global context caused by localized attention, we further propose learnable context tokens appended to each image or frame, allowing the model to maintain holistic visual awareness. Extensive experiments on image and video understanding benchmarks demonstrate that Gaze Attention matches or surpasses dense-attention baselines, while using up to 90% fewer visual KV entries in the attention computation.
Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal signature of hallucination: progressive degradation of text-to-image cross-attention during generation, leading to specific failure patterns like unfocused or biased attention. Existing mitigation strategies are largely outcome-driven and do not explicitly target this failure mode. To address this problem, we propose ADAPT (Attention Dynamics Alignment with Preference Tuning), an attention-based framework that intervenes directly on text-to-image cross-attention dynamics. We propose ADAPT with three key contributions: a cross-attention visual anchor refined from early decoding to provide stable spatial grounding, an attention-supervised inference mechanism that detects and corrects attention drift online, and a Visual Attention Guidance DPO that aligns preferences toward visually grounded responses. Experiments show that each component of ADAPT contributes to hallucination reduction, and the full framework achieves new best results across multiple hallucination benchmarks, reducing hallucination rates by 40%-60% across mainstream backbones while preserving general multimodal capabilities. Our work provides an attention-based perspective on mitigating hallucinations by exploring the model's internal text-to-image cross-attention behaviors. Code is available at https://github.com/yao-ustc/ADAPT
Zhiyuan Yao, Zheren Fu, Zhixiao Zheng +3
University of Science and Technology of China, Hefei, China · State Key Laboratory of Communication Content Cognition, People’s Daily Online, China · Huawei Technologies Ltd., China
Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding, often exhibiting spatial hallucination, attention leakage, and failures in object constancy. In addition, existing approaches typically rely on expensive human annotations or large-scale chain-of-thought (CoT) data generation. We propose Compositional Grounded Contrast (abbr. CGC), a low-cost full framework for boosting fine-grained multi-image understanding of MLLMs. Built on existing single-image grounding annotations, CGC constructs compositional multi-image training instances through Inter-Image Contrast and Intra-Image Contrast, which introduce semantically decoupled distractor contexts for cross-image discrimination and correlated cross-view samples for object constancy, respectively. CGC further introduces a Rule-Based Spatial Reward within the GRPO framework to improve source-image attribution, spatial alignment, and structured output validity under a Think-before-Grounding paradigm. Experiments show that CGC achieves state-of-the-art results on fine-grained multi-image benchmarks, including MIG-Bench and VLM2-Bench. The learned multi-image understanding capability also transfers to broader multimodal understanding and reasoning tasks, yielding consistent gains over the Qwen3-VL-8B base model on MathVista (+2.90), MuirBench (+2.88), MMStar (+1.93), MMMU (+1.77), and BLINK (+1.69).
Lihao Zheng, Zhenwei Shao, Yu Zhou +5
School of Computer Science and Technology, Hangzhou Dianzi University · Base Model, Li Auto