Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocates more resolution to query-relevant content while compressing less informative areas, all while preserving global context. At test time, the approach uses an MLLM's cross-modal attention to perform rectilinear warping of the input image, reallocating spatial resolution toward regions the model deems important, without changing model weights or architecture. This attention-guided warping preserves all original image information but redistributes it non-uniformly, so small objects and subtle relationships become easier for the same model to read while the global layout remains intact. Across five benchmarks (TextVQA, GQA, DocVQA, POPE, MMMU) and four MLLMs (LLaVA, Qwen-VL, InternVL, and InstructBLIP), AttWarp consistently improves accuracy, strengthens compositional reasoning, and reduces hallucinations, outperforming four competitive baselines that manipulate raw images at test time. Together, these results show that attention-guided warping prioritizes information relevant to the query while preserving context, and that the same MLLMs perform better when given such warped inputs.
Figures & tables
Figure 1: AttWarp overview. Given a query, our method extracts cross-modal attention maps from the MLLM’s language decoder, aggregates them into marginal attention profiles, and uses rectilinear warping to expand high-attention regions while compressing low-attention areas. The warped image is then processed by the same MLLM, which now produces the correct answer.
Figure 2: AttWarp-Distill architecture: CLIP vision tokens are FiLM - modulated by text and projected to 1D marginal predictors.
Figure 3: AttWarp improves compositional & spatial reasoning e.g . from GQA dataset (a) correctly identifying zebra behind the path, (b) television below the artwork; text understanding in documents e.g . from DocVQA (d) consumer focus groups, fine-grained recognition of small/occluded objects e.g . from POPE (f) detecting a bottle
Figure 4: AttWarp and prior works of image manipulation on the running example. While plausible, prior works are unable to answer the question correctly.
#
Methods
Key Technique
TextVQA
GQA
MMMU
POPE
DocVQA
LLaVA ( Liu et al., 2024a ) (MLP vision-language connector & open data)
1
Base MLLM
49.3
60.5
36.9
85.3
18.1
2
FGVP-mask ( Yang et al., 2023b )
Green mask overlay
39.4
59.2
36.1
85.3
19.0
3
FGVP-blur ( Yang et al., 2023b )
Blur background
33.9
59.5
35.0
83.1
18.6
4
SoM ( Yang et al., 2023a )
Grounded segments
18.8
54.5
35.6
78.5
15.8
5
API ( Yu et al., 2024 )
Alpha channel fade
49.9
60.6
36.9
85.9
17.4
Table 1: Main results on TextVQA, GQA, MMMU, POPE, and DocVQA datasets in accuracy (%). The Δ Accuracy row reports the absolute improvement of AttWarp-Chain over the base MLLM.
LLaVA
MMVP
BLINK
RealWorldQA
MIA
Base MLLM
48.3
38.3
49.3
65.9
AttWarp-Distilled (ours)
49.3
39.7
51.1
66.4
AttWarp (ours)
50.7
40.4
52.1
67.8
AttWarp-Chains (ours)
51.0
41.2
53.1
68.8
Δ Accuracy
+2.7
+2.9
+3.8
+2.9
Table 2: AttWarp performance (in %) on visual-centric benchmarks. We report individual accuracy for MMVP, as it aligns with the format of other evaluations. We use LLaVA as the base MLLM. For BLINK, we took the single-image subset.
Figure 5: AttWarp-Chain improves on AttWarp
TFLOPs ↓
Peak VRAM ↓
MLLM passes ↓
ViCrop
24.2
22
3
AttWarp-Distill
8.7 (0.4 × )
15
1
Base MLLM
8.5
15
1
Table 3: Comparison of computational overhead. Base MLLM used is LLaVA. Metrics are TFLOPs, peak VRAM (in GB), and number of MLLM passes. Values in brackets show relative cost compared to ViCrop.
Figure 6: ( 6(a) ) Mahalanobis distance to the train Gaussian. ( 6(b) ) Train → Test, AttWarp , Non-Rect. FID/KID summary. ( 6(c) – 6(d) ) Attention–redistribution GT alignment. See Appx. I for setup and additional plots.
Figure 7: Error Characterization and Failure Cases.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Layer-wise localisation quality of LLaVA-1.5-7B cross-attention maps on TextVQA Zhang et al. (2025a) images. Curves report Top-1 Hit Rate ( Zhang et al., 2016 ) (left) and AM@all ( Wang et al., 2020 ) (right) for absolute (blue) and relative (orange) attention.
Figure 9: Below (left to right): (a) Original image, zoomed in at yamaha (answer), (b) to (d) attention maps captured from layers 10, 16, 20 respectively. As can be seen, the attention localization improves drastically from layer 10 to 20, indicating improved query-specific spatial understanding.
LLaVA
Method
TextVQA
GQA
MMMU
POPE
DocVQA
Base MLLM
49.3
60.5
36.9
85.3
18.1
AttWarp (Layer 20)
58.1
63.7
40.4
87.5
25.5
AttWarp (Layer 22)
58.4
62.8
39.1
87.1
24.9
Qwen
Method
TextVQA
GQA
MMMU
POPE
DocVQA
Appendix
Table 4: Performance of AttWarp across different attention layers on LLaVA and Qwen backbones.
Aggregation Methodology
TextVQA Accuracy
Mean over all 32 heads (Ours)
58.1
Max-pooling across 32 heads (token-wise)
55.3
Mean over 8 randomly selected heads
54.6
Random single head (re-sampled per run)
51.9
Appendix
Table 5: Effect of attention head aggregation strategy on TextVQA accuracy (%). Our method of averaging all heads is the most effective and robust.
Table 15
Corruption
LLaVA
LLaVA + AttWarp
Impulse noise
36.8
40.4
Gaussian noise
37.6
41.0
Shot noise
36.0
39.8
Appendix
Table 8: Accuracy (%) of LLaVA and LLaVA+ AttWarp under different image corruptions, following the ImageNet-C protocol.
Figure 10: (top row) the original image, the adversarial perturbation designed to corrupt attention, and the resulting perturbed image in which the task - relevant region is shrunk; (bottom row) outputs of AttWarp at Depth 1–3, illustrating how our method progressively overcomes the interference to refocus on the relevant region.
Task Category
Dataset
LLaVA
AttWarp
Δ
Fine-grained perception
TextVQA (fine-grained)
49.9
59.6
+9.3
DocVQA (table/list/form/handwritten)
13.6
19.5
+5.9
Spatial reasoning
TextVQA (spatial)
54.5
64.9
+10.4
DocVQA (layout)
29.4
37.6
+8.2
MMMU (spatial)
38.2
44.8
+6.6
GQA (relation)
51.5
56.4
+4.9
Appendix
Table 9: Category-level improvements of AttWarp . Accuracy in %. Best results in bold .
Category
LLaVA
AttWarp
Relation
51.5
56.4 (+4.9)
Attribute
67.8
69.3 (+1.5)
Category
51.7
55.1 (+3.4)
Object
86.1
89.4 (+3.3)
Global
62.5
65.5 (+3.0)
Appendix
Table 10: Performance on GQA semantic categories. Accuracy in %.
Dataset
Category
LLaVA
AttWarp
Δ
TextVQA
Positional reasoning
54.5
64.9
+10.4
DocVQA
Layout (positional)
29.4
37.6
+8.2
MMMU
Object shape
36.6
40.7
+4.1
DocVQA
Diagram (object shape)
18.8
22.2
+3.4
Appendix
Table 11: Cross-dataset improvements on positional reasoning and object shape. Accuracy in %.
Category
E[Δlog∣detJ∣]
Attribute
0.19
Category
0.18
Global
0.16
Object
0.19
Relation
0.22
Fine-grained
0.26
Appendix
Table 12: Warping intensity by question type. Values are mean log-change of Jacobian determinants.
Category
LLaVA
ViCrop
AttWarp
Relation
51.5
51.9
56.4
Attribute
67.8
68.2
69.3
Category
51.7
52.0
55.1
Object
86.1
86.3
89.4
Global
62.5
63.4
65.5
Appendix
Table 13: Comparison with ViCrop on GQA semantic categories. Accuracy in %. Best results in bold .
Method
InstructBLIP
InternVL-3 8B
TextVQA
GQA
TextVQA
GQA
Baseline
35.2
49.4
80.2
61.4
ViCrop
46.6
49.7
82.7
63.9
AttWarp
47.8
51.3
84.6
65.9
Appendix
Table 14: Generalization of AttWarp across InstructBLIP and InternVL-3 . Evaluation metric: accuracy (%). Best results in bold .
Method
Overall
Free Text
Table
Layout
Form
Hand-written
Diagram
Others
Image
Y/N
Qwen
77.8
78.1
76.3
85.6
75.8
63.2
79.6
80.0
71.4
82.1
API
68.4
68.9
60.8
79.5
68.9
60.9
72.1
80.0
66.1
64.3
ViCrop
82.3
83.0
78.8
87.2
77.9
68.3
80.8
84.8
76.5
96.4
AttWarp
84.1
86.2
79.1
89.9
83.5
71.4
81.5
86.7
82.1
92.9
Appendix
Table 15: Category-wise accuracy (%) on DocVQA using Qwen2.5-VL-7B as the base MLLM. All results are reported on images resized to 512 × 512 to ensure computational feasibility across methods. AttWarp consistently outperforms all baselines across every document structure category except Yes/No, where ViCrop performs best. Gains span both fine-grained (Form, Hand-written, Table) and global (Layout, Diagram) reasoning categories.
Figure 11: Impact of iterative warping depth on accuracy for TextVQA and GQA . Fixed-length warping initially improves performance but degrades with excessive iterations due to recursive instability. The adaptive AttWarp-Chain (dashed lines), guided by the KL divergence stopping criterion, consistently achieves the best accuracy while avoiding over-warping.
Method (Source of Attention)
TextVQA
GQA
Base LLaVA
49.3
60.5
+ AttWarp (Internal: LLaVA)
58.1
63.7
+ AttWarp (Stable Diffusion Rombach et al. (2022) )
56.0
62.7
+ AttWarp (Qwen-VL Yang et al. (2024a) )
59.3
63.9
Appendix
Table 16: Performance of AttWarp using attention from internal vs. external models. Evaluation metric: accuracy (%). Best results in bold .
Dataset
% Boxes Expanded
Mean Area Increase
TextVQA (single-region)
94.0
+76%
gRef (multi-region)
88.6
+39%
Appendix
Table 17: Expansion of query-relevant bounding boxes under AttWarp .
Model
Original Images
Warped Images
Δ
LISA-LLaVA-7B
54.0
61.0
+7.0
Appendix
Table 18: OVOD performance with LISA-LLaVA-7B. Accuracy in %.
Figure 12: Qualitative OVOD results with LISA-LLaVA-7B. Each example shows the original, warped, and inverse-warped images. Predicted bounding boxes are in blue , ground-truth annotations in red .
Model
Without Warping
With Warping (7B maps)
Δ
LLaVA-1.6-34B
72.6
74.1
+1.5
Appendix
Table 19: TextVQA accuracy (%) of LLaVA-1.6-34B with and without image warping using attention maps from LLaVA-1.5-7B.
Figure 13: The marginal distributions predicted by the student networks closely align with the ground truth marginals, demonstrating robust knowledge transfer.
Figure 14: Comparison of Train → {Test, AttWarp , Non-Rectilinear} FID–KID across two ViT backbones.
University of Science and Technology of China, Hefei, China · State Key Laboratory of Communication Content Cognition, People’s Daily Online, China · Huawei Technologies Ltd., China